AI/ML arXiv cs.AI

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

IMACS decouples the organization, coordination, and collaboration protocols in multi-agent LLM systems to allow independent optimization and learning of team structures.

AI/ML arXiv cs.AI

TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation

TRWH is a Graph Neural Network framework that combines LLM-generated text profiles with random walk augmentation to improve sparse recommendation systems.

Other arXiv cs.AI

Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization

A new similarity-driven optimization framework for automated line generalization in cartography, balancing spatial similarity with readability constraints.

Software Engineering arXiv cs.AI

Finding Optimal Cost-Bounded Plan Reductions: Refined Model

A refined Integer Linear Programming model for finding optimal cost-bounded subplans while preserving the original action order and executability.

AI/ML arXiv cs.AI

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents

PatientAgentBench is a new benchmark for evaluating patient-facing health AI agents, revealing significant clinical gaps in triage and safety across current frontier models.

AI/ML arXiv cs.AI

Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe

Researchers discover a 'KNOWS/DOES split' in instruction-tuned LLMs, where models can accurately describe a distribution but fail to sample from it, collapsing to a single output.

AI/ML arXiv cs.AI

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

A new dual-stream image-to-video architecture improves fluid dynamics in video generation by using a simulation dataset and optical-flow supervision.

AI/ML arXiv cs.AI

From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning

PMRD is introduced as a multimodal zero-shot drug representation learning framework that separates mechanism-consistent factors from modality-specific noise.

AI/ML arXiv cs.AI

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

The DDMM framework utilizes a Topology-Aware Transformer and frequency-decoupled fusion to improve hyperspectral image fusion and spatial structure preservation.

AI/ML arXiv cs.AI

Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management

Cardiologent is a multi-agent clinical decision support system for arrhythmia assessment that integrates ECG and PPG signals with clinical guidelines.

AI/ML arXiv cs.AI

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

EBTE is a mediation layer for AI agents that converts rationales into typed action claims to be verified against server-side facts before execution.

Cybersecurity arXiv cs.AI

AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis

An analysis of AI deployment in the public sector reveals a 'speed asymmetry' and identifies governance gaps in existing frameworks like NIST and ISO.

AI/ML arXiv cs.AI

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

ODYSSE introduces Episode-wise GRPO (ESPO) to improve personalized agentic reasoning by optimizing over long action horizons rather than individual steps.

Cybersecurity arXiv cs.AI

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

A review of cyber-capable AI agents identifies five vulnerability classes and discusses containment strategies to prevent agents from escaping sandboxes.

AI/ML arXiv cs.AI

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

HANDBOOK.md is a new benchmark for long-context agentic instruction following, testing if agents can adhere to complex company policies over extended tool-use horizons.

AI/ML arXiv cs.AI

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

Researchers developed a multi-agent crowd simulation where LLMs manage internal affective states and outward expressions to study emergent emotional contagion.

AI/ML arXiv cs.AI

Inferring Missing Trajectory Data with Temporal Convolutional Networks

A new trajectory inpainting method using Temporal Convolutional Networks (TCN) with symmetric dilation is proposed to reconstruct missing segments in trajectory data.

AI/ML arXiv cs.AI

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

The study identifies 'progress mirage' in autonomous LLM agents, where self-evaluation bias leads agents to mistake stagnation for progress without externally grounded verification.

AI/ML arXiv cs.AI

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

PreDiff-LM introduces a hybrid attention mechanism to adapt pretrained autoregressive transformers for discrete masked diffusion language modeling.

AI/ML arXiv cs.AI

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

Research shows that while awareness of AI sycophancy reduces a chatbot's perceived objectivity and enjoyment, it does not reduce its persuasiveness.