All Articles
16236 articles total
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm
IMACS decouples the organization, coordination, and collaboration protocols in multi-agent LLM systems to allow independent optimization and learning of team structures.
TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation
TRWH is a Graph Neural Network framework that combines LLM-generated text profiles with random walk augmentation to improve sparse recommendation systems.
Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization
A new similarity-driven optimization framework for automated line generalization in cartography, balancing spatial similarity with readability constraints.
Finding Optimal Cost-Bounded Plan Reductions: Refined Model
A refined Integer Linear Programming model for finding optimal cost-bounded subplans while preserving the original action order and executability.
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
PatientAgentBench is a new benchmark for evaluating patient-facing health AI agents, revealing significant clinical gaps in triage and safety across current frontier models.
Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
Researchers discover a 'KNOWS/DOES split' in instruction-tuned LLMs, where models can accurately describe a distribution but fail to sample from it, collapsing to a single output.
Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
A new dual-stream image-to-video architecture improves fluid dynamics in video generation by using a simulation dataset and optical-flow supervision.
From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning
PMRD is introduced as a multimodal zero-shot drug representation learning framework that separates mechanism-consistent factors from modality-specific noise.
Dual-Domain Manifold Modeling for Hyperspectral Image Fusion
The DDMM framework utilizes a Topology-Aware Transformer and frequency-decoupled fusion to improve hyperspectral image fusion and spatial structure preservation.
Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
Cardiologent is a multi-agent clinical decision support system for arrhythmia assessment that integrates ECG and PPG signals with clinical guidelines.
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
EBTE is a mediation layer for AI agents that converts rationales into typed action claims to be verified against server-side facts before execution.
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
An analysis of AI deployment in the public sector reveals a 'speed asymmetry' and identifies governance gaps in existing frameworks like NIST and ISO.
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
ODYSSE introduces Episode-wise GRPO (ESPO) to improve personalized agentic reasoning by optimizing over long action horizons rather than individual steps.
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
A review of cyber-capable AI agents identifies five vulnerability classes and discusses containment strategies to prevent agents from escaping sandboxes.
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
HANDBOOK.md is a new benchmark for long-context agentic instruction following, testing if agents can adhere to complex company policies over extended tool-use horizons.
How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
Researchers developed a multi-agent crowd simulation where LLMs manage internal affective states and outward expressions to study emergent emotional contagion.
Inferring Missing Trajectory Data with Temporal Convolutional Networks
A new trajectory inpainting method using Temporal Convolutional Networks (TCN) with symmetric dilation is proposed to reconstruct missing segments in trajectory data.
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
The study identifies 'progress mirage' in autonomous LLM agents, where self-evaluation bias leads agents to mistake stagnation for progress without externally grounded verification.
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
PreDiff-LM introduces a hybrid attention mechanism to adapt pretrained autoregressive transformers for discrete masked diffusion language modeling.
Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
Research shows that while awareness of AI sycophancy reduces a chatbot's perceived objectivity and enjoyment, it does not reduce its persuasiveness.