All Articles
16246 articles total
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
EBTE is a mediation layer for AI agents that converts rationales into typed action claims to be verified against server-side facts before execution.
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
An analysis of AI deployment in the public sector reveals a 'speed asymmetry' and identifies governance gaps in existing frameworks like NIST and ISO.
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
ODYSSE introduces Episode-wise GRPO (ESPO) to improve personalized agentic reasoning by optimizing over long action horizons rather than individual steps.
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
A review of cyber-capable AI agents identifies five vulnerability classes and discusses containment strategies to prevent agents from escaping sandboxes.
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
HANDBOOK.md is a new benchmark for long-context agentic instruction following, testing if agents can adhere to complex company policies over extended tool-use horizons.
How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
Researchers developed a multi-agent crowd simulation where LLMs manage internal affective states and outward expressions to study emergent emotional contagion.
Inferring Missing Trajectory Data with Temporal Convolutional Networks
A new trajectory inpainting method using Temporal Convolutional Networks (TCN) with symmetric dilation is proposed to reconstruct missing segments in trajectory data.
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
The study identifies 'progress mirage' in autonomous LLM agents, where self-evaluation bias leads agents to mistake stagnation for progress without externally grounded verification.
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
PreDiff-LM introduces a hybrid attention mechanism to adapt pretrained autoregressive transformers for discrete masked diffusion language modeling.
Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
Research shows that while awareness of AI sycophancy reduces a chatbot's perceived objectivity and enjoyment, it does not reduce its persuasiveness.
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection
The authors introduce DebtBench, a persona-enriched benchmark for debt collection negotiation, and DebtGPT, an agent optimized for financial recovery and user experience.
CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models
CADENCE is a framework that uses sparse autoencoders to decompose ECG foundation model embeddings into interpretable physiological cardiac atoms.
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape
This paper explores 'agentic recommendation markets' where LLM agents act on behalf of users, shifting competition from platform-centric to user-centric recommendation.
Many-body Tipping Dynamics of ChatGPT-like AIs
The paper analyzes 'tipping' in LLMs—where they drift into harmful or repetitive content—as a dynamical first passage process driven by many-body token interactions.
ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design
ContractHIL-HLS is a multi-agent workflow for high-level synthesis (HLS) that uses structured contracts and hardware-in-the-loop feedback to improve FPGA design closure.
60 Years Ago, a Submerged Submarine Circled the Globe for the First Time (2020)
A retrospective on the first submerged submarine to circle the globe, occurring 60 years prior to 2020.
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
Research indicates that the final window of pretraining data significantly influences how a model reacts to subsequent alignment and preference optimization.
Addressable Recall Compaction for Long Context-Window Control in AI Agents
Introduces ARC (Addressable Recall Compaction), a framework that improves long-context window control in AI agents by using ID-addressable logs for tool observations.
How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness
Proposes 'Value Router', a system for cost-aware routing in recommenders that balances difficulty and business value when deciding whether to call an LLM.
Towards an Agent Operating System - Lessons from Classical and Cloud OS
Argues for the creation of a standardized 'Agent Operating System' with stable abstractions, drawing parallels to POSIX and Kubernetes.