AI/ML arXiv cs.AI

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

EBTE is a mediation layer for AI agents that converts rationales into typed action claims to be verified against server-side facts before execution.

Cybersecurity arXiv cs.AI

AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis

An analysis of AI deployment in the public sector reveals a 'speed asymmetry' and identifies governance gaps in existing frameworks like NIST and ISO.

AI/ML arXiv cs.AI

ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning

ODYSSE introduces Episode-wise GRPO (ESPO) to improve personalized agentic reasoning by optimizing over long action horizons rather than individual steps.

Cybersecurity arXiv cs.AI

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response

A review of cyber-capable AI agents identifies five vulnerability classes and discusses containment strategies to prevent agents from escaping sandboxes.

AI/ML arXiv cs.AI

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

HANDBOOK.md is a new benchmark for long-context agentic instruction following, testing if agents can adhere to complex company policies over extended tool-use horizons.

AI/ML arXiv cs.AI

How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation

Researchers developed a multi-agent crowd simulation where LLMs manage internal affective states and outward expressions to study emergent emotional contagion.

AI/ML arXiv cs.AI

Inferring Missing Trajectory Data with Temporal Convolutional Networks

A new trajectory inpainting method using Temporal Convolutional Networks (TCN) with symmetric dilation is proposed to reconstruct missing segments in trajectory data.

AI/ML arXiv cs.AI

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

The study identifies 'progress mirage' in autonomous LLM agents, where self-evaluation bias leads agents to mistake stagnation for progress without externally grounded verification.

AI/ML arXiv cs.AI

PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention

PreDiff-LM introduces a hybrid attention mechanism to adapt pretrained autoregressive transformers for discrete masked diffusion language modeling.

AI/ML arXiv cs.AI

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

Research shows that while awareness of AI sycophancy reduces a chatbot's perceived objectivity and enjoyment, it does not reduce its persuasiveness.

AI/ML arXiv cs.AI

Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection

The authors introduce DebtBench, a persona-enriched benchmark for debt collection negotiation, and DebtGPT, an agent optimized for financial recovery and user experience.

AI/ML arXiv cs.AI

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

CADENCE is a framework that uses sparse autoencoders to decompose ECG foundation model embeddings into interpretable physiological cardiac atoms.

AI/ML arXiv cs.AI

The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape

This paper explores 'agentic recommendation markets' where LLM agents act on behalf of users, shifting competition from platform-centric to user-centric recommendation.

AI/ML arXiv cs.AI

Many-body Tipping Dynamics of ChatGPT-like AIs

The paper analyzes 'tipping' in LLMs—where they drift into harmful or repetitive content—as a dynamical first passage process driven by many-body token interactions.

Hardware/Chips arXiv cs.AI

ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design

ContractHIL-HLS is a multi-agent workflow for high-level synthesis (HLS) that uses structured contracts and hardware-in-the-loop feedback to improve FPGA design closure.

Other Hacker News

60 Years Ago, a Submerged Submarine Circled the Globe for the First Time (2020)

A retrospective on the first submerged submarine to circle the globe, occurring 60 years prior to 2020.

AI/ML arXiv cs.AI

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

Research indicates that the final window of pretraining data significantly influences how a model reacts to subsequent alignment and preference optimization.

AI/ML arXiv cs.AI

Addressable Recall Compaction for Long Context-Window Control in AI Agents

Introduces ARC (Addressable Recall Compaction), a framework that improves long-context window control in AI agents by using ID-addressable logs for tool observations.

AI/ML arXiv cs.AI

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

Proposes 'Value Router', a system for cost-aware routing in recommenders that balances difficulty and business value when deciding whether to call an LLM.

AI/ML arXiv cs.AI

Towards an Agent Operating System - Lessons from Classical and Cloud OS

Argues for the creation of a standardized 'Agent Operating System' with stable abstractions, drawing parallels to POSIX and Kubernetes.