AI/ML arXiv cs.AI

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems

The article explores semantic context drift in reasoning-class LLMs and proposes a mathematical model to maintain operator control stability in human-machine decision systems.

AI/ML arXiv cs.AI

Agentic Context Learning with Self-Discovered Specification

Research on context learning in LLMs reveals that acquiring local specifications is more critical than content acquisition, introducing the PSCI intervention to improve performance.

AI/ML arXiv cs.AI

Exploring Agentic Workflows for Generating High Quality Math Visual Aids

Proposes an agentic workflow for creating high-quality K-12 math visual aids through an iterative self-improvement loop using VLMs.

AI/ML arXiv cs.AI

TopoExplore: Topological Discrimination for Archive-Based Exploration

Introduces TopoExplore, a topology-aware exploration method for RL agents that improves efficiency in environments with enclosed structures.

AI/ML arXiv cs.AI

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Presents Who&When Pro, a large-scale benchmark designed to evaluate the ability of LLMs to accurately attribute failures in agentic systems.

AI/ML arXiv cs.AI

A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

Introduces a trace-supervised symbolic neural CPU that allows for interpretable program execution and quantization-simulated writeback.

AI/ML arXiv cs.AI

AgentAbstain: Do LLM Agents Know When Not to Act?

Introduces AgentAbstain, a benchmark and automated pipeline (AbstainGen) to evaluate whether LLM agents know when to abstain from taking irreversible actions.

AI/ML arXiv cs.AI

From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability

Discusses a formal mathematical approach to semantic caching and governed reuse classes to replace similarity-based heuristics in conversational demands.

AI/ML arXiv cs.AI

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

Introduces MAG, a benchmark and harness for multimodal web-agent task execution and guide generation using screenshots.

Software Engineering Hacker News

YouTrackDB is a general-use object-oriented graph database

YouTrackDB is introduced as a general-purpose object-oriented graph database.

AI/ML arXiv cs.AI

LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules

LegalFarePlan is a new route-planning framework for urban rail that handles non-additive fare rules and provides explainable plans.

AI/ML arXiv cs.AI

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

BatteryLake is an agentic, physics-grounded data lakehouse for curating heterogeneous battery aging data, including an open benchmark of 41 datasets.

AI/ML arXiv cs.AI

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm

Research on the 'budgeted placement' of strong corrector agents in a swarm of weak agents to optimize consensus correctness versus cost.

AI/ML arXiv cs.AI

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

A study on robust norm enforcement mechanisms for AI agents to prevent exploitative behavior in multi-agent systems using reliability estimates and escalating penalties.

AI/ML arXiv cs.AI

Verification of Adaptive Agentic Controllers through Finite Rule Revision

A methodological framework for verifying and repairing adaptive agentic controllers through finite rule revision and diagnostic predicates.

AI/ML arXiv cs.AI

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

EvoCUA-1.5 introduces online reinforcement learning (RL) for computer-use agents, utilizing a new policy optimization method called STEPO and a dynamic curriculum.

Software Engineering arXiv cs.AI

From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction

A pipeline using SMT (Satisfiability Modulo Theories) to synthesize paths and construct 2D/3D maze structures from input patterns.

AI/ML arXiv cs.AI

Length Penalties Make Chain-of-Thought Less Monitorable

Research showing that length penalties in RL for Chain-of-Thought reasoning make it harder to monitor the actual influences driving a model's answer.

AI/ML arXiv cs.AI

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

PHITSBench is a new execution-scored benchmark for evaluating AI's ability to generate input for PHITS radiation-transport code.

AI/ML Hacker News

Agents.md – Dumb Human

A discussion on Agents.md, likely a framework or specification for AI agents.