All Articles
17118 articles total
Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems
The article explores semantic context drift in reasoning-class LLMs and proposes a mathematical model to maintain operator control stability in human-machine decision systems.
Agentic Context Learning with Self-Discovered Specification
Research on context learning in LLMs reveals that acquiring local specifications is more critical than content acquisition, introducing the PSCI intervention to improve performance.
Exploring Agentic Workflows for Generating High Quality Math Visual Aids
Proposes an agentic workflow for creating high-quality K-12 math visual aids through an iterative self-improvement loop using VLMs.
TopoExplore: Topological Discrimination for Archive-Based Exploration
Introduces TopoExplore, a topology-aware exploration method for RL agents that improves efficiency in environments with enclosed structures.
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
Presents Who&When Pro, a large-scale benchmark designed to evaluate the ability of LLMs to accurately attribute failures in agentic systems.
A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution
Introduces a trace-supervised symbolic neural CPU that allows for interpretable program execution and quantization-simulated writeback.
AgentAbstain: Do LLM Agents Know When Not to Act?
Introduces AgentAbstain, a benchmark and automated pipeline (AbstainGen) to evaluate whether LLM agents know when to abstain from taking irreversible actions.
From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability
Discusses a formal mathematical approach to semantic caching and governed reuse classes to replace similarity-based heuristics in conversational demands.
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Introduces MAG, a benchmark and harness for multimodal web-agent task execution and guide generation using screenshots.
YouTrackDB is a general-use object-oriented graph database
YouTrackDB is introduced as a general-purpose object-oriented graph database.
LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules
LegalFarePlan is a new route-planning framework for urban rail that handles non-additive fare rules and provides explainable plans.
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking
BatteryLake is an agentic, physics-grounded data lakehouse for curating heterogeneous battery aging data, including an open benchmark of 41 datasets.
How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
Research on the 'budgeted placement' of strong corrector agents in a swarm of weak agents to optimize consensus correctness versus cost.
Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
A study on robust norm enforcement mechanisms for AI agents to prevent exploitative behavior in multi-agent systems using reliability estimates and escalating penalties.
Verification of Adaptive Agentic Controllers through Finite Rule Revision
A methodological framework for verifying and repairing adaptive agentic controllers through finite rule revision and diagnostic predicates.
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
EvoCUA-1.5 introduces online reinforcement learning (RL) for computer-use agents, utilizing a new policy optimization method called STEPO and a dynamic curriculum.
From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction
A pipeline using SMT (Satisfiability Modulo Theories) to synthesize paths and construct 2D/3D maze structures from input patterns.
Length Penalties Make Chain-of-Thought Less Monitorable
Research showing that length penalties in RL for Chain-of-Thought reasoning make it harder to monitor the actual influences driving a model's answer.
PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
PHITSBench is a new execution-scored benchmark for evaluating AI's ability to generate input for PHITS radiation-transport code.
Agents.md – Dumb Human
A discussion on Agents.md, likely a framework or specification for AI agents.