All Articles
17128 articles total
LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules
LegalFarePlan is a new route-planning framework for urban rail that handles non-additive fare rules and provides explainable plans.
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking
BatteryLake is an agentic, physics-grounded data lakehouse for curating heterogeneous battery aging data, including an open benchmark of 41 datasets.
How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
Research on the 'budgeted placement' of strong corrector agents in a swarm of weak agents to optimize consensus correctness versus cost.
Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
A study on robust norm enforcement mechanisms for AI agents to prevent exploitative behavior in multi-agent systems using reliability estimates and escalating penalties.
Verification of Adaptive Agentic Controllers through Finite Rule Revision
A methodological framework for verifying and repairing adaptive agentic controllers through finite rule revision and diagnostic predicates.
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
EvoCUA-1.5 introduces online reinforcement learning (RL) for computer-use agents, utilizing a new policy optimization method called STEPO and a dynamic curriculum.
From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction
A pipeline using SMT (Satisfiability Modulo Theories) to synthesize paths and construct 2D/3D maze structures from input patterns.
Length Penalties Make Chain-of-Thought Less Monitorable
Research showing that length penalties in RL for Chain-of-Thought reasoning make it harder to monitor the actual influences driving a model's answer.
PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
PHITSBench is a new execution-scored benchmark for evaluating AI's ability to generate input for PHITS radiation-transport code.
Agents.md – Dumb Human
A discussion on Agents.md, likely a framework or specification for AI agents.
Our Amish Language
A discussion about a specific constructed or community language, 'Our Amish Language'.
AGM-like Paraconsistent Partial Meet Abductive Expansion Operation
Introduces a new paraconsistent AGM-like abductive expansion operation (AGMpabd) for handling contradictory hypotheses in reasoning systems.
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
Proposes an evaluation-unsupervised prompt subset selection method (coreset selection) to efficiently compress LLM benchmarks without needing model evaluations.
A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles
Presents DSiGAT, a dynamic scene graph attention framework for predicting lane-change intentions and trajectories of multiple vehicles in autonomous driving.
Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets
Investigates how structured reasoning interventions (scaffolding) affect the strategic economic reasoning of standard vs. reasoning-optimized LLMs.
A Theory of Least Autonomy in AI
Proposes the 'theory of least autonomy' for AI agents as a generalization of the principle of least privilege to manage agentic influence and collusion.
SupplyNetPy: An Open-Source Python Library for High-Fidelity Modeling and Simulation of Arbitrary Supply Chain and Inventory Networks
Introduces SupplyNetPy, an open-source Python library for high-fidelity modeling and simulation of arbitrary supply chain and inventory networks.
Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems
Proposes Epistemic State Replication (ESR), a framework for distributed agentic systems to agree on semantic belief rather than bitwise state.
Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks
Presents TCSDG, a task-conditioned synthetic data generation algorithm for improving ML performance in agricultural prediction tasks.
From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation
A new framework for retinal diagnosis that decomposes ML predictions into structured argumentation components (claim, grounds, warrant, etc.) using MedGemma and MedSigLip to improve interpretability for medical experts.