AI/ML arXiv cs.AI

LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules

LegalFarePlan is a new route-planning framework for urban rail that handles non-additive fare rules and provides explainable plans.

AI/ML arXiv cs.AI

BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

BatteryLake is an agentic, physics-grounded data lakehouse for curating heterogeneous battery aging data, including an open benchmark of 41 datasets.

AI/ML arXiv cs.AI

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm

Research on the 'budgeted placement' of strong corrector agents in a swarm of weak agents to optimize consensus correctness versus cost.

AI/ML arXiv cs.AI

Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems

A study on robust norm enforcement mechanisms for AI agents to prevent exploitative behavior in multi-agent systems using reliability estimates and escalating penalties.

AI/ML arXiv cs.AI

Verification of Adaptive Agentic Controllers through Finite Rule Revision

A methodological framework for verifying and repairing adaptive agentic controllers through finite rule revision and diagnostic predicates.

AI/ML arXiv cs.AI

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

EvoCUA-1.5 introduces online reinforcement learning (RL) for computer-use agents, utilizing a new policy optimization method called STEPO and a dynamic curriculum.

Software Engineering arXiv cs.AI

From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction

A pipeline using SMT (Satisfiability Modulo Theories) to synthesize paths and construct 2D/3D maze structures from input patterns.

AI/ML arXiv cs.AI

Length Penalties Make Chain-of-Thought Less Monitorable

Research showing that length penalties in RL for Chain-of-Thought reasoning make it harder to monitor the actual influences driving a model's answer.

AI/ML arXiv cs.AI

PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language

PHITSBench is a new execution-scored benchmark for evaluating AI's ability to generate input for PHITS radiation-transport code.

AI/ML Hacker News

Agents.md – Dumb Human

A discussion on Agents.md, likely a framework or specification for AI agents.

Other Hacker News

Our Amish Language

A discussion about a specific constructed or community language, 'Our Amish Language'.

AI/ML arXiv cs.AI

AGM-like Paraconsistent Partial Meet Abductive Expansion Operation

Introduces a new paraconsistent AGM-like abductive expansion operation (AGMpabd) for handling contradictory hypotheses in reasoning systems.

AI/ML arXiv cs.AI

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

Proposes an evaluation-unsupervised prompt subset selection method (coreset selection) to efficiently compress LLM benchmarks without needing model evaluations.

AI/ML arXiv cs.AI

A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles

Presents DSiGAT, a dynamic scene graph attention framework for predicting lane-change intentions and trajectories of multiple vehicles in autonomous driving.

AI/ML arXiv cs.AI

Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets

Investigates how structured reasoning interventions (scaffolding) affect the strategic economic reasoning of standard vs. reasoning-optimized LLMs.

AI/ML arXiv cs.AI

A Theory of Least Autonomy in AI

Proposes the 'theory of least autonomy' for AI agents as a generalization of the principle of least privilege to manage agentic influence and collusion.

Open Source arXiv cs.AI

SupplyNetPy: An Open-Source Python Library for High-Fidelity Modeling and Simulation of Arbitrary Supply Chain and Inventory Networks

Introduces SupplyNetPy, an open-source Python library for high-fidelity modeling and simulation of arbitrary supply chain and inventory networks.

AI/ML arXiv cs.AI

Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems

Proposes Epistemic State Replication (ESR), a framework for distributed agentic systems to agree on semantic belief rather than bitwise state.

AI/ML arXiv cs.AI

Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks

Presents TCSDG, a task-conditioned synthetic data generation algorithm for improving ML performance in agricultural prediction tasks.

AI/ML arXiv cs.AI

From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation

A new framework for retinal diagnosis that decomposes ML predictions into structured argumentation components (claim, grounds, warrant, etc.) using MedGemma and MedSigLip to improve interpretability for medical experts.