AI/ML arXiv cs.AI

NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts

Introduces NEST, a regime-oriented Mixture-of-Experts framework for handling dataset-level distribution shifts in long-term forecasting.

AI/ML arXiv cs.AI

InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs

Researchers propose InductWave, a wavelet-based inductive embedding method for logical query answering on large knowledge graphs that outperforms baselines with fewer resources.

AI/ML arXiv cs.AI

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

The paper identifies a 'false-pass' bias in LLM judges that silently disables skill retirement in self-evolving agents, leading to a failure in behavioral safety.

AI/ML arXiv cs.AI

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

SpaCellAgent is introduced as an autonomous LLM multi-agent framework that automates end-to-end spatiotemporal analysis and narrative generation for computational biology.

AI/ML arXiv cs.AI

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

The Pyligent framework teaches LLMs correction-aware reasoning by supervising failed-branch search trees, significantly improving solve rates in structured reasoning tasks.

AI/ML arXiv cs.AI

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

An ablation study reveals that LLM-generated skills do not reliably improve performance for data science workflows compared to basic prompting.

AI/ML arXiv cs.AI

RL Post-Training Builds Compositional Reasoning Strategies

Research shows that RL post-training can compose primitive skills into higher-level reasoning strategies rather than just amplifying latent abilities.

AI/ML arXiv cs.AI

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

A comprehensive survey of recursive self-improvement in AI, distinguishing between bounded self-refinement and open-ended recursive self-improvement.

AI/ML arXiv cs.AI

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

SkillCenter is introduced as a massive, source-grounded open skill library for AI agents containing over 216,000 structured skills.

AI/ML arXiv cs.AI

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

The 'institutional red-teaming' methodology demonstrates how deployment rules, not just model capabilities, causally shape safety outcomes in multi-agent AI systems.

AI/ML arXiv cs.AI

Can Reinforcement Learning Efficiently Discover Price Manipulation?

An investigation into whether RL agents can discover price manipulation strategies in financial markets, finding RL often outperforms model-based approaches in noisy environments.

AI/ML arXiv cs.AI

Learning social norms enhances compatibility in dynamic human-AI coordination

Researchers developed a framework to improve human-AI coordination by formalizing tacit social norms (predictability, value alignment, and advantage awareness) into quantifiable principles for LLMs.

AI/ML arXiv cs.AI

Measuring Intelligence Beyond Human Scale

A new paradigm for measuring AI intelligence beyond human capability proposes using models to generate adversarial, verifiable challenges to evaluate other systems.

AI/ML arXiv cs.AI

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

This study analyzes safety failures in multi-agent LLM pipelines, identifying 'operational reframing' and 'approval-framed delegation' as key risk factors that can bypass safety filters.

AI/ML arXiv cs.AI

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

The ImagingBench benchmark reveals a significant gap between the semantic visual abilities of agentic AI and their ability to handle the physics of computational imaging tasks.

AI/ML arXiv cs.AI

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

A new framework for 'reasoning consistency scanning' allows auditors to detect when an AI's chain-of-thought reasoning is logically inconsistent with its final answer.

AI/ML arXiv cs.AI

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

EvoSOP is a framework that allows LLM agents to self-evolve by synthesizing atomic actions into reusable Standard Operating Procedures (SOPs) to reduce reasoning overhead.

AI/ML arXiv cs.AI

Physics-Audited Agentic Discovery in Scientific Machine Learning

PA-SciML introduces a verification-first workflow for agentic scientific machine learning to ensure discovered surrogate models satisfy fundamental physics requirements.

AI/ML arXiv cs.AI

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

MIRA-Math is a benchmark designed to evaluate an AI's ability to identify and request specifically missing atomic facts needed to solve mathematical problems.

AI/ML arXiv cs.AI

Agentic Data Environments

A proposal for 'Agentic Data Environments' reframes data systems from passive stores into active substrates that can amplify agent capabilities while enforcing safety guarantees.