All Articles
17524 articles total
NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts
Introduces NEST, a regime-oriented Mixture-of-Experts framework for handling dataset-level distribution shifts in long-term forecasting.
InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs
Researchers propose InductWave, a wavelet-based inductive embedding method for logical query answering on large knowledge graphs that outperforms baselines with fewer resources.
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
The paper identifies a 'false-pass' bias in LLM judges that silently disables skill retirement in self-evolving agents, leading to a failure in behavioral safety.
SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
SpaCellAgent is introduced as an autonomous LLM multi-agent framework that automates end-to-end spatiotemporal analysis and narrative generation for computational biology.
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning
The Pyligent framework teaches LLMs correction-aware reasoning by supervising failed-branch search trees, significantly improving solve rates in structured reasoning tasks.
Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
An ablation study reveals that LLM-generated skills do not reliably improve performance for data science workflows compared to basic prompting.
RL Post-Training Builds Compositional Reasoning Strategies
Research shows that RL post-training can compose primitive skills into higher-level reasoning strategies rather than just amplifying latent abilities.
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
A comprehensive survey of recursive self-improvement in AI, distinguishing between bounded self-refinement and open-ended recursive self-improvement.
SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents
SkillCenter is introduced as a massive, source-grounded open skill library for AI agents containing over 216,000 structured skills.
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
The 'institutional red-teaming' methodology demonstrates how deployment rules, not just model capabilities, causally shape safety outcomes in multi-agent AI systems.
Can Reinforcement Learning Efficiently Discover Price Manipulation?
An investigation into whether RL agents can discover price manipulation strategies in financial markets, finding RL often outperforms model-based approaches in noisy environments.
Learning social norms enhances compatibility in dynamic human-AI coordination
Researchers developed a framework to improve human-AI coordination by formalizing tacit social norms (predictability, value alignment, and advantage awareness) into quantifiable principles for LLMs.
Measuring Intelligence Beyond Human Scale
A new paradigm for measuring AI intelligence beyond human capability proposes using models to generate adversarial, verifiable challenges to evaluate other systems.
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
This study analyzes safety failures in multi-agent LLM pipelines, identifying 'operational reframing' and 'approval-framed delegation' as key risk factors that can bypass safety filters.
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
The ImagingBench benchmark reveals a significant gap between the semantic visual abilities of agentic AI and their ability to handle the physics of computational imaging tasks.
Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations
A new framework for 'reasoning consistency scanning' allows auditors to detect when an AI's chain-of-thought reasoning is logically inconsistent with its final answer.
From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
EvoSOP is a framework that allows LLM agents to self-evolve by synthesizing atomic actions into reusable Standard Operating Procedures (SOPs) to reduce reasoning overhead.
Physics-Audited Agentic Discovery in Scientific Machine Learning
PA-SciML introduces a verification-first workflow for agentic scientific machine learning to ensure discovered surrogate models satisfy fundamental physics requirements.
MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning
MIRA-Math is a benchmark designed to evaluate an AI's ability to identify and request specifically missing atomic facts needed to solve mathematical problems.
Agentic Data Environments
A proposal for 'Agentic Data Environments' reframes data systems from passive stores into active substrates that can amplify agent capabilities while enforcing safety guarantees.