AI/ML arXiv cs.AI

SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills

Introduction of SkillDisCo, a framework that distills successful agent traces into reusable procedural skills represented as control-flow subgraphs.

AI/ML arXiv cs.AI

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

A detailed, transparent post-training pipeline called NebulaExp for 8B-scale LLMs, focusing on instruction adherence and complex reasoning.

AI/ML arXiv cs.AI

Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data

This paper evaluates confidence interval methods for performance metrics in text classification, proposing a pseudo-count regularized bootstrap for improved accuracy, especially for F1 scores.

AI/ML arXiv cs.AI

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

Researchers argue that data-driven machine learning cannot reach symbolic-level logical reasoning due to methodological limitations in training data and target mapping.

AI/ML arXiv cs.AI

MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation

The authors introduce MKG-RAG-Bench, a benchmark for evaluating retrieval in multimodal knowledge graph-augmented generation, highlighting retrieval as a critical bottleneck.

AI/ML arXiv cs.AI

auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation

The auto-psych system uses nested agent-based discovery loops to automate hypothesis generation, experimental design, and data collection in computational cognitive science.

AI/ML arXiv cs.AI

Clinical Harness for Governable Medical AI Skill Ecosystems

The paper proposes the Clinical Harness, a runtime governance architecture for managing and monitoring AI-enabled clinical capabilities in medical care.

AI/ML arXiv cs.AI

Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

A study on Large Reasoning Models (LRMs) reveals that they spend more tokens on failures than successes, contrasting with human behavior of disengaging from difficult tasks.

AI/ML arXiv cs.AI

NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications

A tutorial on NeuraDock Agent, an open-source EEG workflow for analyzing visual cognitive load and Alpha dynamics in real-time.

AI/ML arXiv cs.AI

Boundary-Aware Context Grounding for A Low-Channel EEG Agent

NeuraDock Agent is presented as an open-source architecture that separates a deterministic EEG engine from a hardware-aware LLM layer to improve scientific grounding.

AI/ML arXiv cs.AI

Radical AI Interpretability

This work develops a framework for AI interpretability, treating AI systems as agents to solve for beliefs and desires using mechanistic interpretability tools.

AI/ML arXiv cs.AI

PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting

PMDformer is introduced as a transformer-based model for long-term time series forecasting that uses patch-mean decoupling to improve shape similarity modeling.

AI/ML arXiv cs.AI

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

This research discusses the 'verification horizon' in coding agents, arguing that verifying complex solutions is becoming harder than generating them and requires co-evolving reward functions.

AI/ML arXiv cs.AI

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?

An empirical study on tool-augmented LLM agents performing real-world energy analytics tasks, including market data retrieval and quantitative modeling.

AI/ML arXiv cs.AI

What We are Missing in Multimodal LLM Evaluation?

A critical review of multimodal LLM evaluation benchmarks, highlighting gaps in temporal-spatial coherence and physical world understanding.

AI/ML arXiv cs.AI

OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents

Introduction of OpenFinGym, a unified gym environment for evaluating quantitative finance agents across forecasting, trading, and fraud detection.

AI/ML arXiv cs.AI

Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems

The paper defines 'compositional behavioral leakage' (CBL), where editing one prompt module in an agentic system silently affects others due to transformer self-attention.

AI/ML arXiv cs.AI

Accelerating Returns and the Qualitative Engine for Science

Discusses the limitations of the accelerating returns thesis in scientific discovery and proposes the Qualitative Engine for Science (QES) to address qualitative reasoning gaps.

AI/ML arXiv cs.AI

Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models

Presents Narration-of-Thought (NoT), a system prompt scaffolding technique that improves ethical reasoning and reduces stakeholder collapse in LLMs.

AI/ML arXiv cs.AI

Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry

Proposes a Geometry-Aware Monte Carlo Tree Search (MCTS) framework to solve extremal problems in combinatorial geometry more efficiently.