All Articles
17889 articles total
SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills
Introduction of SkillDisCo, a framework that distills successful agent traces into reusable procedural skills represented as control-flow subgraphs.
NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research
A detailed, transparent post-training pipeline called NebulaExp for 8B-scale LLMs, focusing on instruction adherence and complex reasoning.
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
This paper evaluates confidence interval methods for performance metrics in text classification, proposing a pseudo-count regularized bootstrap for improved accuracy, especially for F1 scores.
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
Researchers argue that data-driven machine learning cannot reach symbolic-level logical reasoning due to methodological limitations in training data and target mapping.
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
The authors introduce MKG-RAG-Bench, a benchmark for evaluating retrieval in multimodal knowledge graph-augmented generation, highlighting retrieval as a critical bottleneck.
auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
The auto-psych system uses nested agent-based discovery loops to automate hypothesis generation, experimental design, and data collection in computational cognitive science.
Clinical Harness for Governable Medical AI Skill Ecosystems
The paper proposes the Clinical Harness, a runtime governance architecture for managing and monitoring AI-enabled clinical capabilities in medical care.
Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
A study on Large Reasoning Models (LRMs) reveals that they spend more tokens on failures than successes, contrasting with human behavior of disengaging from difficult tasks.
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications
A tutorial on NeuraDock Agent, an open-source EEG workflow for analyzing visual cognitive load and Alpha dynamics in real-time.
Boundary-Aware Context Grounding for A Low-Channel EEG Agent
NeuraDock Agent is presented as an open-source architecture that separates a deterministic EEG engine from a hardware-aware LLM layer to improve scientific grounding.
Radical AI Interpretability
This work develops a framework for AI interpretability, treating AI systems as agents to solve for beliefs and desires using mechanistic interpretability tools.
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
PMDformer is introduced as a transformer-based model for long-term time series forecasting that uses patch-mean decoupling to improve shape similarity modeling.
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
This research discusses the 'verification horizon' in coding agents, arguing that verifying complex solutions is becoming harder than generating them and requires co-evolving reward functions.
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?
An empirical study on tool-augmented LLM agents performing real-world energy analytics tasks, including market data retrieval and quantitative modeling.
What We are Missing in Multimodal LLM Evaluation?
A critical review of multimodal LLM evaluation benchmarks, highlighting gaps in temporal-spatial coherence and physical world understanding.
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
Introduction of OpenFinGym, a unified gym environment for evaluating quantitative finance agents across forecasting, trading, and fraud detection.
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
The paper defines 'compositional behavioral leakage' (CBL), where editing one prompt module in an agentic system silently affects others due to transformer self-attention.
Accelerating Returns and the Qualitative Engine for Science
Discusses the limitations of the accelerating returns thesis in scientific discovery and proposes the Qualitative Engine for Science (QES) to address qualitative reasoning gaps.
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Presents Narration-of-Thought (NoT), a system prompt scaffolding technique that improves ethical reasoning and reduces stakeholder collapse in LLMs.
Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry
Proposes a Geometry-Aware Monte Carlo Tree Search (MCTS) framework to solve extremal problems in combinatorial geometry more efficiently.