All Articles
17543 articles total
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
DT-Guard is a low-latency safety guardrail for LLMs that uses reasoning-active training to internalize safety judgments without needing reasoning traces at inference time.
Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
This research integrates unsupervised dictionary learning into end-to-end autonomous driving models to make their decision-making process interpretable and correctable.
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting
TopoBrick is a training-free, zero-shot IoT forecasting framework that leverages building knowledge graphs and agentic topology sampling for better accuracy.
A Definition and Roadmap for World Models
A comprehensive perspective article that defines 'World Models' in AI and provides a staged roadmap for their future development.
ExplAIner: A Declarative Query Language for Explaining Classification Models
ExplAIner is a declarative query language designed to uniformly specify, combine, and analyze explanations for classification models.
Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports
A study evaluates the Nimblemind Multi-Agent System (nMAS) for evidence-linked extraction of H. pylori infections from medical biopsy reports.
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
Danus is an open-source orchestration system that uses a shared fact-graph memory to coordinate mathematical reasoning agents for long-horizon research problems.
A Physics-Informed Neural Network Framework for Elastodynamic Wave Propagation in Bimaterial Systems
A new PINN-based framework for modeling elastodynamic wave propagation in bimaterial systems, providing a computationally efficient surrogate model for solid mechanics.
Automate Excel with Python: From manual grind to one-click workflow
A discussion on automating Excel workflows using Python to transition from manual processes to one-click automation.
AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round
AI chip maker SambaNova has raised $1B at an $11B valuation, signaling continued high investor interest in AI hardware.
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test
A pre-registered experiment exploring information limits and attractor dynamics in small economies of LLM agents, confirming specific quantitative predictions about wealth growth and population misalignment.
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Introduction of PolyWorkBench, a benchmark for evaluating LLM agents' performance on multilingual, long-horizon workplace workflows across various domains.
Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency
A study proposing a reward-density heuristic (Efficiency heuristic) for dynamic multi-vehicle routing that matches the quality of complex metaheuristics with significantly lower compute costs.
When do prophets profit in prediction markets?
Research establishing a formal equivalence between predictive accuracy and profitability in prediction markets, with a live deployment achieving high ROI on Kalshi.
A toy framework for single and multi-agent human-AI curiosity ecosystems
A conceptual toy framework for analyzing curiosity as an ecosystem in both single and multi-agent human-AI systems.
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Introduction of IGRPO, a policy optimization framework for LLM agents that uses information gain to allocate rollout budgets more efficiently for long-horizon search tasks.
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
Demonstration of TOFFEE, a system that uses MCTS and adaptive model selection to synthesize high-quality data agent trajectories for finetuning and in-context learning.
From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution
A theoretical framework proposing a native meta-architecture for AI to move beyond stateless LLMs by introducing structural tension, recurrent loops, and inference-time plasticity.
GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos
Researchers successfully tricked GitHub's AI agent into leaking private repository information, highlighting critical security vulnerabilities in AI-integrated developer tools.
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
TurnOPD introduces a turn-level budgeting strategy for on-policy distillation to improve the efficiency and accuracy of training long-horizon AI agents.