All Articles
18316 articles total
DARPA Heavy Life Challenge
Details regarding the DARPA Heavy Life Challenge, focusing on robotics and physical tasks.
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
Introduces ENPIRE, a harness framework that allows coding agents to autonomously improve robotic policies through a real-world feedback loop of reset, execution, and verification.
Reward as An Agent for Embodied World Models
Proposes 'Reward as An Agent' and DynDiff-GRPO to enable broader exploration in embodied world models while mitigating reward hacking via robust verification.
Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
Evaluates multi-agent orchestration for enterprise AI, introducing a Task Manager to handle priority and event merging to reduce latency at scale.
Process-Verified Reinforcement Learning for Theorem Proving via Lean
Develops a process-verified RL approach using the Lean proof assistant as a symbolic reward oracle to improve theorem proving.
Residual-Space Evolutionary Optimization via Flow-based Generative Models
Introduces residual-space evolutionary optimization combining flow-based generative models with evolutionary algorithms for non-differentiable objectives.
Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing
Integrates multi-head attention with Soft Actor-Critic (SAC) to optimize process parameters and predict porosity in additive manufacturing.
ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research
Presents ScaffoldAgent, a framework that dynamically optimizes report outlines using utility-guided feedback for open-ended deep research.
eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization
Introduces eCNNTO, a CNN-based approach to accelerate density-based Topology Optimization by reducing iterations by up to 97% using residual connections and a novel training strategy.
The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self
Explores the concept of autotelic AI, focusing on how agents can generate their own goals and the philosophical and technical implications of the 'self' in agency.
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Presents PhysDrift, a framework for humanoid co-speech motion generation that bypasses human-body representations to create physically executable robot trajectories.
Advancing DialNav through Automatic Embodied Dialog Augmentation
Introduces the RAINbow dataset, a large-scale dataset for DialNav to improve embodied agent dialogue and navigation performance through automatic augmentation.
Gribouille 0.3.0: A Grammar of Graphics for Typst
Gribouille 0.3.0 is a 'Grammar of Graphics' implementation specifically designed for the Typst typesetting system.
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA
AgentFinVQA is a multi-agent pipeline for auditable and on-premise financial chart question answering, utilizing open-weights models like Qwen for data residency.
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
ORAgentBench is a new execution-grounded benchmark for evaluating autonomous LLM agents on end-to-end operations research tasks.
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
CombEval is a dynamic benchmark for evaluating the ability of LLMs to perform combinatorial counting through a typed specification framework.
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
The SEVRA framework introduces selective verification for reasoning allocation to optimize the cost-accuracy trade-off in test-time reasoning for LLMs.
Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery
A proposed four-layer verification architecture for AI-assisted legal discovery aims to reduce 'trajectory collapse' and privilege-waiver risk using human-on-the-loop orchestration.
TelcoAgent: A Scalable 5G Multi-KPM Forecasting With 3GPP-Grounded Explainability
TelcoAgent is a foundation model framework for scalable and explainable 5G KPM forecasting using a 3GPP-grounded knowledge graph.
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models
A systematic evaluation of black-box uncertainty estimation methods for LLMs provides practical guidance for building trustworthy models when API logits are unavailable.