Other Hacker News

DARPA Heavy Life Challenge

Details regarding the DARPA Heavy Life Challenge, focusing on robotics and physical tasks.

AI/ML arXiv cs.AI

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

Introduces ENPIRE, a harness framework that allows coding agents to autonomously improve robotic policies through a real-world feedback loop of reset, execution, and verification.

AI/ML arXiv cs.AI

Reward as An Agent for Embodied World Models

Proposes 'Reward as An Agent' and DynDiff-GRPO to enable broader exploration in embodied world models while mitigating reward hacking via robust verification.

AI/ML arXiv cs.AI

Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale

Evaluates multi-agent orchestration for enterprise AI, introducing a Task Manager to handle priority and event merging to reduce latency at scale.

AI/ML arXiv cs.AI

Process-Verified Reinforcement Learning for Theorem Proving via Lean

Develops a process-verified RL approach using the Lean proof assistant as a symbolic reward oracle to improve theorem proving.

AI/ML arXiv cs.AI

Residual-Space Evolutionary Optimization via Flow-based Generative Models

Introduces residual-space evolutionary optimization combining flow-based generative models with evolutionary algorithms for non-differentiable objectives.

AI/ML arXiv cs.AI

Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing

Integrates multi-head attention with Soft Actor-Critic (SAC) to optimize process parameters and predict porosity in additive manufacturing.

AI/ML arXiv cs.AI

ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research

Presents ScaffoldAgent, a framework that dynamically optimizes report outlines using utility-guided feedback for open-ended deep research.

AI/ML arXiv cs.AI

eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization

Introduces eCNNTO, a CNN-based approach to accelerate density-based Topology Optimization by reducing iterations by up to 97% using residual connections and a novel training strategy.

AI/ML arXiv cs.AI

The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self

Explores the concept of autotelic AI, focusing on how agents can generate their own goals and the philosophical and technical implications of the 'self' in agency.

AI/ML arXiv cs.AI

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

Presents PhysDrift, a framework for humanoid co-speech motion generation that bypasses human-body representations to create physically executable robot trajectories.

AI/ML arXiv cs.AI

Advancing DialNav through Automatic Embodied Dialog Augmentation

Introduces the RAINbow dataset, a large-scale dataset for DialNav to improve embodied agent dialogue and navigation performance through automatic augmentation.

Open Source Hacker News

Gribouille 0.3.0: A Grammar of Graphics for Typst

Gribouille 0.3.0 is a 'Grammar of Graphics' implementation specifically designed for the Typst typesetting system.

AI/ML arXiv cs.AI

AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

AgentFinVQA is a multi-agent pipeline for auditable and on-premise financial chart question answering, utilizing open-weights models like Qwen for data residency.

AI/ML arXiv cs.AI

ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

ORAgentBench is a new execution-grounded benchmark for evaluating autonomous LLM agents on end-to-end operations research tasks.

AI/ML arXiv cs.AI

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

CombEval is a dynamic benchmark for evaluating the ability of LLMs to perform combinatorial counting through a typed specification framework.

AI/ML arXiv cs.AI

Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

The SEVRA framework introduces selective verification for reasoning allocation to optimize the cost-accuracy trade-off in test-time reasoning for LLMs.

AI/ML arXiv cs.AI

Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery

A proposed four-layer verification architecture for AI-assisted legal discovery aims to reduce 'trajectory collapse' and privilege-waiver risk using human-on-the-loop orchestration.

AI/ML arXiv cs.AI

TelcoAgent: A Scalable 5G Multi-KPM Forecasting With 3GPP-Grounded Explainability

TelcoAgent is a foundation model framework for scalable and explainable 5G KPM forecasting using a 3GPP-grounded knowledge graph.

AI/ML arXiv cs.AI

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

A systematic evaluation of black-box uncertainty estimation methods for LLMs provides practical guidance for building trustworthy models when API logits are unavailable.