AI/ML arXiv cs.AI

SEAGym: An Evaluation Environment for Self-Evolving LLM Agents

Introduces SEAGym, an evaluation environment designed to measure the evolution and reliability of LLM agent harnesses.

AI/ML arXiv cs.AI

DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack

Presents DeepInsight, a unified evaluation infrastructure for the entire physical AI stack, from LLMs to humanoid control.

AI/ML arXiv cs.AI

Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow

Demonstrates a foundation model-orchestrated workflow for accelerating pedestrian protection design in automotive crash safety.

AI/ML arXiv cs.AI

Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning

Proposes a three-layer architecture for verbal reinforcement learning to prevent forgetting and improve insight governance in non-stationary environments.

Other Hacker News

Working in Glass

A discussion thread regarding working in Glass, though content is limited to comments.

Software Engineering Hacker News

NetNewsWire Status

A status update or discussion regarding NetNewsWire, a RSS reader.

AI/ML arXiv cs.AI

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

Introduction of SpeechDx, a large-scale benchmark for clinical speech AI designed to evaluate generalization across diverse health conditions.

AI/ML arXiv cs.AI

Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes

Proposed architecture for distributed general-purpose agent networks using P2P overlays to enable autonomous agent discovery and cooperation.

AI/ML arXiv cs.AI

Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation

An adaptive clinical decision support system combining Treatment Effect estimation, Digital Twins, and Reinforcement Learning for personalized medicine.

AI/ML arXiv cs.AI

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

Study on 'Incumbent Advantage' in LLM recommendation systems, highlighting how brand bias and Generative Engine Optimization (GEO) affect market competition.

AI/ML arXiv cs.AI

A Machine-Learned Comorbidity Index

Proposal of a Machine-Learned Comorbidity Index (MLCI) that uses nHSIC to better capture nonlinear risk-outcome relationships in clinical data.

AI/ML arXiv cs.AI

MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors

Introduction of MapSatisfyBench, a benchmark to evaluate how well map agents can identify and satisfy implicit user decision factors.

AI/ML arXiv cs.AI

Dissecting model behavior through agent trajectories

Research on the 'intent-execution gap' in AI agents, introducing 'Simple Strands Agent' (SSA) to analyze and improve model-harness alignment.

AI/ML arXiv cs.AI

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

CEO-Bench, a new benchmark evaluating LLMs' ability to handle strategic resource reallocation and synthesize conflicting C-suite advice.

Other Hacker News

Stop Killing Games fails to secure EU law despite 1.3M signatures

An initiative to prevent the permanent deletion of games failed to secure new EU laws despite significant public support.

Other Hacker News

The Amphibious Villagers of Indonesia

A feature on the amphibious villagers of Indonesia.

Tech Business/VC Hacker News

Leaked OpenAI financials show $38.5B loss and compute burn

Leaked financial documents reveal that OpenAI has incurred massive losses and significant compute costs.

AI/ML arXiv cs.AI

Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search

Researchers introduce DivInit, a training-free intervention that improves agentic search by ensuring diverse initial queries to reduce redundancy.

AI/ML arXiv cs.AI

When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval

A new self-evolving framework for rule-driven query rewriting is proposed to enhance BM25 retrieval for legal case search without parameter training.

AI/ML arXiv cs.AI

SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions

SkillChain-Gym is introduced as a benchmark for production-inventory control that accounts for workforce skill decay and reskilling needs.