All Articles
18677 articles total
SEAGym: An Evaluation Environment for Self-Evolving LLM Agents
Introduces SEAGym, an evaluation environment designed to measure the evolution and reliability of LLM agent harnesses.
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Presents DeepInsight, a unified evaluation infrastructure for the entire physical AI stack, from LLMs to humanoid control.
Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow
Demonstrates a foundation model-orchestrated workflow for accelerating pedestrian protection design in automotive crash safety.
Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning
Proposes a three-layer architecture for verbal reinforcement learning to prevent forgetting and improve insight governance in non-stationary environments.
Working in Glass
A discussion thread regarding working in Glass, though content is limited to comments.
NetNewsWire Status
A status update or discussion regarding NetNewsWire, a RSS reader.
SpeechDx: A Multi-Task Benchmark for Clinical Speech AI
Introduction of SpeechDx, a large-scale benchmark for clinical speech AI designed to evaluate generalization across diverse health conditions.
Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes
Proposed architecture for distributed general-purpose agent networks using P2P overlays to enable autonomous agent discovery and cooperation.
Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation
An adaptive clinical decision support system combining Treatment Effect estimation, Digital Twins, and Reinforcement Learning for personalized medicine.
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
Study on 'Incumbent Advantage' in LLM recommendation systems, highlighting how brand bias and Generative Engine Optimization (GEO) affect market competition.
A Machine-Learned Comorbidity Index
Proposal of a Machine-Learned Comorbidity Index (MLCI) that uses nHSIC to better capture nonlinear risk-outcome relationships in clinical data.
MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors
Introduction of MapSatisfyBench, a benchmark to evaluate how well map agents can identify and satisfy implicit user decision factors.
Dissecting model behavior through agent trajectories
Research on the 'intent-execution gap' in AI agents, introducing 'Simple Strands Agent' (SSA) to analyze and improve model-harness alignment.
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
CEO-Bench, a new benchmark evaluating LLMs' ability to handle strategic resource reallocation and synthesize conflicting C-suite advice.
Stop Killing Games fails to secure EU law despite 1.3M signatures
An initiative to prevent the permanent deletion of games failed to secure new EU laws despite significant public support.
The Amphibious Villagers of Indonesia
A feature on the amphibious villagers of Indonesia.
Leaked OpenAI financials show $38.5B loss and compute burn
Leaked financial documents reveal that OpenAI has incurred massive losses and significant compute costs.
Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
Researchers introduce DivInit, a training-free intervention that improves agentic search by ensuring diverse initial queries to reduce redundancy.
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
A new self-evolving framework for rule-driven query rewriting is proposed to enhance BM25 retrieval for legal case search without parameter training.
SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions
SkillChain-Gym is introduced as a benchmark for production-inventory control that accounts for workforce skill decay and reskilling needs.