All Articles
15976 articles total
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
MerchantBench is a new 365-day simulation benchmark for testing the long-term coherence of LLM agents in complex e-commerce operations.
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
SciDisco is a scalable framework for training scientific discovery agents using process-verifiable environments and turn-level credit assignment.
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
MMShopBench introduces a real-log benchmark and an offline shopping sandbox for evaluating multimodal, multi-turn shopping agents.
Evidence-Grounded Constraint Checking in Construction Documents
This research investigates the trade-offs between resolution and breadth when using RAG-based pipelines for evidence-grounded constraint checking in construction documents.
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
The study explores how activation steering vectors can improve the faithfulness of Chain-of-Thought reasoning across different LLMs.
A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation
The authors present a Generalized-Bayes perspective on counterfactual explanations, introducing new decision rules for more interpretable ML model outputs.
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration
WILC is a framework for coordinating multiple LLMs through iterative collaboration and complementarity-driven model selection to achieve higher collective intelligence.
OpenAI's super PAC is funding AI-generated news site attacking industry critics
OpenAI's super PAC is allegedly funding an AI-generated news site used to attack critics of the AI industry.
Cro – elegant reactive services in Raku
Introduction to Cro, a framework for building elegant reactive services using the Raku programming language.
AI migrated legacy COBOL programs to Java, bugs included
A report on the failure of AI to migrate legacy COBOL programs to Java without introducing bugs.
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
A research paper presenting STL-GO, a formalism for multi-agent planning with spatio-temporal and topological constraints using MIP and SMT encodings.
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
An analysis of LSR-Synth, a benchmark designed to measure symbolic discovery in scientific equations while preventing model memorization.
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
A validity audit of current agent-safety benchmarks, arguing that many scores correlate more with general capability than actual safety.
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Introduction of SciToolAgent-Evo, an agent that self-evolves to acquire scientific tools, and OpenSciToolBench, a new evaluation benchmark.
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
EarlyDx is introduced as a benchmark for open-ended generation of evidence-supported emergency department encounter diagnoses.
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
A proposal for an interaction-centric taxonomy to better localize and repair failures in AI agents across different components.
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
A study using the ANCHOR audit to evaluate persona collapse and behavioral drift in long-horizon AI companions.
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
Proposes a layered architecture for Agentic AI using Ollama for inference and OpenClaw for orchestration to enable autonomous, goal-driven AI agents.
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
Introduces a benchmarking protocol using multi-model LLM peer-review to evaluate autonomous AI research generation systems.
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
Presents a three-stage pipeline for discovering mathematical conjectures, utilizing Lean 4 for formal validation of the findings.