AI/ML arXiv cs.AI

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

MerchantBench is a new 365-day simulation benchmark for testing the long-term coherence of LLM agents in complex e-commerce operations.

AI/ML arXiv cs.AI

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

SciDisco is a scalable framework for training scientific discovery agents using process-verifiable environments and turn-level credit assignment.

AI/ML arXiv cs.AI

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

MMShopBench introduces a real-log benchmark and an offline shopping sandbox for evaluating multimodal, multi-turn shopping agents.

AI/ML arXiv cs.AI

Evidence-Grounded Constraint Checking in Construction Documents

This research investigates the trade-offs between resolution and breadth when using RAG-based pipelines for evidence-grounded constraint checking in construction documents.

AI/ML arXiv cs.AI

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

The study explores how activation steering vectors can improve the faithfulness of Chain-of-Thought reasoning across different LLMs.

AI/ML arXiv cs.AI

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

The authors present a Generalized-Bayes perspective on counterfactual explanations, introducing new decision rules for more interpretable ML model outputs.

AI/ML arXiv cs.AI

Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration

WILC is a framework for coordinating multiple LLMs through iterative collaboration and complementarity-driven model selection to achieve higher collective intelligence.

Tech Business/VC Hacker News

OpenAI's super PAC is funding AI-generated news site attacking industry critics

OpenAI's super PAC is allegedly funding an AI-generated news site used to attack critics of the AI industry.

Software Engineering Hacker News

Cro – elegant reactive services in Raku

Introduction to Cro, a framework for building elegant reactive services using the Raku programming language.

Software Engineering Hacker News

AI migrated legacy COBOL programs to Java, bugs included

A report on the failure of AI to migrate legacy COBOL programs to Java without introducing bugs.

AI/ML arXiv cs.AI

Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO

A research paper presenting STL-GO, a formalism for multi-agent planning with spatio-temporal and topological constraints using MIP and SMT encodings.

AI/ML arXiv cs.AI

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

An analysis of LSR-Synth, a benchmark designed to measure symbolic discovery in scientific equations while preventing model memorization.

AI/ML arXiv cs.AI

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

A validity audit of current agent-safety benchmarks, arguing that many scores correlate more with general capability than actual safety.

AI/ML arXiv cs.AI

SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition

Introduction of SciToolAgent-Evo, an agent that self-evolves to acquire scientific tools, and OpenSciToolBench, a new evaluation benchmark.

AI/ML arXiv cs.AI

EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses

EarlyDx is introduced as a benchmark for open-ended generation of evidence-supported emergency department encounter diagnoses.

AI/ML arXiv cs.AI

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

A proposal for an interaction-centric taxonomy to better localize and repair failures in AI agents across different components.

AI/ML arXiv cs.AI

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

A study using the ANCHOR audit to evaluate persona collapse and behavioral drift in long-horizon AI companions.

AI/ML arXiv cs.AI

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

Proposes a layered architecture for Agentic AI using Ollama for inference and OpenClaw for orchestration to enable autonomous, goal-driven AI agents.

AI/ML arXiv cs.AI

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Introduces a benchmarking protocol using multi-model LLM peer-review to evaluate autonomous AI research generation systems.

AI/ML arXiv cs.AI

LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis

Presents a three-stage pipeline for discovering mathematical conjectures, utilizing Lean 4 for formal validation of the findings.