All Articles
17730 articles total
US Supreme Court Just Blew Up EU-US Data Transfers
The US Supreme Court has issued a ruling that significantly impacts EU-US data transfers, potentially disrupting existing data flow agreements.
How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions
Research on using simple self-play AI methods to determine optimal due diligence levels in takeover auctions, finding that modest diligence is often sufficient.
Agent-Computer Observation Interfaces Enable Dynamic Computer Use
Introduction of the Agent-Computer Observation Interface (AOI), a perception layer that improves computer-use agents' ability to handle dynamic UI events and audio.
Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
An audit of Lean theorem-proving benchmarks revealing thousands of defects, and the release of a suite of automated checkers to improve evaluation reliability.
Cognitive World Models for Process-Level Social Influence Evaluation
The Cognitive World Model (CogWM) is introduced to evaluate social influence dialogue by tracking the evolution of a user's internal cognitive state.
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation
UCOB is a new framework for reinforcement learning agents to utilize and evolve agentic skills via credit-aware bidirectional self-distillation.
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
OSWorld 2.0 introduces a benchmark of 108 long-horizon computer-use workflows to test frontier agents' ability to handle professional-level real-world tasks.
Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing
A study on the gap between theory-informed roles and learned coordination conventions in multi-agent reinforcement learning (MARL).
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
SCARCE is proposed as a scalable method for characterizing rare-event failures in AI systems using latent representations and geometric rulers.
SFBench: The SciFy Scientific Feasibility Benchmark
SFBench is a new benchmark dataset for evaluating the feasibility of scientific claims in materials science, created by subject matter experts.
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
Researchers investigate how LLM-based context compression in financial analysis can distort decision-making by removing critical caveats and qualifiers.
The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling
The Complexity Ceiling Benchmark (CCB) evaluates how LLM reasoning performance decays as the number of sequential steps increases across multiple domains.
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners
Introduces PASS, a middleware for process-supervised reinforcement learning that addresses structural pathologies in Group Relative Policy Optimization (GRPO) for LLM reasoners.
Hierarchical Experimentalist Agents
HExA is a training-free framework that enables agents to learn from active experimentation and build reusable skill libraries to solve complex long-horizon tasks.
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
Proposes Privileged Hidden Flow (PHF), a method for on-policy self-distillation that aligns teacher-student hidden state trajectories rather than just output distributions.
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Introduces Communicative Language Symbolism Routing (CLSR), allowing multi-agent LLMs to invent compact symbolic languages to reduce token costs and latency.
Diagnosing and Repairing Factual Errors in RAG under Budget Constraints
D2R-RAG is a resource-aware framework that diagnoses and repairs factual errors in RAG systems under specific latency and VRAM constraints.
LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents
A planning-based LLM agent system for multi-hop reasoning over multimodal nuclear regulatory documents, outperforming standard RAG approaches.
Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning
Introduces Mixture of Debaters (MoD), using an MoE-style architecture to enable a single model to perform dynamic self-debate with lower latency and token use.
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
FADE is a training-free method to mitigate hallucinations in Large Vision-Language Models by attenuating FFN outputs to reduce language-prior dominance.