Other Hacker News

US Supreme Court Just Blew Up EU-US Data Transfers

The US Supreme Court has issued a ruling that significantly impacts EU-US data transfers, potentially disrupting existing data flow agreements.

AI/ML arXiv cs.AI

How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions

Research on using simple self-play AI methods to determine optimal due diligence levels in takeover auctions, finding that modest diligence is often sufficient.

AI/ML arXiv cs.AI

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

Introduction of the Agent-Computer Observation Interface (AOI), a perception layer that improves computer-use agents' ability to handle dynamic UI events and audio.

AI/ML arXiv cs.AI

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

An audit of Lean theorem-proving benchmarks revealing thousands of defects, and the release of a suite of automated checkers to improve evaluation reliability.

AI/ML arXiv cs.AI

Cognitive World Models for Process-Level Social Influence Evaluation

The Cognitive World Model (CogWM) is introduced to evaluate social influence dialogue by tracking the evolution of a user's internal cognitive state.

AI/ML arXiv cs.AI

UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

UCOB is a new framework for reinforcement learning agents to utilize and evolve agentic skills via credit-aware bidirectional self-distillation.

AI/ML arXiv cs.AI

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

OSWorld 2.0 introduces a benchmark of 108 long-horizon computer-use workflows to test frontier agents' ability to handle professional-level real-world tasks.

AI/ML arXiv cs.AI

Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing

A study on the gap between theory-informed roles and learned coordination conventions in multi-agent reinforcement learning (MARL).

AI/ML arXiv cs.AI

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

SCARCE is proposed as a scalable method for characterizing rare-event failures in AI systems using latent representations and geometric rulers.

AI/ML arXiv cs.AI

SFBench: The SciFy Scientific Feasibility Benchmark

SFBench is a new benchmark dataset for evaluating the feasibility of scientific claims in materials science, created by subject matter experts.

AI/ML arXiv cs.AI

When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

Researchers investigate how LLM-based context compression in financial analysis can distort decision-making by removing critical caveats and qualifiers.

AI/ML arXiv cs.AI

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

The Complexity Ceiling Benchmark (CCB) evaluates how LLM reasoning performance decays as the number of sequential steps increases across multiple domains.

AI/ML arXiv cs.AI

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners

Introduces PASS, a middleware for process-supervised reinforcement learning that addresses structural pathologies in Group Relative Policy Optimization (GRPO) for LLM reasoners.

AI/ML arXiv cs.AI

Hierarchical Experimentalist Agents

HExA is a training-free framework that enables agents to learn from active experimentation and build reusable skill libraries to solve complex long-horizon tasks.

AI/ML arXiv cs.AI

PHF: Privileged Hidden Flow for On-Policy Self-Distillation

Proposes Privileged Hidden Flow (PHF), a method for on-policy self-distillation that aligns teacher-student hidden state trajectories rather than just output distributions.

AI/ML arXiv cs.AI

When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning

Introduces Communicative Language Symbolism Routing (CLSR), allowing multi-agent LLMs to invent compact symbolic languages to reduce token costs and latency.

AI/ML arXiv cs.AI

Diagnosing and Repairing Factual Errors in RAG under Budget Constraints

D2R-RAG is a resource-aware framework that diagnoses and repairs factual errors in RAG systems under specific latency and VRAM constraints.

AI/ML arXiv cs.AI

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

A planning-based LLM agent system for multi-hop reasoning over multimodal nuclear regulatory documents, outperforming standard RAG approaches.

AI/ML arXiv cs.AI

Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning

Introduces Mixture of Debaters (MoD), using an MoE-style architecture to enable a single model to perform dynamic self-debate with lower latency and token use.

AI/ML arXiv cs.AI

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

FADE is a training-free method to mitigate hallucinations in Large Vision-Language Models by attenuating FFN outputs to reduce language-prior dominance.