AI/ML arXiv cs.AI

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Presents UrbanDS, a graph-guided multi-agent system designed to automate data-intensive urban data science tasks.

AI/ML arXiv cs.AI

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

A causal audit of latent communication in multi-agent LLM systems to determine if receivers actually use task-relevant information.

AI/ML arXiv cs.AI

Property-driven Causal Abstractions for Markov Decision Processes

Introduces a property-driven causal abstraction technique to reduce state space blowup in Markov Decision Processes (MDPs).

AI/ML arXiv cs.AI

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

Introduces Pegasus, a framework that converts human manipulation videos into robot-learnable data via structured knowledge transfer.

Cybersecurity arXiv cs.AI

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Researches the detection of AI agents using browser automation, identifying key behavioral features that distinguish them from humans.

AI/ML arXiv cs.AI

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

Proposes a Belief-Guided architecture for Computer Go to reduce reliance on costly MCTS search on consumer-grade hardware.

AI/ML arXiv cs.AI

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Introduces Setoka, a benchmark for evaluating hierarchical user understanding in personalized agents over heterogeneous data.

AI/ML arXiv cs.AI

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Proposes Routing-based On-Policy Distillation (ROPD) for LLM safety realignment to mitigate template-mismatch risks and jailbreaking.

AI/ML arXiv cs.AI

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Researchers introduce StatMechBench-v0 to test if LLM-based agents can discover statistical mechanical mappings in theoretical physics, finding that while numerical feedback helps, agents still struggle with structural discovery.

AI/ML arXiv cs.AI

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

CaM-Wolf is a multimodal AI agent for social deduction games like Werewolf that uses video perception and causal-aware reasoning to better simulate human social interaction.

AI/ML arXiv cs.AI

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

The CG-World dataset provides a large-scale, structured world-state protocol derived from computer graphics pipelines to improve world models and embodied AI.

AI/ML arXiv cs.AI

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

MultivationBench is a new benchmark designed to evaluate how multimodal LLMs reason about evolving human motivations within sequential visual narratives.

AI/ML arXiv cs.AI

EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

EvoPINN is an agentic framework that automates the discovery of executable algorithms for Physics-Informed Neural Networks (PINNs), successfully inventing a new architecture called SLRC-PINN.

AI/ML arXiv cs.AI

Evidence-Ledger Adjudication for Claim-Evidence Traceability

The Evidence-Ledger Adjudication workflow improves claim-evidence traceability in AI-assisted writing by routing unsupported claims back to authors via an auditable ledger.

AI/ML arXiv cs.AI

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

Eco3S is a simulation framework for socio-economic research that uses LLM-based agent-based modeling and structural causal mechanisms to analyze policy and economic phenomena.

Software Engineering arXiv cs.AI

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

The CAPA benchmark evaluates how AI coding assistants adapt to personalized ambiguity patterns across different sessions to reduce the need for repeated clarifications.

AI/ML arXiv cs.AI

AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining

AlphaSchema introduces a structured semantic space for LLM-based alpha mining in trading, decoupling the exploration of trading factors from their actual implementation.

AI/ML arXiv cs.AI

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

SkillBoost is a framework that enables LLM agents to evolve skills through a constrained exploration-exploitation process, preventing skill overfitting while improving performance.

Tech Business/VC Hacker News

The AI trade now runs on borrowed money, and the lenders are repricing it

A discussion on how the AI investment boom is increasingly funded by borrowed money, leading to higher costs for lenders.

Software Engineering Hacker News

The Session You Cannot take with you

A technical discussion regarding session management and the inability to persist certain sessions across different environments.