AI/ML arXiv cs.AI

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail

DT-Guard is a low-latency safety guardrail for LLMs that uses reasoning-active training to internalize safety judgments without needing reasoning traces at inference time.

AI/ML arXiv cs.AI

Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models

This research integrates unsupervised dictionary learning into end-to-end autonomous driving models to make their decision-making process interpretable and correctable.

AI/ML arXiv cs.AI

TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting

TopoBrick is a training-free, zero-shot IoT forecasting framework that leverages building knowledge graphs and agentic topology sampling for better accuracy.

AI/ML arXiv cs.AI

A Definition and Roadmap for World Models

A comprehensive perspective article that defines 'World Models' in AI and provides a staged roadmap for their future development.

AI/ML arXiv cs.AI

ExplAIner: A Declarative Query Language for Explaining Classification Models

ExplAIner is a declarative query language designed to uniformly specify, combine, and analyze explanations for classification models.

AI/ML arXiv cs.AI

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

A study evaluates the Nimblemind Multi-Agent System (nMAS) for evidence-linked extraction of H. pylori infections from medical biopsy reports.

AI/ML arXiv cs.AI

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

Danus is an open-source orchestration system that uses a shared fact-graph memory to coordinate mathematical reasoning agents for long-horizon research problems.

AI/ML arXiv cs.AI

A Physics-Informed Neural Network Framework for Elastodynamic Wave Propagation in Bimaterial Systems

A new PINN-based framework for modeling elastodynamic wave propagation in bimaterial systems, providing a computationally efficient surrogate model for solid mechanics.

Software Engineering Hacker News

Automate Excel with Python: From manual grind to one-click workflow

A discussion on automating Excel workflows using Python to transition from manual processes to one-click automation.

Tech Business/VC TechCrunch

AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round

AI chip maker SambaNova has raised $1B at an $11B valuation, signaling continued high investor interest in AI hardware.

AI/ML arXiv cs.AI

Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

A pre-registered experiment exploring information limits and attractor dynamics in small economies of LLM agents, confirming specific quantitative predictions about wealth growth and population misalignment.

AI/ML arXiv cs.AI

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents

Introduction of PolyWorkBench, a benchmark for evaluating LLM agents' performance on multilingual, long-horizon workplace workflows across various domains.

AI/ML arXiv cs.AI

Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency

A study proposing a reward-density heuristic (Efficiency heuristic) for dynamic multi-vehicle routing that matches the quality of complex metaheuristics with significantly lower compute costs.

AI/ML arXiv cs.AI

When do prophets profit in prediction markets?

Research establishing a formal equivalence between predictive accuracy and profitability in prediction markets, with a live deployment achieving high ROI on Kalshi.

AI/ML arXiv cs.AI

A toy framework for single and multi-agent human-AI curiosity ecosystems

A conceptual toy framework for analyzing curiosity as an ecosystem in both single and multi-agent human-AI systems.

AI/ML arXiv cs.AI

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Introduction of IGRPO, a policy optimization framework for LLM agents that uses information gain to allocate rollout budgets more efficiently for long-horizon search tasks.

AI/ML arXiv cs.AI

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

Demonstration of TOFFEE, a system that uses MCTS and adaptive model selection to synthesize high-quality data agent trajectories for finetuning and in-context learning.

AI/ML arXiv cs.AI

From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution

A theoretical framework proposing a native meta-architecture for AI to move beyond stateless LLMs by introducing structural tension, recurrent loops, and inference-time plasticity.

Cybersecurity Hacker News

GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos

Researchers successfully tricked GitHub's AI agent into leaking private repository information, highlighting critical security vulnerabilities in AI-integrated developer tools.

AI/ML arXiv cs.AI

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

TurnOPD introduces a turn-level budgeting strategy for on-policy distillation to improve the efficiency and accuracy of training long-horizon AI agents.