AI/ML arXiv cs.AI

Kairos: A Native World Model Stack for Physical AI

Kairos is presented as a native world model stack for Physical AI, utilizing a cross-embodiment data curriculum and hybrid linear temporal attention.

AI/ML arXiv cs.AI

The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements

The authors introduce Bidirectional Provability Fingerprinting (BPF) to certify semantic equivalence between natural-language and formal mathematical statements.

AI/ML arXiv cs.AI

ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning

ROSA-RL is a reinforcement learning framework for uncertainty-aware speed advisory in roundabouts to improve autonomous driving safety.

AI/ML arXiv cs.AI

TNODEV: Toolbox for Neural ODE Verification

TNODEV is introduced as the first sound formal verifier for neural ordinary differential equations, integrating falsification and reachability analysis.

AI/ML arXiv cs.AI

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

The ARB4WM framework is introduced to benchmark the adversarial robustness of world models in continuous control systems under visual perturbations.

Other Hacker News

New way of making espresso with ultrasound

Exploration of using ultrasound technology to create espresso, potentially changing the extraction process.

Software Engineering Hacker News

Getting Creative with Perlin Noise Fields

A technical look at using Perlin Noise Fields for creative generative art and programming.

Open Source Hacker News

Trinket.io shutting down, so we saved it and hosted it a trinket.strivemath.org

The community saved the Trinket.io platform by hosting a mirror at trinket.strivemath.org after its shutdown.

Hardware/Chips Hacker News

Commodore Releases Flip Phone

Commodore has released a new flip phone, attempting to revive the brand in the mobile hardware market.

AI/ML arXiv cs.AI

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

Research showing that LLM tool-selection failures happen at the decision readout stage rather than due to 'lost-in-the-middle' context issues.

AI/ML arXiv cs.AI

Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions

Introduction of 'Posterior Twins,' a digital-twin approach for distributional behavioral simulation to improve enterprise decision-making.

AI/ML arXiv cs.AI

When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting

A framework for quantifying and insuring autonomous AI risk using 'trace-economic underwriting' based on tool-use traces.

AI/ML arXiv cs.AI

Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent LLM Planning

Tensor-Coord uses multilinear algebra (CP and Tucker decompositions) to prevent conflicts in multi-agent LLM planning.

AI/ML arXiv cs.AI

Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents

EC-Script is a framework for generating narratives with controlled emotional trajectories to assist in AI-driven art therapy.

AI/ML arXiv cs.AI

Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing

METIS introduces a loss-aware many-shot model merging protocol to reduce task interference and information erasure in multi-task LLMs.

Other The Verge

After resurrecting an iconic PC brand, Commodore is getting into flip phones

Commodore is being revived by a retro gaming YouTuber, starting with a modern version of the Commodore 64 and expanding into flip phones.

AI/ML arXiv cs.AI

Latent Thought Flow: Efficient Latent Reasoning in Large Language Models

Researchers propose Latent Thought Flow (LTF), a method that moves LLM reasoning into continuous space to reduce inference overhead compared to traditional Chain-of-Thought.

AI/ML arXiv cs.AI

SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

SpecAlign is introduced as a framework to align LLMs using synthetic data generated directly from structured provider specifications rather than abstract principles.

AI/ML arXiv cs.AI

State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs

StateGen is a synthetic data generation platform for tool-augmented LLMs that uses an authoritative state manager to eliminate tool-call hallucinations.

AI/ML arXiv cs.AI

Architectural Wisdom: A Framework for Governing Optimization in AI Systems

The authors propose 'architectural wisdom,' a governance layer for AI systems designed to interrogate objectives rather than just optimizing them to prevent structural failures.