AI/ML arXiv cs.AI

Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization

DUPO is proposed to handle delayed feedback in reinforcement learning using diffusion models to estimate state discrepancies.

Other Hacker News

Historic Photos of NASA's Cavernous Wind Tunnels

A collection of historic photographs documenting NASA's large-scale wind tunnel facilities.

Software Engineering Hacker News

Show HN: Fast, native Mac file manager (filters, fuzzy find, 9 MB, no Electron)

A new native Mac file manager featuring fuzzy finding and filters, built without Electron to maintain a small footprint (9 MB).

Other Hacker News

Why migrants come to Germany for work and then leave again

An exploration of the factors causing migrants to leave Germany after initially arriving for work.

Other The Verge

Are you ready for what it takes to stop ghost guns?

An investigation into the use of 3D printers to manufacture illegal firearm components and automatic weapon switches.

Hardware/Chips The Verge

Nothing’s first B-series phone is also skipping the US

Nothing launches the Phone 4B, a budget-tier device skipping the US market.

Hardware/Chips The Verge

Nothing’s new earbuds can record calls and what you’re listening to

Nothing introduces new budget-friendly Ear 3A earbuds with the ability to record audio directly on the device.

AI/ML arXiv cs.AI

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

AgenticPD is a stage-aware agentic framework designed to optimize physical design Quality-of-Results (QoR) in EDA flows.

AI/ML arXiv cs.AI

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

CARL is a novel reinforcement learning framework that improves LLM planning by strengthening intrinsic constraint awareness.

AI/ML arXiv cs.AI

Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation

Medi-Gemma is a hybrid clinical decision support system that integrates deterministic EMR analytics with RAG to reduce hallucinations in healthcare.

AI/ML arXiv cs.AI

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

STAPO introduces a hierarchical RL framework using normalized entropy to solve the problem of trajectory neglect in long-horizon LLM agent training.

Cybersecurity Hacker News

Microsoft Can Track Users via a Windows Device ID

Discussion on Microsoft's ability to track Windows users via a unique device ID.

Other Hacker News

In Praise of Observational Evidence

An exploration and praise of the importance of observational evidence in scientific or technical contexts.

Other Hacker News

Dropping in on Gottfried Leibniz (2013)

A retrospective look or discussion involving the philosopher and mathematician Gottfried Leibniz.

Software Engineering Hacker News

The Art of Computer Programming by Donald E. Knuth

Discussion regarding Donald Knuth's seminal work, 'The Art of Computer Programming'.

Hardware/Chips TechCrunch

The first American autonomous ground vehicles are fighting in Ukraine

Forterra has deployed over 100 autonomous ground vehicles for combat operations in Ukraine.

AI/ML arXiv cs.AI

MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents

Introduction of MRMS, a multi-resolution memory substrate designed to provide continuity and personalization for long-lived AI agents.

AI/ML arXiv cs.AI

Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs

Formal Disco is a distributed system for generating large-scale synthetic datasets of formally verified programs in languages like Dafny and Verus.

AI/ML arXiv cs.AI

Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas

Research proposing a new utility function (AFP) to induce mutual cooperation among agents in sequential social dilemmas using reinforcement learning.

Cybersecurity arXiv cs.AI

FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents

Presentation of FORGE, an attack method that hijacks the reasoning trajectory of deep research agents via planning-layer poisoning.