All Articles
17556 articles total
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
DUPO is proposed to handle delayed feedback in reinforcement learning using diffusion models to estimate state discrepancies.
Historic Photos of NASA's Cavernous Wind Tunnels
A collection of historic photographs documenting NASA's large-scale wind tunnel facilities.
Show HN: Fast, native Mac file manager (filters, fuzzy find, 9 MB, no Electron)
A new native Mac file manager featuring fuzzy finding and filters, built without Electron to maintain a small footprint (9 MB).
Why migrants come to Germany for work and then leave again
An exploration of the factors causing migrants to leave Germany after initially arriving for work.
Are you ready for what it takes to stop ghost guns?
An investigation into the use of 3D printers to manufacture illegal firearm components and automatic weapon switches.
Nothing’s first B-series phone is also skipping the US
Nothing launches the Phone 4B, a budget-tier device skipping the US market.
Nothing’s new earbuds can record calls and what you’re listening to
Nothing introduces new budget-friendly Ear 3A earbuds with the ability to record audio directly on the device.
AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization
AgenticPD is a stage-aware agentic framework designed to optimize physical design Quality-of-Results (QoR) in EDA flows.
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
CARL is a novel reinforcement learning framework that improves LLM planning by strengthening intrinsic constraint awareness.
Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation
Medi-Gemma is a hybrid clinical decision support system that integrates deterministic EMR analytics with RAG to reduce hallucinations in healthcare.
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
STAPO introduces a hierarchical RL framework using normalized entropy to solve the problem of trajectory neglect in long-horizon LLM agent training.
Microsoft Can Track Users via a Windows Device ID
Discussion on Microsoft's ability to track Windows users via a unique device ID.
In Praise of Observational Evidence
An exploration and praise of the importance of observational evidence in scientific or technical contexts.
Dropping in on Gottfried Leibniz (2013)
A retrospective look or discussion involving the philosopher and mathematician Gottfried Leibniz.
The Art of Computer Programming by Donald E. Knuth
Discussion regarding Donald Knuth's seminal work, 'The Art of Computer Programming'.
The first American autonomous ground vehicles are fighting in Ukraine
Forterra has deployed over 100 autonomous ground vehicles for combat operations in Ukraine.
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
Introduction of MRMS, a multi-resolution memory substrate designed to provide continuity and personalization for long-lived AI agents.
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
Formal Disco is a distributed system for generating large-scale synthetic datasets of formally verified programs in languages like Dafny and Verus.
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
Research proposing a new utility function (AFP) to induce mutual cooperation among agents in sequential social dilemmas using reinforcement learning.
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
Presentation of FORGE, an attack method that hijacks the reasoning trajectory of deep research agents via planning-layer poisoning.