All Articles
17108 articles total
Pinwheel launches a retro-inspired landline phone for kids
Pinwheel has launched a retro-inspired landline phone for children to reduce smartphone distractions.
New York becomes the first state to enact a data center moratorium
New York has implemented the first statewide moratorium on new hyperscale data centers over 50 megawatts to address energy and environmental concerns.
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
A research paper proposing a roadmap for autonomous medical agents, emphasizing clinical environment scaling and self-evolution over simple parameter scaling.
SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
Introduction of ScaleCUA, a framework for scaling computer use agents using verifiable task synthesis and efficient online reinforcement learning.
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
Research identifying two distinct planning abilities in LLMs: operational reasoning and structural enumeration, arguing that scaling alone does not improve the latter.
PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection
PREF-Gate is presented as an auditable decision framework for graph fraud detection that constraints relational evidence based on provenance.
Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory
A study on heterogeneous agent cohorts using specialized roles (Disrupter, Validator, Broker) and 'Scars' (cached constraint patches) for safe open-ended exploration.
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
Initial validation of the ARC-AGI benchmark as a measure of human fluid intelligence, supporting the role of rule induction in cognitive ability.
Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought
A paper diagnosing latent inefficiency in Chain-of-Thought reasoning and introducing CAID and PACE to compress reasoning chains without losing accuracy.
Alternative(s) to run CUDA on non-Nvidia hardware
A discussion on alternative ways to run CUDA-like workloads on non-Nvidia hardware, focusing on compatibility layers and hardware abstractions.
Zero Knowledge Tolstoyan Art
An exploration of using Zero Knowledge proofs in the creation and verification of art, blending cryptography and creativity.
Indian scientists produce most detailed 3D atlas of the human brainstem
Indian scientists have developed a highly detailed 3D atlas of the human brainstem, advancing neurological mapping.
The Anatomy of an Instruction Pipeline Hazard
A technical deep-dive into the anatomy of instruction pipeline hazards in CPU architecture.
Show HN: Benchmark your eng team's AI agent maturity in 5 minutes
A tool designed to benchmark the AI agent maturity of engineering teams within a short timeframe.
OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping
OS-Pruner is a lightweight framework that treats Chain-of-Thought pruning as an optimal stopping problem to reduce LLM latency and cost.
A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery
Proposes a hierarchical, stack-based architecture for LLM agents to manage tool orchestration and prevent decision-space explosion.
NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management
NextFund is a performance tracking platform for agentic portfolio management, providing observability into the decision paths of financial AI agents.
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
Introduces AgentFootprint, a benchmark to measure the persistent storage costs and data footprints left by LLM agents.
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
STAMP introduces provenance-guided credit assignment to improve the training of deep search agents by targeting rewards to the actions that surface critical evidence.
First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)
Researchers extended a methodology for embedding first-order modal logic into Isabelle/HOL, providing automated faithfulness proofs and a mechanization of the downward Lowenheim-Skolem theorem.