All Articles
17562 articles total
The Art of Computer Programming by Donald E. Knuth
Discussion regarding Donald Knuth's seminal work, 'The Art of Computer Programming'.
The first American autonomous ground vehicles are fighting in Ukraine
Forterra has deployed over 100 autonomous ground vehicles for combat operations in Ukraine.
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
Introduction of MRMS, a multi-resolution memory substrate designed to provide continuity and personalization for long-lived AI agents.
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
Formal Disco is a distributed system for generating large-scale synthetic datasets of formally verified programs in languages like Dafny and Verus.
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
Research proposing a new utility function (AFP) to induce mutual cooperation among agents in sequential social dilemmas using reinforcement learning.
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
Presentation of FORGE, an attack method that hijacks the reasoning trajectory of deep research agents via planning-layer poisoning.
FM-ChangeNet: Learning Change through Pathwise Feature Transport
FM-ChangeNet introduces a pathwise-supervised framework for change detection in remote sensing by treating bi-temporal reasoning as continuous transport in feature space.
Lago (YC S21) Is Hiring for Our GTM Team
Lago is hiring for their Go-To-Market (GTM) team.
Inkfield
Discussion or announcement regarding Inkfield.
Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical Innovation
Research proposing that mathematical innovation in humans and AI relies on cross-domain patterns from the natural world rather than pure logical deduction.
Compressing the Validation Bottleneck: An Agentic Self-Driving Lab for Scientific Discovery
Introduction of an agentic self-driving lab (SDL) that optimizes scientific discovery by reducing the number and cost of validation experiments.
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
A method for optimizing the language conditioning of frozen Vision-Language-Action (VLA) models using RL to improve robot manipulation tasks.
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Analysis of how software-agent benchmarks' 'harnesses' can bias LLM agent beliefs, introducing a no-training protocol called BIWM to align belief trajectories.
Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
Introduction of the Heaviside Continuity of Rolling Coefficients (HCRC), a framework that uses a predicate-gated execution to eliminate errors in LLM reasoning.
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
Study on 'answer-driven reasoning' in AI tutors, using Truncated Chain-of-Thought Auditing (TRACE) to detect when models use hidden answers to fake reasoning.
Attention Limited Reward Learning
Research showing that RLHF reward modeling can be distorted by 'rational inattention,' where difficulty in detecting differences is mistaken for indifference.
Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority
Proposed 'Governed Individuation' framework to cryptographically decouple an agent's learning from its authority, ensuring safety invariants during deployment.
Dolosse – a South African invention used over the world
A discussion about Dolosse, a South African invention used worldwide for coastal protection.
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Introduction of HAS-Bench, a benchmark for evaluating Human-Agent Systems based on a graph-based framework representing humans and agents as first-class participants.
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Research investigating whether multimodal GUI agents rely more on visual pixels or structural data (DOM) when forming beliefs about interface states.