All Articles
18336 articles total
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
A new framework called Rubric-Conditioned Self-Distillation that uses fine-grained rubrics instead of scalar rewards to improve the reasoning of LLMs.
I Hate Compilers
A Hacker News discussion thread centered around the frustrations and complexities of compiler design and usage.
Magic Buffers and io_uring Registered Buffers
A technical exploration of 'magic buffers' and the use of registered buffers within the io_uring interface for high-performance I/O in Linux.
Scaling opencomputer from 1 VM to 1 million sandboxes
An analysis of scaling the 'opencomputer' project from a single VM to a million sandboxes, focusing on isolation and scale.
Skill-Guided Continuation Distillation for GUI Agents
Proposes Skill-Guided Continuation Distillation (SGCD) to improve GUI agents by providing supervision for off-trajectory states during execution.
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
Introduces SciRisk-Bench, a safety benchmark for AI in science that evaluates LLMs across various scientific disciplines and risk dimensions.
Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents
Presents Decoupled Search Grounding (DSG), a vendor-agnostic architecture that separates search retrieval from LLM reasoning to reduce cost and latency.
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
Introduces RTSGameBench and RTSGameAgent, utilizing the game 'Beyond All Reason' to evaluate and improve strategic reasoning in Vision-Language Models.
ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection
Presents ThinkDeception, an interpretable multimodal framework for deception detection using MLLMs and a progressive RL training strategy.
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
Proposes RODS, a reward-driven online data synthesis method to prevent sample depletion in multi-turn tool-use RL for AI agents.
ARIADNE: Agnostic Routing for Inference-time Adapter DyNamic sElection
Introduces ARIADNE, a training-free routing framework that dynamically selects the best PEFT adapter at inference time using embedding centroids.
Nim Conf 2026 (Online, Sat June 20)
Announcement of the Nim Conference 2026, scheduled for June 20th.
SteamOS Linux 3.8 released as stable
SteamOS Linux 3.8 has been released as a stable version.
Show HN: Local personal data redaction for any AI tools
A new tool for local personal data redaction is introduced to help users maintain privacy when using AI tools.
R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning
Introduction of R2D-RL, a reinforcement learning environment connecting robot soccer simulations to Python-based MARL workflows.
ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch
ProfiLLM is an agentic LLM data pipeline designed to optimize user profiling for industrial ride-hailing dispatch systems.
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
The WorldLines benchmark and ObsMem framework are introduced to improve long-horizon stateful memory for embodied agents in household settings.
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
Xcientist is a research harness that externalizes AI synthesis and validation processes to ensure scientific accountability and traceability.
Generative-Model Predictive Planning for Navigation in Partially Observable Environments
BeliefDiffusion combines diffusion models and Model Predictive Control for robust navigation in partially observable environments.
Local Qwen isn't a worse Opus, it's a different tool
A discussion on the differing utility of local LLMs like Qwen compared to high-end proprietary models like Claude Opus, emphasizing that local models are different tools for different needs.