AI/ML arXiv cs.AI

Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation

A new framework called Rubric-Conditioned Self-Distillation that uses fine-grained rubrics instead of scalar rewards to improve the reasoning of LLMs.

Software Engineering Hacker News

I Hate Compilers

A Hacker News discussion thread centered around the frustrations and complexities of compiler design and usage.

Software Engineering Hacker News

Magic Buffers and io_uring Registered Buffers

A technical exploration of 'magic buffers' and the use of registered buffers within the io_uring interface for high-performance I/O in Linux.

Software Engineering Hacker News

Scaling opencomputer from 1 VM to 1 million sandboxes

An analysis of scaling the 'opencomputer' project from a single VM to a million sandboxes, focusing on isolation and scale.

AI/ML arXiv cs.AI

Skill-Guided Continuation Distillation for GUI Agents

Proposes Skill-Guided Continuation Distillation (SGCD) to improve GUI agents by providing supervision for off-trajectory states during execution.

AI/ML arXiv cs.AI

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Introduces SciRisk-Bench, a safety benchmark for AI in science that evaluates LLMs across various scientific disciplines and risk dimensions.

AI/ML arXiv cs.AI

Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents

Presents Decoupled Search Grounding (DSG), a vendor-agnostic architecture that separates search retrieval from LLM reasoning to reduce cost and latency.

AI/ML arXiv cs.AI

RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

Introduces RTSGameBench and RTSGameAgent, utilizing the game 'Beyond All Reason' to evaluate and improve strategic reasoning in Vision-Language Models.

AI/ML arXiv cs.AI

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection

Presents ThinkDeception, an interpretable multimodal framework for deception detection using MLLMs and a progressive RL training strategy.

AI/ML arXiv cs.AI

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

Proposes RODS, a reward-driven online data synthesis method to prevent sample depletion in multi-turn tool-use RL for AI agents.

AI/ML arXiv cs.AI

ARIADNE: Agnostic Routing for Inference-time Adapter DyNamic sElection

Introduces ARIADNE, a training-free routing framework that dynamically selects the best PEFT adapter at inference time using embedding centroids.

Software Engineering Hacker News

Nim Conf 2026 (Online, Sat June 20)

Announcement of the Nim Conference 2026, scheduled for June 20th.

Open Source Hacker News

SteamOS Linux 3.8 released as stable

SteamOS Linux 3.8 has been released as a stable version.

AI/ML Hacker News

Show HN: Local personal data redaction for any AI tools

A new tool for local personal data redaction is introduced to help users maintain privacy when using AI tools.

AI/ML arXiv cs.AI

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

Introduction of R2D-RL, a reinforcement learning environment connecting robot soccer simulations to Python-based MARL workflows.

AI/ML arXiv cs.AI

ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch

ProfiLLM is an agentic LLM data pipeline designed to optimize user profiling for industrial ride-hailing dispatch systems.

AI/ML arXiv cs.AI

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

The WorldLines benchmark and ObsMem framework are introduced to improve long-horizon stateful memory for embodied agents in household settings.

AI/ML arXiv cs.AI

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Xcientist is a research harness that externalizes AI synthesis and validation processes to ensure scientific accountability and traceability.

AI/ML arXiv cs.AI

Generative-Model Predictive Planning for Navigation in Partially Observable Environments

BeliefDiffusion combines diffusion models and Model Predictive Control for robust navigation in partially observable environments.

AI/ML Hacker News

Local Qwen isn't a worse Opus, it's a different tool

A discussion on the differing utility of local LLMs like Qwen compared to high-end proprietary models like Claude Opus, emphasizing that local models are different tools for different needs.