AI/ML arXiv cs.AI

ARIADNE: Agnostic Routing for Inference-time Adapter DyNamic sElection

Introduces ARIADNE, a training-free routing framework that dynamically selects the best PEFT adapter at inference time using embedding centroids.

Software Engineering Hacker News

Nim Conf 2026 (Online, Sat June 20)

Announcement of the Nim Conference 2026, scheduled for June 20th.

Open Source Hacker News

SteamOS Linux 3.8 released as stable

SteamOS Linux 3.8 has been released as a stable version.

AI/ML Hacker News

Show HN: Local personal data redaction for any AI tools

A new tool for local personal data redaction is introduced to help users maintain privacy when using AI tools.

AI/ML arXiv cs.AI

R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning

Introduction of R2D-RL, a reinforcement learning environment connecting robot soccer simulations to Python-based MARL workflows.

AI/ML arXiv cs.AI

ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch

ProfiLLM is an agentic LLM data pipeline designed to optimize user profiling for industrial ride-hailing dispatch systems.

AI/ML arXiv cs.AI

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

The WorldLines benchmark and ObsMem framework are introduced to improve long-horizon stateful memory for embodied agents in household settings.

AI/ML arXiv cs.AI

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

Xcientist is a research harness that externalizes AI synthesis and validation processes to ensure scientific accountability and traceability.

AI/ML arXiv cs.AI

Generative-Model Predictive Planning for Navigation in Partially Observable Environments

BeliefDiffusion combines diffusion models and Model Predictive Control for robust navigation in partially observable environments.

AI/ML Hacker News

Local Qwen isn't a worse Opus, it's a different tool

A discussion on the differing utility of local LLMs like Qwen compared to high-end proprietary models like Claude Opus, emphasizing that local models are different tools for different needs.

Hardware/Chips Hacker News

I restarted a 10 year old Xeon 174 times to delete 12 flags and gain 4 TPS

A technical deep-dive into optimizing a 10-year-old Xeon processor by deleting specific flags to marginally increase transactions per second (TPS).

AI/ML arXiv cs.AI

NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation

NAVI-Orbital demonstrates the first in-orbit autonomous Earth observation using a local vision-language model (Gemma 3) and LangGraph to reduce downlink bandwidth via semantic compression.

AI/ML arXiv cs.AI

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

Introduction of CaVe-VLM-CoT, a modular reflection-based agentic-RAG framework designed to reduce hallucinations in vision-language models through grounded reasoning.

AI/ML arXiv cs.AI

Searching for Synergy in Shared Workspace Human-AI Collaboration

A study on shared-workspace human-AI collaboration, exploring how structured coordination and simulated HITL gates improve performance in complex tasks.

AI/ML arXiv cs.AI

CEO-Bench: Can Agents Play the Long Game?

CEO-Bench evaluates long-horizon agent capabilities by simulating the operation of a startup over 500 days via a programmable Python interface.

AI/ML arXiv cs.AI

DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models

DeFAb is a verifiable benchmark for defeasible abduction in foundation models, testing the ability to construct hypotheses that explain anomalies while maintaining logical rigor.

Other arXiv cs.AI

Optimizing Lithium Production Decisions under Geological, Demand, and Pricing Uncertainties: A POMDP Framework for Multi-Objective Decision Making

A POMDP framework for optimizing lithium production decisions by managing geological, demand, and pricing uncertainties.

AI/ML arXiv cs.AI

ForecastBench-Sim: A Simulated-World Forecasting Benchmark

ForecastBench-Sim uses game rollouts from Freeciv to create a simulated-world forecasting benchmark for studying probabilistic reasoning in AI.

AI/ML arXiv cs.AI

What Must Generalist Agents Remember?

A formal theoretical account of what generalist agents must store in memory to maintain near-optimal performance across diverse environments.

Tech Business/VC Hacker News

Continue has been acquired by Cursor

The AI-powered coding assistant Continue has been acquired by Cursor.