AI/ML arXiv cs.AI

scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology

scBench-Long is a new verifiable benchmark for long-horizon single-cell biology, testing an AI agent's ability to recover scientific conclusions from raw data.

AI/ML arXiv cs.AI

IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control

IDEA is a sim-to-real transfer method for multi-agent control that uses effect alignment to remain insensitive to dynamics mismatch.

Other Hacker News

Why have papers by one of history's most famous physicists been retracted?

Discussion regarding the retraction of scientific papers by a famous physicist.

Other Hacker News

A forgotten social media post may hold key clues to Covid-19's origin

Analysis of a social media post potentially providing clues about the origins of Covid-19.

AI/ML Hacker News

The AI backlash is only getting started

Opinion piece discussing the growing backlash against AI integration and adoption.

Tech Business/VC The Verge

Of course Meta thinks gambling is the future

Meta is reportedly developing a prediction market app, continuing its trend of cloning successful social mechanics.

Hardware/Chips Ars Technica

SpaceX plans to launch Starlink mobile service in the US

SpaceX plans to introduce Starlink mobile services in the US to compete in the mass-market phone business.

AI/ML arXiv cs.AI

Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs

Research on the 'Evaluation-Strategy Gap' in fault diagnosis for deep learning programs, highlighting how within-program cross-validation can be inadequate.

AI/ML arXiv cs.AI

Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge

Introduction of MemStrata, a retrieval memory system for AI agents that eliminates stale-fact errors in evolving knowledge bases without requiring LLM calls.

AI/ML arXiv cs.AI

Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting

A new framework combining Gated Bottleneck Latent ODEs and multi-path fine-tuning for forecasting biopharmaceutical cell culture processes.

AI/ML arXiv cs.AI

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

Researchers identify the 'Inattentional Gap,' where task-conditioned AI models ignore safety-critical signals they are otherwise capable of detecting.

AI/ML arXiv cs.AI

\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

DiARC is proposed as a method to improve LLM reasoning on ARC-like tasks by training models to distinguish between positive and negative samples.

Cybersecurity Hacker News

Incident CVE-2026-LGTM

Discussion regarding a security incident identified as CVE-2026-LGTM.

Other Hacker News

Ultrasound Imaging of the Brain

Discussion on the application of ultrasound imaging for brain visualization.

Software Engineering Hacker News

The best thing that has ever happened for multiplayer games

Community discussion regarding a significant positive development for multiplayer gaming architecture.

Tech Business/VC The Verge

Anthropic’s Mythos mess is only getting worse

Anthropic's Mythos-class models remain offline following a government ultimatum, with no clear timeline for their return.

AI/ML arXiv cs.AI

3D Spatial Pattern Matching

Researchers propose extending spatial pattern matching to 3D to better handle real-world entities with height, introducing a new subgraph matching algorithm.

AI/ML arXiv cs.AI

Localizing RL-Induced Tool Use to a Single Crosscoder Feature

The study introduces Dedicated Feature Crosscoders (DFC) to isolate RL-specific features in LLMs, enabling better control and capability transfer for tool-use behaviors.

AI/ML arXiv cs.AI

Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks

A study on Retrieval-Warmed Energy-Based Reasoning (RW-EBR) uses a five-arm ablation methodology to analyze diffusion-as-inference for structured reasoning tasks.

Cybersecurity arXiv cs.AI

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

Research evaluates 'out-of-band' deterministic defenses against prompt injection in LLM agents, finding them more resilient than 'in-band' detection.