AI/ML arXiv cs.AI

Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization

A new SCM-GRPO framework grounds multi-hop reasoning in structural causal models to reduce hallucinations in fact verification.

AI/ML arXiv cs.AI

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

BioMedArena is an open-source toolkit for the standardized building and evaluation of biomedical deep research agents.

AI/ML arXiv cs.AI

2.5-D Decomposition for LLM-Based Spatial Construction

A 2.5-D decomposition pipeline improves LLM spatial reasoning for autonomous construction by separating horizontal planning from vertical execution.

AI/ML arXiv cs.AI

Efficient Test-time Inference for Generative Planning Models with OCL Search

The OCL search algorithm improves generative planning models by optimizing the inference process through efficient exploration control.

Software Engineering Hacker News

15 sorting algorithms in 6 minutes (2013) [video]

A classic 2013 video demonstrating 15 different sorting algorithms through visual animations.

AI/ML arXiv cs.AI

InSight: Self-Guided Skill Acquisition via Steerable VLAs

InSight is a framework for vision-language-action (VLA) models to autonomously acquire new manipulation skills by making primitives steerable and using a VLM-guided data flywheel.

AI/ML arXiv cs.AI

Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data

Random Rule Forest (RRF) uses LLMs to generate simple YES/NO questions as weak learners for an interpretable, auditable ensemble predicting success from unstructured data.

AI/ML arXiv cs.AI

TIP-Search: Time-Predictable Inference Scheduling for Market Prediction under Uncertain Load

TIP-Search is a time-predictable inference scheduling system for market prediction that trades accuracy for deadline satisfaction under uncertain load.

AI/ML arXiv cs.AI

From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control

MERA is a meta-cognitive reasoning framework that decouples reasoning from control to reduce overthinking and improve efficiency in Large Reasoning Models (LRMs).

AI/ML arXiv cs.AI

A global log for medical AI

MedLog is a proposed universal protocol for event-level logging of medical AI interactions to enable auditing, bias detection, and performance monitoring.

AI/ML arXiv cs.AI

Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs

RILKE is a method for lifelong knowledge memory control in LLMs that uses representation-space interventions to update knowledge without full retraining.

AI/ML arXiv cs.AI

Evolving Programmatic Skill Networks

Programmatic Skill Networks (PSN) allow agents in embodied environments to build and evolve a library of executable symbolic programs for skill reuse and adaptation.

AI/ML arXiv cs.AI

BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding

BioPIE is a new biomedical protocol information extraction dataset providing procedure-centric knowledge graphs for better experimental understanding.

AI/ML arXiv cs.AI

LLM-MINE: Large Language Model based Alzheimer's Disease and Related Dementias Phenotypes Mining from Clinical Notes

LLM-MINE is a framework for automatically extracting Alzheimer's disease phenotypes from unstructured clinical notes using LLMs.

Other Hacker News

Writers and Drugs

A discussion thread regarding the intersection of writers and drug use.

AI/ML Hacker News

What I'm Finding About LLM Code Style and Token Costs

An analysis of how LLM code style and token costs vary and their implications for development.

AI/ML arXiv cs.AI

Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce

Proposes a micro-transaction market for verified product information in agentic e-commerce to improve trust and competition.

AI/ML arXiv cs.AI

Grad Detect: Gradient-Based Hallucination Detection in LLMs

Introduces Grad Detect, a gradient-based approach to detect LLM hallucinations by analyzing layer-wise gradient patterns during inference.

AI/ML arXiv cs.AI

EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

Presents the EG-VQA benchmark for verifiable video question answering and the EG-Reasoner model for grounded reasoning.

AI/ML arXiv cs.AI

OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis

Introduces OrbitForge, a framework that converts text-generated videos into consistent 3D Gaussian Splatting scenes.