AI/ML arXiv cs.AI

AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agent Collaboration

AdaSTORM is a new framework that enables LLMs to reason on thousand-node dynamic graphs by using adaptive partitioning and multi-agent collaboration.

AI/ML arXiv cs.AI

Exploiting Search in Symbolic Numeric Planning with Patterns

This paper presents a numeric planning procedure based on Symbolic Pattern Planning (SPP) that uses symbolic search to refine patterns for better goal reachability.

AI/ML arXiv cs.AI

Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery

A phase-aware guidance injection framework for recurrent MAPPO is proposed to improve disruption recovery in industrial assembly lines using external knowledge.

AI/ML arXiv cs.AI

Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules

Medical Heuristic Learning (MHL) is an LLM-driven framework that produces interpretable, versioned Python decision rules for clinical tabular prediction instead of black-box models.

AI/ML arXiv cs.AI

Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection

An audit of LLM hotel recommendations reveals that guest ratings and price dominate selections, while list position also significantly influences outcomes.

Software Engineering Hacker News

Show HN: SharkClean MCP

A Show HN post introducing SharkClean MCP, a tool likely related to Model Context Protocol (MCP) for cleaning data or context.

AI/ML arXiv cs.AI

Thinking with Visual Grounding

Introduces visually grounded thinking for Vision-Language Models (VLMs), allowing models to interleave natural-language reasoning with explicit point or box groundings of visual evidence.

AI/ML arXiv cs.AI

VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models

VibeThinker-3B is a compact 3B parameter model that achieves frontier-level verifiable reasoning performance, matching larger models like DeepSeek V3.2 on complex tasks.

AI/ML arXiv cs.AI

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis

LiteOdyssey is a lightweight reasoning AI agent framework designed for interpretable rare-disease diagnosis using a clinical genetics workflow without requiring large-scale fine-tuning.

AI/ML arXiv cs.AI

The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning

The authors identify a 'Quality-Utility Paradox' where high-reward data from stronger models can impair small model mathematical reasoning due to distributional drift, proposing Style-Aligned Refinement as a solution.

AI/ML arXiv cs.AI

AI Pluralism and the Worlds It Misses

Explores AI pluralism through the lens of 'ontological flattening' and proposes the Pluralistic Lifecycle Governance (PLG) framework for auditing AI systems' epistemic inclusion.

AI/ML arXiv cs.AI

TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting

Introduces TimeVista, a VLM-as-a-Judge benchmark that uses Vision-Language Models to evaluate time series forecasting by analyzing plots and textual information.

AI/ML arXiv cs.AI

PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums

Introduces PAL-Bench, a controlled benchmark for evidence-grounded profile reconstruction from longitudinal personal albums, highlighting the gap between summarization and faithful social reconstruction.

AI/ML arXiv cs.AI

Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact

A diagnostic study arguing that LLM tutoring benchmarks should separate task-solving ability from pedagogy-oriented learning support to better measure educational impact.

AI/ML arXiv cs.AI

Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients

Proposes OQ-TSAE, a framework for sensor-conditioned representation learning that ensures latent geometry preserves sensing-justified scene distinctions while suppressing nuisance factors.

Other Hacker News

Understanding the rationale behind a rule when trying to circumvent it

A discussion on the importance of understanding the underlying rationale of rules before attempting to find ways around them.

AI/ML arXiv cs.AI

LLM-as-Code Agentic Programming for Agent Harness

Proposes 'Agentic Programming', a paradigm where a deterministic program governs control flow and the LLM acts as an adaptive component ('LLM-as-Code') to reduce hallucinations and token explosion.

AI/ML arXiv cs.AI

UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics

Introduces UrbanWell, a large-scale benchmark for evaluating the spatio-temporal reasoning capabilities of multimodal LLMs in urban wellbeing analytics.

AI/ML arXiv cs.AI

Agentic Framework for Deep Learning workload migration via In-Context Learning

Presents an autonomous agentic system using In-Context Learning and an execution oracle to automate the migration of deep learning models from PyTorch to JAX.

AI/ML arXiv cs.AI

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

Explores the ability of LLMs to generate mathematical equations from scientific texts and evaluates the gap between LLM-based and human judgments of semantic accuracy.