All Articles
19192 articles total
AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agent Collaboration
AdaSTORM is a new framework that enables LLMs to reason on thousand-node dynamic graphs by using adaptive partitioning and multi-agent collaboration.
Exploiting Search in Symbolic Numeric Planning with Patterns
This paper presents a numeric planning procedure based on Symbolic Pattern Planning (SPP) that uses symbolic search to refine patterns for better goal reachability.
Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery
A phase-aware guidance injection framework for recurrent MAPPO is proposed to improve disruption recovery in industrial assembly lines using external knowledge.
Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules
Medical Heuristic Learning (MHL) is an LLM-driven framework that produces interpretable, versioned Python decision rules for clinical tabular prediction instead of black-box models.
Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection
An audit of LLM hotel recommendations reveals that guest ratings and price dominate selections, while list position also significantly influences outcomes.
Show HN: SharkClean MCP
A Show HN post introducing SharkClean MCP, a tool likely related to Model Context Protocol (MCP) for cleaning data or context.
Thinking with Visual Grounding
Introduces visually grounded thinking for Vision-Language Models (VLMs), allowing models to interleave natural-language reasoning with explicit point or box groundings of visual evidence.
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
VibeThinker-3B is a compact 3B parameter model that achieves frontier-level verifiable reasoning performance, matching larger models like DeepSeek V3.2 on complex tasks.
LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis
LiteOdyssey is a lightweight reasoning AI agent framework designed for interpretable rare-disease diagnosis using a clinical genetics workflow without requiring large-scale fine-tuning.
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
The authors identify a 'Quality-Utility Paradox' where high-reward data from stronger models can impair small model mathematical reasoning due to distributional drift, proposing Style-Aligned Refinement as a solution.
AI Pluralism and the Worlds It Misses
Explores AI pluralism through the lens of 'ontological flattening' and proposes the Pluralistic Lifecycle Governance (PLG) framework for auditing AI systems' epistemic inclusion.
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
Introduces TimeVista, a VLM-as-a-Judge benchmark that uses Vision-Language Models to evaluate time series forecasting by analyzing plots and textual information.
PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums
Introduces PAL-Bench, a controlled benchmark for evidence-grounded profile reconstruction from longitudinal personal albums, highlighting the gap between summarization and faithful social reconstruction.
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
A diagnostic study arguing that LLM tutoring benchmarks should separate task-solving ability from pedagogy-oriented learning support to better measure educational impact.
Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients
Proposes OQ-TSAE, a framework for sensor-conditioned representation learning that ensures latent geometry preserves sensing-justified scene distinctions while suppressing nuisance factors.
Understanding the rationale behind a rule when trying to circumvent it
A discussion on the importance of understanding the underlying rationale of rules before attempting to find ways around them.
LLM-as-Code Agentic Programming for Agent Harness
Proposes 'Agentic Programming', a paradigm where a deterministic program governs control flow and the LLM acts as an adaptive component ('LLM-as-Code') to reduce hallucinations and token explosion.
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
Introduces UrbanWell, a large-scale benchmark for evaluating the spatio-temporal reasoning capabilities of multimodal LLMs in urban wellbeing analytics.
Agentic Framework for Deep Learning workload migration via In-Context Learning
Presents an autonomous agentic system using In-Context Learning and an execution oracle to automate the migration of deep learning models from PyTorch to JAX.
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
Explores the ability of LLMs to generate mathematical equations from scientific texts and evaluates the gap between LLM-based and human judgments of semantic accuracy.