All Articles
18727 articles total
Show HN: SharkClean MCP
A Show HN post introducing SharkClean MCP, a tool likely related to Model Context Protocol (MCP) for cleaning data or context.
Thinking with Visual Grounding
Introduces visually grounded thinking for Vision-Language Models (VLMs), allowing models to interleave natural-language reasoning with explicit point or box groundings of visual evidence.
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
VibeThinker-3B is a compact 3B parameter model that achieves frontier-level verifiable reasoning performance, matching larger models like DeepSeek V3.2 on complex tasks.
LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis
LiteOdyssey is a lightweight reasoning AI agent framework designed for interpretable rare-disease diagnosis using a clinical genetics workflow without requiring large-scale fine-tuning.
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
The authors identify a 'Quality-Utility Paradox' where high-reward data from stronger models can impair small model mathematical reasoning due to distributional drift, proposing Style-Aligned Refinement as a solution.
AI Pluralism and the Worlds It Misses
Explores AI pluralism through the lens of 'ontological flattening' and proposes the Pluralistic Lifecycle Governance (PLG) framework for auditing AI systems' epistemic inclusion.
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
Introduces TimeVista, a VLM-as-a-Judge benchmark that uses Vision-Language Models to evaluate time series forecasting by analyzing plots and textual information.
PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums
Introduces PAL-Bench, a controlled benchmark for evidence-grounded profile reconstruction from longitudinal personal albums, highlighting the gap between summarization and faithful social reconstruction.
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
A diagnostic study arguing that LLM tutoring benchmarks should separate task-solving ability from pedagogy-oriented learning support to better measure educational impact.
Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients
Proposes OQ-TSAE, a framework for sensor-conditioned representation learning that ensures latent geometry preserves sensing-justified scene distinctions while suppressing nuisance factors.
Understanding the rationale behind a rule when trying to circumvent it
A discussion on the importance of understanding the underlying rationale of rules before attempting to find ways around them.
LLM-as-Code Agentic Programming for Agent Harness
Proposes 'Agentic Programming', a paradigm where a deterministic program governs control flow and the LLM acts as an adaptive component ('LLM-as-Code') to reduce hallucinations and token explosion.
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
Introduces UrbanWell, a large-scale benchmark for evaluating the spatio-temporal reasoning capabilities of multimodal LLMs in urban wellbeing analytics.
Agentic Framework for Deep Learning workload migration via In-Context Learning
Presents an autonomous agentic system using In-Context Learning and an execution oracle to automate the migration of deep learning models from PyTorch to JAX.
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
Explores the ability of LLMs to generate mathematical equations from scientific texts and evaluates the gap between LLM-based and human judgments of semantic accuracy.
Auditing Reward Hackability in Code RL Training Environments
Analyzes 'reward hackability' in Code RL environments, finding that many test suites allow incorrect patches to pass, and proposes a hardening procedure using a Docker gold-sanity gate.
Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games
Introduces Mind-Studio, a framework that uses LLMs to synthesize executable world models (pygame-style) from interaction trajectories for partially observable games.
Rhythm of the Deep: A Computational-Linguistic Test of Duality of Patterning in Sperm Whale Codas
A computational-linguistic study analyzing sperm whale codas to identify a two-tier architectural structure similar to the duality of patterning in human language.
RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation
Presents RecourseBench, a modular and reproducible evaluation framework for algorithmic recourse methods, providing a standardized way to compare counterfactual explanations.
Know Your Limits : On the Faithfulness of LLMs as Solvers and Autoformalizers in Legal Reasoning
Examines the faithfulness of LLMs in legal reasoning, identifying 'scope laundering' where models claim logical grounding without actually executing formal reasoning.