AI/ML arXiv cs.AI

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

OmniFood-Bench is a new benchmark for evaluating VLMs on nutrient reasoning and health advice, revealing significant gaps in mass estimation and safety-critical advice.

Other Hacker News

Harman and Dr. Sean Olive are reshaping headphone sound

A discussion on how Harman and Dr. Sean Olive are influencing the standards for headphone sound quality and target curves.

AI/ML arXiv cs.AI

AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution

Introduces AutoPersonas, a framework designed to prevent 'self-locking' in long-term AI persona agents by separating environment and persona states.

AI/ML arXiv cs.AI

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

A study on using multiple AI teachers to create a verifiable curriculum for student models, finding that RLVR is more effective than SFT for coding improvement.

AI/ML arXiv cs.AI

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Introduces MentalHospital, a virtual environment and MentalEval evaluators for assessing LLM performance in psychiatric clinical encounters.

AI/ML arXiv cs.AI

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

Explores on-device distillation from reasoning teachers (like DeepSeek-R1) to small students, finding that reasoning capability transfers more effectively than scale.

AI/ML arXiv cs.AI

PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

Presents PolyUQuest, a verifiable, structure-aware web RAG framework using heterogeneous graphs to improve retrieval from HTML-based web pages.

AI/ML arXiv cs.AI

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Introduces PredicateLongBench, a benchmark to stress-test long-context reasoning by requiring models to identify contiguous subsequences satisfying specific predicates.

AI/ML arXiv cs.AI

Psychological Competence as a Missing Dimension in AI Evaluation

Proposes 'psychological competence' as a critical missing dimension for evaluating human-facing AI systems beyond technical performance.

AI/ML arXiv cs.AI

INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis

Proposes the INTENT framework using LSTM to predict vehicle intentions at intersections for autonomous driving safety.

AI/ML arXiv cs.AI

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Introduces blind-spots-bench, a diagnostic benchmark to expose tasks that are trivial for humans but challenging for modern AI models.

AI/ML arXiv cs.AI

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

AegisDx is a safety-oriented framework that uses specialized LLM components and verification gates to improve the accuracy and safety of AI-assisted clinical differential diagnosis.

AI/ML arXiv cs.AI

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

Research shows that agreement between different LLMs or within a single model's samples is a weak predictor of correctness and can be driven by shared biases.

AI/ML arXiv cs.AI

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

The study finds that Chain-of-Thought monitoring can be bypassed by persuasion attacks, but combining different model families for fact-checking and monitoring reduces this vulnerability.

AI/ML arXiv cs.AI

PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction

PARA-PV is a physics-aware retrieval-augmented framework for photovoltaic power forecasting that combines physical knowledge with time-series foundation models.

AI/ML arXiv cs.AI

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

CausalDS is a new benchmark for evaluating the causal reasoning capabilities of data-science agents using synthetic natural-language stories and structural causal models.

AI/ML arXiv cs.AI

Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

A new neurosymbolic reasoning methodology integrates answer set programming (ASP) with energy-based models for robust end-to-end training in dynamic domains.

AI/ML arXiv cs.AI

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

The 'overthinking' technique uses reasoning task vectors to amplify a model's propensity to think out loud, making it easier to extract hidden secrets from black-box models.

AI/ML arXiv cs.AI

ASMR: Agentic Schema Generation for Ship Maintenance Report Writing

ASMR is an agentic framework that automatically generates schemas for ship maintenance reports by extracting semantic concepts and optimizing the structure via RL.

AI/ML arXiv cs.AI

A First-Principles Theory of Slow Thinking and Active Perception

This paper proposes a mathematical theory of 'active lifting' to provide a first-principles formulation of slow thinking and active perception in LLMs.