AI/ML arXiv cs.AI

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Research into the 'Knowing--Using Gap' in LLMs reveals that memorized knowledge often fails to generalize because it is not routed to computation-effective layers.

AI/ML arXiv cs.AI

Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination

G-Frame is a multi-agent framework using game theory to reduce LLM hallucinations in scientific domains, leading to the creation of the OmniChem model for chemistry.

AI/ML arXiv cs.AI

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

OmniFood-Bench is a new benchmark for evaluating VLMs on nutrient reasoning and health advice, revealing significant gaps in mass estimation and safety-critical advice.

Other Hacker News

Harman and Dr. Sean Olive are reshaping headphone sound

A discussion on how Harman and Dr. Sean Olive are influencing the standards for headphone sound quality and target curves.

AI/ML arXiv cs.AI

AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution

Introduces AutoPersonas, a framework designed to prevent 'self-locking' in long-term AI persona agents by separating environment and persona states.

AI/ML arXiv cs.AI

Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation

A study on using multiple AI teachers to create a verifiable curriculum for student models, finding that RLVR is more effective than SFT for coding improvement.

AI/ML arXiv cs.AI

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Introduces MentalHospital, a virtual environment and MentalEval evaluators for assessing LLM performance in psychiatric clinical encounters.

AI/ML arXiv cs.AI

Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

Explores on-device distillation from reasoning teachers (like DeepSeek-R1) to small students, finding that reasoning capability transfers more effectively than scale.

AI/ML arXiv cs.AI

PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

Presents PolyUQuest, a verifiable, structure-aware web RAG framework using heterogeneous graphs to improve retrieval from HTML-based web pages.

AI/ML arXiv cs.AI

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

Introduces PredicateLongBench, a benchmark to stress-test long-context reasoning by requiring models to identify contiguous subsequences satisfying specific predicates.

AI/ML arXiv cs.AI

Psychological Competence as a Missing Dimension in AI Evaluation

Proposes 'psychological competence' as a critical missing dimension for evaluating human-facing AI systems beyond technical performance.

AI/ML arXiv cs.AI

INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis

Proposes the INTENT framework using LSTM to predict vehicle intentions at intersections for autonomous driving safety.

AI/ML arXiv cs.AI

Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

Introduces blind-spots-bench, a diagnostic benchmark to expose tasks that are trivial for humans but challenging for modern AI models.

AI/ML arXiv cs.AI

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

AegisDx is a safety-oriented framework that uses specialized LLM components and verification gates to improve the accuracy and safety of AI-assisted clinical differential diagnosis.

AI/ML arXiv cs.AI

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

Research shows that agreement between different LLMs or within a single model's samples is a weak predictor of correctness and can be driven by shared biases.

AI/ML arXiv cs.AI

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

The study finds that Chain-of-Thought monitoring can be bypassed by persuasion attacks, but combining different model families for fact-checking and monitoring reduces this vulnerability.

AI/ML arXiv cs.AI

PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction

PARA-PV is a physics-aware retrieval-augmented framework for photovoltaic power forecasting that combines physical knowledge with time-series foundation models.

AI/ML arXiv cs.AI

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

CausalDS is a new benchmark for evaluating the causal reasoning capabilities of data-science agents using synthetic natural-language stories and structural causal models.

AI/ML arXiv cs.AI

Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models

A new neurosymbolic reasoning methodology integrates answer set programming (ASP) with energy-based models for robust end-to-end training in dynamic domains.

AI/ML arXiv cs.AI

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

The 'overthinking' technique uses reasoning task vectors to amplify a model's propensity to think out loud, making it easier to extract hidden secrets from black-box models.