All Articles
17327 articles total
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
OmniFood-Bench is a new benchmark for evaluating VLMs on nutrient reasoning and health advice, revealing significant gaps in mass estimation and safety-critical advice.
Harman and Dr. Sean Olive are reshaping headphone sound
A discussion on how Harman and Dr. Sean Olive are influencing the standards for headphone sound quality and target curves.
AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution
Introduces AutoPersonas, a framework designed to prevent 'self-locking' in long-term AI persona agents by separating environment and persona states.
Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation
A study on using multiple AI teachers to create a verifiable curriculum for student models, finding that RLVR is more effective than SFT for coding improvement.
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
Introduces MentalHospital, a virtual environment and MentalEval evaluators for assessing LLM performance in psychiatric clinical encounters.
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment
Explores on-device distillation from reasoning teachers (like DeepSeek-R1) to small students, finding that reasoning capability transfers more effectively than scale.
PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs
Presents PolyUQuest, a verifiable, structure-aware web RAG framework using heterogeneous graphs to improve retrieval from HTML-based web pages.
Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench
Introduces PredicateLongBench, a benchmark to stress-test long-context reasoning by requiring models to identify contiguous subsequences satisfying specific predicates.
Psychological Competence as a Missing Dimension in AI Evaluation
Proposes 'psychological competence' as a critical missing dimension for evaluating human-facing AI systems beyond technical performance.
INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis
Proposes the INTENT framework using LSTM to predict vehicle intentions at intersections for autonomous driving safety.
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Introduces blind-spots-bench, a diagnostic benchmark to expose tasks that are trivial for humans but challenging for modern AI models.
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
AegisDx is a safety-oriented framework that uses specialized LLM components and verification gates to improve the accuracy and safety of AI-assisted clinical differential diagnosis.
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Research shows that agreement between different LLMs or within a single model's samples is a weak predictor of correctness and can be driven by shared biases.
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
The study finds that Chain-of-Thought monitoring can be bypassed by persuasion attacks, but combining different model families for fact-checking and monitoring reduces this vulnerability.
PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction
PARA-PV is a physics-aware retrieval-augmented framework for photovoltaic power forecasting that combines physical knowledge with time-series foundation models.
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
CausalDS is a new benchmark for evaluating the causal reasoning capabilities of data-science agents using synthetic natural-language stories and structural causal models.
Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models
A new neurosymbolic reasoning methodology integrates answer set programming (ASP) with energy-based models for robust end-to-end training in dynamic domains.
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
The 'overthinking' technique uses reasoning task vectors to amplify a model's propensity to think out loud, making it easier to extract hidden secrets from black-box models.
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
ASMR is an agentic framework that automatically generates schemas for ship maintenance reports by extracting semantic concepts and optimizing the structure via RL.
A First-Principles Theory of Slow Thinking and Active Perception
This paper proposes a mathematical theory of 'active lifting' to provide a first-principles formulation of slow thinking and active perception in LLMs.