All Articles
17509 articles total
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
Research into the 'Knowing--Using Gap' in LLMs reveals that memorized knowledge often fails to generalize because it is not routed to computation-effective layers.
Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination
G-Frame is a multi-agent framework using game theory to reduce LLM hallucinations in scientific domains, leading to the creation of the OmniChem model for chemistry.
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
OmniFood-Bench is a new benchmark for evaluating VLMs on nutrient reasoning and health advice, revealing significant gaps in mass estimation and safety-critical advice.
Harman and Dr. Sean Olive are reshaping headphone sound
A discussion on how Harman and Dr. Sean Olive are influencing the standards for headphone sound quality and target curves.
AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution
Introduces AutoPersonas, a framework designed to prevent 'self-locking' in long-term AI persona agents by separating environment and persona states.
Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation
A study on using multiple AI teachers to create a verifiable curriculum for student models, finding that RLVR is more effective than SFT for coding improvement.
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
Introduces MentalHospital, a virtual environment and MentalEval evaluators for assessing LLM performance in psychiatric clinical encounters.
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment
Explores on-device distillation from reasoning teachers (like DeepSeek-R1) to small students, finding that reasoning capability transfers more effectively than scale.
PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs
Presents PolyUQuest, a verifiable, structure-aware web RAG framework using heterogeneous graphs to improve retrieval from HTML-based web pages.
Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench
Introduces PredicateLongBench, a benchmark to stress-test long-context reasoning by requiring models to identify contiguous subsequences satisfying specific predicates.
Psychological Competence as a Missing Dimension in AI Evaluation
Proposes 'psychological competence' as a critical missing dimension for evaluating human-facing AI systems beyond technical performance.
INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis
Proposes the INTENT framework using LSTM to predict vehicle intentions at intersections for autonomous driving safety.
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Introduces blind-spots-bench, a diagnostic benchmark to expose tasks that are trivial for humans but challenging for modern AI models.
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
AegisDx is a safety-oriented framework that uses specialized LLM components and verification gates to improve the accuracy and safety of AI-assisted clinical differential diagnosis.
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Research shows that agreement between different LLMs or within a single model's samples is a weak predictor of correctness and can be driven by shared biases.
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
The study finds that Chain-of-Thought monitoring can be bypassed by persuasion attacks, but combining different model families for fact-checking and monitoring reduces this vulnerability.
PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction
PARA-PV is a physics-aware retrieval-augmented framework for photovoltaic power forecasting that combines physical knowledge with time-series foundation models.
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
CausalDS is a new benchmark for evaluating the causal reasoning capabilities of data-science agents using synthetic natural-language stories and structural causal models.
Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models
A new neurosymbolic reasoning methodology integrates answer set programming (ASP) with energy-based models for robust end-to-end training in dynamic domains.
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
The 'overthinking' technique uses reasoning task vectors to amplify a model's propensity to think out loud, making it easier to extract hidden secrets from black-box models.