All Articles
17337 articles total
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Introduces blind-spots-bench, a diagnostic benchmark to expose tasks that are trivial for humans but challenging for modern AI models.
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
AegisDx is a safety-oriented framework that uses specialized LLM components and verification gates to improve the accuracy and safety of AI-assisted clinical differential diagnosis.
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Research shows that agreement between different LLMs or within a single model's samples is a weak predictor of correctness and can be driven by shared biases.
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
The study finds that Chain-of-Thought monitoring can be bypassed by persuasion attacks, but combining different model families for fact-checking and monitoring reduces this vulnerability.
PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction
PARA-PV is a physics-aware retrieval-augmented framework for photovoltaic power forecasting that combines physical knowledge with time-series foundation models.
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
CausalDS is a new benchmark for evaluating the causal reasoning capabilities of data-science agents using synthetic natural-language stories and structural causal models.
Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models
A new neurosymbolic reasoning methodology integrates answer set programming (ASP) with energy-based models for robust end-to-end training in dynamic domains.
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
The 'overthinking' technique uses reasoning task vectors to amplify a model's propensity to think out loud, making it easier to extract hidden secrets from black-box models.
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
ASMR is an agentic framework that automatically generates schemas for ship maintenance reports by extracting semantic concepts and optimizing the structure via RL.
A First-Principles Theory of Slow Thinking and Active Perception
This paper proposes a mathematical theory of 'active lifting' to provide a first-principles formulation of slow thinking and active perception in LLMs.
Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction
ZendoWorld is an interactive environment designed to challenge AI agents to perform active visual concept induction and test hypotheses about hidden logical rules.
What Big Food Did to Ice Cream
An article discussing the industrialization and changes in the ice cream industry.
Common prefix skipping, adaptive sort
A discussion on common prefix skipping and adaptive sorting algorithms.
After Apple, India’s smartphone manufacturing boom enters new phase with Vivo JV
Vivo is establishing a joint venture in India, potentially creating a new template for Chinese smartphone manufacturers.
Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Introduction of Feedback Manipulation Regularization (FMR), an algorithm-agnostic method to improve imitation learning alignment in sequential decision-making environments.
Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer
The release of the Nigeria Machinery Usage and Failures Dataset and a method for building chain-of-thought reasoning examples for industrial machinery.
Persona Cartography: Charting Language Model Personality Traits in Weight Space
Research on 'Persona Cartography,' using the OCEAN framework to map and control LLM personality traits in weight space via low-rank adapters.
Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies
A study on optimizing frame rates and neural network architectures (LSTM/GRU) for the automated detection of autism-related self-stimulatory behaviors from video.
Agentic Neural Architecture Search
AgentNAS is a new pipeline that combines LLM-driven seed architecture generation with NAS-driven search to optimize neural architectures without manual engineering.
Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models
Introduction of Concretized Proposition Prompting (CPP) to resolve the tension between compositionality and knowledge in LLM reasoning.