AI/ML arXiv cs.AI

From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

Researchers present a factorised study on probe-based uncertainty estimation in LLMs to detect hallucinations, finding that structured features are more robust under distribution shift.

AI/ML arXiv cs.AI

CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence

CBD is introduced as an API-only black-box unlearning framework to remove sensitive or harmful data from LLMs without requiring internal model access.

AI/ML arXiv cs.AI

Mitigating LLM-based p-Hacking by Preregistering for the Next LLM

A new protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible future model released.

AI/ML arXiv cs.AI

Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting

Study on multi-horizon volatility forecasting showing that inference-time deployment policies and rollout rules can significantly impact accuracy and cost.

AI/ML arXiv cs.AI

Halt Fast! Early Stopping for Certified Robustness

A meta-learning framework for certified robustness that uses a lightweight learner to reduce sample complexity by 20-fold in Randomized Smoothing.

AI/ML arXiv cs.AI

Class-frequency Guided Noise Schedule for Diffusion Models

The Class-frequency Guided (CFRG) noise schedule is proposed for diffusion models to improve the generation quality of low-frequency classes in imbalanced datasets.

AI/ML arXiv cs.AI

What Was That Again? Certified Robustness for Automatic Speech Recognition

A certification-inspired diagnostic pipeline for Automatic Speech Recognition (ASR) that reduces Word Error Rate and improves acoustic security.

Cybersecurity arXiv cs.AI

Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

A high-throughput simulation framework for over-the-air acoustic attacks on voice control systems, highlighting risks in Whisper and wav2vec models.

AI/ML arXiv cs.AI

Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning

Research on safer empathetic fine-tuning for LLMs using a persona-driven rewriting pipeline to reduce jailbreak susceptibility without sacrificing warmth.

AI/ML arXiv cs.AI

The Simulacrum: Decision-Theoretic Pretraining for Near-Optimal Time-Series Forecasting and Inference

Introduces 'The Simulacrum', a decision-theoretic pretraining framework for neural time-series estimators that outperforms traditional statistical baselines.

AI/ML arXiv cs.AI

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF

Introduces Retroactive Advantage Correction (RAC) to address asynchronous reward signals in RLHF, reducing policy bias in production environments.

AI/ML arXiv cs.AI

SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction

Presents SceneBot, a motion-tracking framework for humanoid robots that unifies free-space locomotion and contact-rich interactions with the environment.

AI/ML arXiv cs.AI

CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance

Introduces CoIn, a framework combining 2D inpainting diffusion models and 3D Gaussian Splatting for consistent 3D scene reconstruction and editing.

AI/ML arXiv cs.AI

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

Proposes Fox, a training-free inference-time framework that reduces hallucinations in Large Vision-Language Models by severing pathological shortcut paths in attention.

AI/ML arXiv cs.AI

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing

Explores Narrative-UFET, using automatically generated short narratives to improve ultra-fine entity typing, especially for long-tail types.

AI/ML arXiv cs.AI

Global Explanations for Multivariate Time Series Forecasting Models via $K$-Order Markov Approximations

Introduces KARMA, a method for providing global explanations for multivariate time series forecasting models using K-order Markov approximations.

AI/ML arXiv cs.AI

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

Proposes HybridCodec, a speech language model architecture that combines discrete tokens with continuous residuals to improve speaker characteristic retention.

AI/ML arXiv cs.AI

Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining

Presents a dual-threshold hard example mining strategy to improve the cross-platform performance of Chinese offensive comment detection.

Other arXiv cs.AI

Reconstructing the Developmental Trajectory of Adipocytes in Human Adipose Tissue Using Single-Cell RNA Sequencing

Uses single-cell RNA sequencing to map the developmental trajectory of human adipocytes, identifying key signaling pathways for obesity treatment.

AI/ML arXiv cs.AI

Explainable AI for Biodiversity Monitoring and Ecological Image Analysis

Argues for the integration of Explainable AI (XAI) into biodiversity monitoring to ensure ecological models are based on meaningful signals rather than artifacts.