AI/ML arXiv cs.AI

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

Introduces Symbolic Mechanistic Data Attribution (SMDA), a tool to trace high-level model behaviors back to specific training examples using SAE features.

AI/ML arXiv cs.AI

Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection

Presents Anomaly Factory 3D (AF3AD), a modular framework for synthesizing pseudo-anomalies to improve unsupervised 3D anomaly detection.

AI/ML arXiv cs.AI

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

Introduces AIOps2025 and RCA100, benchmarks for evaluating LLM agents in diagnosing microservice failures based on reasoning processes.

AI/ML arXiv cs.AI

Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering

Proposes MoRE (Mode Redirection), a method to distill desired behavior modes into policy weights to remove unsafe behaviors without inference-time overhead.

AI/ML TechCrunch

Trump drops restrictions on Anthropic’s Mythos and Fable models

Anthropic is restoring access to its Fable models following the removal of certain restrictions.

Tech Business/VC TechCrunch

Wayve launches $85M employee tender offer at $8.5B valuation

AI startup Wayve launches an $85M employee tender offer at an $8.5B valuation to attract and retain talent.

AI/ML arXiv cs.AI

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

Researchers propose the Operational Turing Test (OTT) to demonstrate that tabular foundation models cannot reason about running systems without access to governing rules.

AI/ML arXiv cs.AI

Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs

The paper introduces a normal-fan geometry for non-stationary adversarial MDPs to better analyze the cost of non-stationarity in reward sequences.

AI/ML arXiv cs.AI

Unified Complex-valued Neural Network: A Magnitude-Phase Computational Model for Event-Driven Neuromorphic Learning

A new Unified Complex-valued Neuron (UCN) model is introduced to integrate continuous activation and phase-driven event generation for neuromorphic learning.

AI/ML arXiv cs.AI

BTI-Net: Bidirectional Decoder-Level Task Interaction via Uncertainty-Aware Gating for Multi-Task Medical Image Analysis

BTI-Net is introduced for multi-task medical image analysis, using bidirectional decoder-level interaction and uncertainty-aware gating to improve segmentation and classification.

AI/ML arXiv cs.AI

A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment

The Enhanced Neurological Disorder Detection Network (End-Net) is proposed for multi-class MRI classification using multiscale features and WGAN-GP for class imbalance.

AI/ML arXiv cs.AI

LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics

Researchers develop a semantic signaling game framework to analyze strategic communication, deception, and awareness shaping in LLM-mediated interactions.

AI/ML arXiv cs.AI

When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems

An analysis of autonomous vehicle (AV) failure incidents suggests that current 'stopping' fallback behaviors are insufficient and need to be replaced with human-interactive autonomy.

AI/ML arXiv cs.AI

Knowing in Advance When an Evolutionary Outer Loop Will Not Help: A Pre-Registered Cheap-Baseline Screening Rule

A pre-registered screening rule is introduced to determine if an evolutionary outer loop for neural network parameters is worth the computational expense compared to a cheap single-shot alternative.

AI/ML arXiv cs.AI

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

A new pipeline for spatio-temporal grounding in long videos using second-level tracking and RL verification to balance efficiency and localization quality.

AI/ML arXiv cs.AI

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

Research on leveraging synthetic speech for ASR training, demonstrating that room impulse responses (RIRs) can narrow the gap between synthetic and real audio data.

AI/ML arXiv cs.AI

The strength of clinical evidence is recoverable from language model representations but not from their stated grades

Study finding that LLMs possess recoverable evidence-strength signals in their representations but fail to accurately state those grades explicitly.

AI/ML arXiv cs.AI

Metric Aggregation Divergence: A Hidden Validity Threat in Agent-Based Policy Optimization and a Contractual Remedy

Identification of 'Metric Aggregation Divergence' in agent-based policy optimization and the introduction of 'metric contracts' to ensure pipeline consistency.

AI/ML arXiv cs.AI

Flow Matching in Feature Space for Stochastic World Modeling

Introduction of FlowWM, a stochastic world model that performs flow matching within high-dimensional pretrained feature spaces for better perception and robustness.

Cybersecurity arXiv cs.AI

Fairness Attacks on Recommender Systems

A novel reinforcement learning-based attack method designed to exacerbate unfairness and bias within recommender systems.