AI/ML arXiv cs.AI

Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA

Evaluation of various Vision-Language Models (VLMs) for detecting geometry clipping in video game QA, finding they currently act better as high-recall filters than standalone detectors.

AI/ML arXiv cs.AI

Face De-Identification: A Domain-Centric Survey from Capture to Processing

A comprehensive survey on face de-identification techniques spanning physical, sensor, and digital domains to protect privacy in AI.

AI/ML arXiv cs.AI

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Development of ClinMM-Bench, a large multi-turn multimodal clinical diagnostic evaluation benchmark to better assess MLLMs in real-world medical scenarios.

AI/ML arXiv cs.AI

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

Introduction of Modus, a decoder-only any-to-any multimodal model that treats all modalities symmetrically without task-specific heads.

AI/ML arXiv cs.AI

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

Introduction of the Kontrast framework for detecting and explaining knowledge inconsistencies across text, tables, and knowledge graphs.

AI/ML arXiv cs.AI

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

PRISM-AH is a framework for recognizing ambivalence and hesitancy in health behavior change through multimodal conflict detection across video streams.

AI/ML arXiv cs.AI

Reinforcement Learning for Code Optimization

A study on using Reinforcement Learning (RL) to optimize code execution time, introducing the DMC-Optim dataset and calibrated sandboxes to overcome measurement noise.

Software Engineering Hacker News

Recursive Filters: SMA, EMA, Low‑Pass, and a Tiny Kalman

A technical discussion or guide covering various recursive filters including Simple Moving Average (SMA), Exponential Moving Average (EMA), Low-Pass filters, and a minimal Kalman filter implementation.

AI/ML arXiv cs.AI

OmniQEC: discovering practical quantum error-correcting codes by an AI scientist

Introduction of OmniQEC, an AI-driven framework that uses LLMs and a slow-fast synergistic workflow to discover practical quantum error-correcting codes.

AI/ML arXiv cs.AI

How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair

An empirical study on how LLMs attend to bug reports during automated program repair, finding that diffused attention across diagnostics correlates with success.

AI/ML arXiv cs.AI

A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks

A2TTA is proposed as a test-time adaptation framework to handle evolving topology and temporal shifts in traffic sensor networks for better forecasting.

AI/ML arXiv cs.AI

Stemma: Induced Decision Regions Reveal LLM Provenance

Stemma is a black-box LLM fingerprinting method that uses induced decision regions to determine the provenance and lineage of a suspect model.

AI/ML arXiv cs.AI

A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields

A machine learning workflow combining performance curve forecasting and Bayesian Optimization to optimize gas lift injection rates in unconventional oil fields.

AI/ML arXiv cs.AI

Device Invariance using Domain Adaptation on Acoustic Scene Classification

An evaluation of DANN and CDAN domain adaptation techniques for acoustic scene classification, highlighting how they interact with CNN and transformer features.

AI/ML arXiv cs.AI

Depression Markers in Speech: An Approach based on Tract Variables Dynamics

Research identifying new depression biomarkers in speech by analyzing the dynamical properties of tract variables using entropy and Lyapunov exponents.

AI/ML arXiv cs.AI

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

A study on suppressing specific internal latents (like evaluation-awareness) in LLMs via prompt optimization, arguing that internal readability does not equal behavioral control.

AI/ML arXiv cs.AI

AnnoBench: A Benchmark for Visualization Annotation Generation

Introduction of AnnoBench, a benchmark for evaluating the automated generation of annotations for data visualizations using VLM-as-a-judge.

AI/ML Hacker News

Show HN: A local merge queue for parallel Claude Code agents

A project introducing a local merge queue designed to coordinate parallel Claude Code agents for improved software development workflows.

Other Hacker News

The Productivity Mirage

An exploration of the 'productivity mirage', discussing how perceived productivity gains in modern work environments may be illusory.

Tech Business/VC TechCrunch

Microsoft is openly competing with OpenAI, Anthropic more than ever

Microsoft is shifting its AI strategy to more directly compete with partners OpenAI and Anthropic by developing its own homegrown models.