All Articles
16186 articles total
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
Evaluation of various Vision-Language Models (VLMs) for detecting geometry clipping in video game QA, finding they currently act better as high-recall filters than standalone detectors.
Face De-Identification: A Domain-Centric Survey from Capture to Processing
A comprehensive survey on face de-identification techniques spanning physical, sensor, and digital domains to protect privacy in AI.
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Development of ClinMM-Bench, a large multi-turn multimodal clinical diagnostic evaluation benchmark to better assess MLLMs in real-world medical scenarios.
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Introduction of Modus, a decoder-only any-to-any multimodal model that treats all modalities symmetrically without task-specific heads.
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs
Introduction of the Kontrast framework for detecting and explaining knowledge inconsistencies across text, tables, and knowledge graphs.
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
PRISM-AH is a framework for recognizing ambivalence and hesitancy in health behavior change through multimodal conflict detection across video streams.
Reinforcement Learning for Code Optimization
A study on using Reinforcement Learning (RL) to optimize code execution time, introducing the DMC-Optim dataset and calibrated sandboxes to overcome measurement noise.
Recursive Filters: SMA, EMA, Low‑Pass, and a Tiny Kalman
A technical discussion or guide covering various recursive filters including Simple Moving Average (SMA), Exponential Moving Average (EMA), Low-Pass filters, and a minimal Kalman filter implementation.
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist
Introduction of OmniQEC, an AI-driven framework that uses LLMs and a slow-fast synergistic workflow to discover practical quantum error-correcting codes.
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
An empirical study on how LLMs attend to bug reports during automated program repair, finding that diffused attention across diagnostics correlates with success.
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
A2TTA is proposed as a test-time adaptation framework to handle evolving topology and temporal shifts in traffic sensor networks for better forecasting.
Stemma: Induced Decision Regions Reveal LLM Provenance
Stemma is a black-box LLM fingerprinting method that uses induced decision regions to determine the provenance and lineage of a suspect model.
A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields
A machine learning workflow combining performance curve forecasting and Bayesian Optimization to optimize gas lift injection rates in unconventional oil fields.
Device Invariance using Domain Adaptation on Acoustic Scene Classification
An evaluation of DANN and CDAN domain adaptation techniques for acoustic scene classification, highlighting how they interact with CNN and transformer features.
Depression Markers in Speech: An Approach based on Tract Variables Dynamics
Research identifying new depression biomarkers in speech by analyzing the dynamical properties of tract variables using entropy and Lyapunov exponents.
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
A study on suppressing specific internal latents (like evaluation-awareness) in LLMs via prompt optimization, arguing that internal readability does not equal behavioral control.
AnnoBench: A Benchmark for Visualization Annotation Generation
Introduction of AnnoBench, a benchmark for evaluating the automated generation of annotations for data visualizations using VLM-as-a-judge.
Show HN: A local merge queue for parallel Claude Code agents
A project introducing a local merge queue designed to coordinate parallel Claude Code agents for improved software development workflows.
The Productivity Mirage
An exploration of the 'productivity mirage', discussing how perceived productivity gains in modern work environments may be illusory.
Microsoft is openly competing with OpenAI, Anthropic more than ever
Microsoft is shifting its AI strategy to more directly compete with partners OpenAI and Anthropic by developing its own homegrown models.