AI/ML arXiv cs.AI

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

ProMSA is a progressive multimodal search agent for Knowledge-Based Visual Question Answering that iteratively chooses between image and text search to improve retrieval accuracy.

AI/ML arXiv cs.AI

SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval

SHARD is a retrieval-preserving embedding transform that protects vector stores from alignment attacks by sharding residuals into keyed cells.

AI/ML arXiv cs.AI

Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation

Parallel Rollout Approximation (PRA) improves pixel-space autoregressive image generation by generating low-dimensional intermediate states, achieving new SOTA results on ImageNet-1K.

AI/ML arXiv cs.AI

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection

A multimodal hybrid NLP pipeline was developed for insurance fraud detection, combining ASR, diarisation, NER, and LLM-RAG to flag inconsistencies in dialogues.

AI/ML arXiv cs.AI

MLVC: Multi-platform Learned Video Codec for Real-World Deployment

MLVC is a hardware-robust neural video codec that ensures cross-platform consistency and real-time performance on NPUs from Apple, Intel, and Qualcomm.

AI/ML arXiv cs.AI

Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution

A study on cross-sensor diffusion super-resolution highlights a significant domain gap between synthetic training data and real satellite imagery.

AI/ML arXiv cs.AI

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

DG^VoiC is a voice clustering framework designed for fraud investigation in call centres to identify repeated speakers across different customer profiles.

AI/ML arXiv cs.AI

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

Research indicates that LLMs may not be better at judging generated answers than creating them, challenging the common assumption in self-evaluation pipelines.

AI/ML arXiv cs.AI

MultiHashFormer: Hash-based Generative Language Models

MultiHashFormer introduces a hash-based generative language model framework that reduces parameter footprint and allows multilingual expansion with constant size.

AI/ML arXiv cs.AI

ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

ToolPrivacyBench provides a benchmark to evaluate purpose-bound privacy in tool-using LLM agents, auditing whether private data is over-disclosed during multi-tool trajectories.

Software Engineering Hacker News

Let's Decode the Mystery Bytes [video]

A video discussing the process of decoding mystery bytes, likely focusing on low-level data analysis or reverse engineering.

Other Hacker News

Pollen (CEO Negus-Fancey, CTO Wright) tried to remove article, and Google helped

A report on Pollen's attempts to remove an article with alleged assistance from Google.

AI/ML arXiv cs.AI

Every Step of the Way: Video-based Parkinsonian Turning Step Counting

Researchers propose a video-based framework for counting turning steps in Parkinson's patients using 3D human mesh recovery and motion encoders.

AI/ML arXiv cs.AI

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

Reflect-R1 introduces an evidence-driven self-correction framework for long video understanding, utilizing a decoupled reinforcement learning algorithm called SD-GRPO.

AI/ML arXiv cs.AI

Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design

Home3D 1.0 is a modular image-to-3D asset generation system designed for high-fidelity interior design and e-commerce assets.

Cybersecurity arXiv cs.AI

Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy

A study demonstrating how agentic AI can be used to autonomously re-identify individuals from mobility microdata by cross-referencing public records.

AI/ML arXiv cs.AI

Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition

A two-stage fine-tuning pipeline using reinforcement learning to generate protein sequences with specific amino-acid compositions for nutritional design.

AI/ML arXiv cs.AI

VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring

VASAE is a new method for naming Sparse Autoencoder dictionary directions by aligning them with the Transformer's token vocabulary.

Software Engineering arXiv cs.AI

Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering

A call for a shift in software engineering from data-driven prediction to causal reasoning to better support complex system development.

AI/ML arXiv cs.AI

From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection

A multi-stage explainable framework that uses LLaMA-3.1-70B to translate black-box speech models into clinical insights for cognitive impairment detection.