All Articles
17790 articles total
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
ProMSA is a progressive multimodal search agent for Knowledge-Based Visual Question Answering that iteratively chooses between image and text search to improve retrieval accuracy.
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
SHARD is a retrieval-preserving embedding transform that protects vector stores from alignment attacks by sharding residuals into keyed cells.
Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
Parallel Rollout Approximation (PRA) improves pixel-space autoregressive image generation by generating low-dimensional intermediate states, achieving new SOTA results on ImageNet-1K.
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
A multimodal hybrid NLP pipeline was developed for insurance fraud detection, combining ASR, diarisation, NER, and LLM-RAG to flag inconsistencies in dialogues.
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
MLVC is a hardware-robust neural video codec that ensures cross-platform consistency and real-time performance on NPUs from Apple, Intel, and Qualcomm.
Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution
A study on cross-sensor diffusion super-resolution highlights a significant domain gap between synthetic training data and real satellite imagery.
DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions
DG^VoiC is a voice clustering framework designed for fraud investigation in call centres to identify repeated speakers across different customer profiles.
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
Research indicates that LLMs may not be better at judging generated answers than creating them, challenging the common assumption in self-evaluation pipelines.
MultiHashFormer: Hash-based Generative Language Models
MultiHashFormer introduces a hash-based generative language model framework that reduces parameter footprint and allows multilingual expansion with constant size.
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
ToolPrivacyBench provides a benchmark to evaluate purpose-bound privacy in tool-using LLM agents, auditing whether private data is over-disclosed during multi-tool trajectories.
Let's Decode the Mystery Bytes [video]
A video discussing the process of decoding mystery bytes, likely focusing on low-level data analysis or reverse engineering.
Pollen (CEO Negus-Fancey, CTO Wright) tried to remove article, and Google helped
A report on Pollen's attempts to remove an article with alleged assistance from Google.
Every Step of the Way: Video-based Parkinsonian Turning Step Counting
Researchers propose a video-based framework for counting turning steps in Parkinson's patients using 3D human mesh recovery and motion encoders.
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Reflect-R1 introduces an evidence-driven self-correction framework for long video understanding, utilizing a decoupled reinforcement learning algorithm called SD-GRPO.
Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design
Home3D 1.0 is a modular image-to-3D asset generation system designed for high-fidelity interior design and e-commerce assets.
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
A study demonstrating how agentic AI can be used to autonomously re-identify individuals from mobility microdata by cross-referencing public records.
Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition
A two-stage fine-tuning pipeline using reinforcement learning to generate protein sequences with specific amino-acid compositions for nutritional design.
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
VASAE is a new method for naming Sparse Autoencoder dictionary directions by aligning them with the Transformer's token vocabulary.
Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering
A call for a shift in software engineering from data-driven prediction to causal reasoning to better support complex system development.
From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection
A multi-stage explainable framework that uses LLaMA-3.1-70B to translate black-box speech models into clinical insights for cognitive impairment detection.