All Articles
16910 articles total
MASPRM: Multi-Agent System Process Reward Model
MASPRM is a process reward model for multi-agent systems that improves inference-time search (like MCTS) by scoring intermediate agent messages without needing human annotations.
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
An evaluation of long-context LLMs reveals critical failure modes like 'Distributional Collapse' and a 'Safety Tax' where anti-hallucination prompts increase refusal rates.
Mind the Gap: Action Rebinding Attacks against Android GUI Agents
Researchers uncover a 'Action Rebinding' attack that allows malicious Android apps to hijack GUI agents to perform privileged operations by exploiting reasoning latency.
ELF: A Family of Encoder-Free ECG-Language Models
ELF is a family of encoder-free ECG-Language Models that simplifies the architecture of automated ECG interpretation while maintaining state-of-the-art performance.
With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots
The ARGUS pipeline identifies and remedies 'blind spots' in neural retrievers for RAG systems using uncertainty scoring and targeted document augmentation.
Left-right asymmetry in predicting brain activity from LLMs' representations emerges with their formal linguistic competence
Study finds that the ability of LLM representations to predict brain activity in the left hemisphere emerges alongside the model's formal linguistic competence.
The Human-in-the-Loop Is Tired
A discussion on Hacker News about the fatigue associated with the 'human-in-the-loop' requirement in AI systems.
SpaceX scrubs Starship launch after some of its engines didn't start
SpaceX postponed a Starship launch attempt after several engines failed to ignite.
Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation
Introduces DoLQ, a multi-agent LLM framework for discovering ordinary differential equations from observational data through qualitative and quantitative evaluation.
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents
Research on using context-calibrated internal monitoring and entropy to identify and mitigate reward-hacking risks in LLM agents.
Koopman-driven grip force prediction through EMG sensing
A study on using Koopman operator theory and EMG sensing to predict grip force in real-time for robotic rehabilitation.
PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
Presents PersGuard, a framework that uses model backdoors to prevent unauthorized personalization of text-to-image diffusion models.
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
Introduces NSNQuant, a calibration-free vector quantization method for low-bit KV cache compression in LLMs, improving throughput by up to 3x.
Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials
Proposes deep weighted polynomial approximants for functions with asymmetric growth and decay, with applications in option-pricing.
Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery
A ViT-based framework for segmenting disaster-affected areas using Sentinel-2 and Formosat-5 satellite imagery.
Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
Introduces Inverse-LLaVA, a multimodal architecture that projects text embeddings into visual space, reducing the need for explicit alignment pre-training.
Google Kills Custom Search API on Jan 1, 2027
Google is shutting down its Custom Search API on January 1, 2027.
Lingbot-map: A 3D foundation model for reconstructing scenes from streaming data
Lingbot-map is a 3D foundation model designed to reconstruct scenes from streaming data.
Canada says bridge tolls won't be split with U.S. until $6.4B of debt is repaid
Canada has stated that bridge tolls will not be split with the U.S. until $6.4 billion in debt is repaid.
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
Researchers developed a multi-expert routing system for low-resource Manchu OCR that handles various writing styles with high accuracy.