AI/ML arXiv cs.AI

MASPRM: Multi-Agent System Process Reward Model

MASPRM is a process reward model for multi-agent systems that improves inference-time search (like MCTS) by scoring intermediate agent messages without needing human annotations.

AI/ML arXiv cs.AI

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

An evaluation of long-context LLMs reveals critical failure modes like 'Distributional Collapse' and a 'Safety Tax' where anti-hallucination prompts increase refusal rates.

Cybersecurity arXiv cs.AI

Mind the Gap: Action Rebinding Attacks against Android GUI Agents

Researchers uncover a 'Action Rebinding' attack that allows malicious Android apps to hijack GUI agents to perform privileged operations by exploiting reasoning latency.

AI/ML arXiv cs.AI

ELF: A Family of Encoder-Free ECG-Language Models

ELF is a family of encoder-free ECG-Language Models that simplifies the architecture of automated ECG interpretation while maintaining state-of-the-art performance.

AI/ML arXiv cs.AI

With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots

The ARGUS pipeline identifies and remedies 'blind spots' in neural retrievers for RAG systems using uncertainty scoring and targeted document augmentation.

AI/ML arXiv cs.AI

Left-right asymmetry in predicting brain activity from LLMs' representations emerges with their formal linguistic competence

Study finds that the ability of LLM representations to predict brain activity in the left hemisphere emerges alongside the model's formal linguistic competence.

AI/ML Hacker News

The Human-in-the-Loop Is Tired

A discussion on Hacker News about the fatigue associated with the 'human-in-the-loop' requirement in AI systems.

Other Ars Technica

SpaceX scrubs Starship launch after some of its engines didn't start

SpaceX postponed a Starship launch attempt after several engines failed to ignite.

AI/ML arXiv cs.AI

Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation

Introduces DoLQ, a multi-agent LLM framework for discovering ordinary differential equations from observational data through qualitative and quantitative evaluation.

AI/ML arXiv cs.AI

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

Research on using context-calibrated internal monitoring and entropy to identify and mitigate reward-hacking risks in LLM agents.

AI/ML arXiv cs.AI

Koopman-driven grip force prediction through EMG sensing

A study on using Koopman operator theory and EMG sensing to predict grip force in real-time for robotic rehabilitation.

AI/ML arXiv cs.AI

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

Presents PersGuard, a framework that uses model backdoors to prevent unauthorized personalization of text-to-image diffusion models.

AI/ML arXiv cs.AI

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

Introduces NSNQuant, a calibration-free vector quantization method for low-bit KV cache compression in LLMs, improving throughput by up to 3x.

AI/ML arXiv cs.AI

Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials

Proposes deep weighted polynomial approximants for functions with asymmetric growth and decay, with applications in option-pricing.

AI/ML arXiv cs.AI

Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery

A ViT-based framework for segmenting disaster-affected areas using Sentinel-2 and Formosat-5 satellite imagery.

AI/ML arXiv cs.AI

Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping

Introduces Inverse-LLaVA, a multimodal architecture that projects text embeddings into visual space, reducing the need for explicit alignment pre-training.

Tech Business/VC Hacker News

Google Kills Custom Search API on Jan 1, 2027

Google is shutting down its Custom Search API on January 1, 2027.

AI/ML Hacker News

Lingbot-map: A 3D foundation model for reconstructing scenes from streaming data

Lingbot-map is a 3D foundation model designed to reconstruct scenes from streaming data.

Other Hacker News

Canada says bridge tolls won't be split with U.S. until $6.4B of debt is repaid

Canada has stated that bridge tolls will not be split with the U.S. until $6.4 billion in debt is repaid.

AI/ML arXiv cs.AI

Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study

Researchers developed a multi-expert routing system for low-resource Manchu OCR that handles various writing styles with high accuracy.