Software Engineering Hacker News

SQLite improving performance with pre-sort

SQLite is introducing pre-sort functionality to enhance query performance.

AI/ML arXiv cs.AI

ELF: Embedded Language Flows

Embedded Language Flows (ELF) is a new class of diffusion models for language modeling that operates in continuous embedding space, outperforming existing discrete DLMs.

AI/ML arXiv cs.AI

EAGT: Echocardiography Augmentation for Generalisability and Transferability

A study on 2D left ventricular segmentation in echocardiography evaluates 29 data augmentation techniques to improve cross-dataset generalisability.

AI/ML arXiv cs.AI

FormalASR: End-to-End Spoken Chinese to Formal Text

FormalASR introduces compact end-to-end models (0.6B and 1.7B) to directly transcribe spoken Chinese into formal written text, eliminating the need for post-processing LLMs.

AI/ML arXiv cs.AI

Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning

Mirage is a representation-level auditing framework that reveals limitations in visual unlearning within Vertical Federated Learning, challenging output-level certification.

AI/ML arXiv cs.AI

Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution

ASASR is a framework for image super-resolution that uses adversarial Sobolev alignment to improve spectral consistency and structural fidelity.

AI/ML arXiv cs.AI

The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection

Student-Centric Answer Sampling (SCAS) proposes selecting teacher-generated answers based on the student's learning cost rather than just the teacher's performance.

AI/ML arXiv cs.AI

Energy-Structured Low-Rank Adaptation for Continual Learning

E2-LoRA is a new adaptation method for continual learning that concentrates knowledge into leading ranks to mitigate task interference and free capacity.

AI/ML arXiv cs.AI

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

PromptPO explores using LLMs as black-box policy optimizers for RL tasks, finding they can match standard RL baselines in several domains but struggle with continuous control.

AI/ML arXiv cs.AI

$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems

tau-Rec is a verifiable benchmark for agentic recommender systems that uses verifiable rewards to replace subjective LLM-as-a-judge evaluations.

Cybersecurity Ars Technica

US offers $10 million for info on group behind Signal and WhatsApp hacking spree

The US government is offering a $10 million reward for information on Russian state-sponsored groups responsible for hacking Signal and WhatsApp.

AI/ML arXiv cs.AI

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

HiMu is a training-free framework for compositional multimodal frame selection in long video QA, improving accuracy over uniform sampling without needing retraining.

AI/ML arXiv cs.AI

IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models

IWP proposes a training-free token pruning framework for Large Vision Language Models based on the dual form perspective of attention to optimize efficiency.

AI/ML arXiv cs.AI

Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior

Researchers developed a privacy-preserving pipeline using OpenPose and Gaze-LLE with an LLM to analyze student attention in classrooms without storing identifiable video.

AI/ML arXiv cs.AI

From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

This study introduces the Spectral Sensitivity Theorem to explain hallucinations in Whisper ASR models, linking them to rank-1 collapse in deep networks.

AI/ML arXiv cs.AI

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

LiveClawBench is a new benchmark for LLM agents focusing on real-world assistant tasks with reproducible full-stack mock applications and stateful execution.

AI/ML arXiv cs.AI

GenMatter: Perceiving Physical Objects with Generative Matter Models

GenMatter is a generative model for motion-based scene interpretation that groups motion cues into particles to perceive physical objects across diverse settings.

AI/ML arXiv cs.AI

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

This research examines the 'alignment target problem,' finding that humans judge AI behavior differently when its human designers' agency is made visible.

AI/ML arXiv cs.AI

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

Uncertainty-Aware Reward Discounting (UARD) is a framework that reduces reward hacking in RLHF systems by modeling both epistemic and aleatoric uncertainty.

AI/ML arXiv cs.AI

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Driver-WM is a driver-centric latent world model that forecasts in-cabin dynamics conditioned on external traffic context for safer driving automation.