AI/ML arXiv cs.AI

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

Introduction of EgoSafetyBench, a diagnostic egocentric video benchmark designed to evaluate Vision-Language Models (VLMs) as runtime safety guards for embodied AI.

AI/ML arXiv cs.AI

A Category Theory Account of AI Identity

A category-theoretic formalization of AI identity, providing a structured hierarchy of criteria to determine when an AI system remains the same over time and deployments.

AI/ML arXiv cs.AI

Adaptive Perturbation Selection for Contrastive Audio Decoding

Explores adaptive perturbation selection for contrastive audio decoding in Large Audio-Language Models (LALMs) to reduce hallucinations.

AI/ML arXiv cs.AI

Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring

Presents UPADNet, an unrolled network that uses amplitude and phase decomposition to improve image deblurring, especially in high-noise environments.

AI/ML arXiv cs.AI

Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification

Introduces a particle-based diversification framework for Test-Time Adaptation (TTA) to mitigate underspecification and improve model robustness under distribution shifts.

AI/ML arXiv cs.AI

Validating Causal Abstraction Metrics on Simulated Complex Systems

Introduces the Causal Abstraction Error (CAE), a continuous validity metric for discovering and validating high-level causal explanations of complex systems.

AI/ML arXiv cs.AI

ASPIRE: Agentic /Skills Discovery for Robotics

Presents ASPIRE, a continual learning system for robotics that autonomously writes and refines control programs using a code-as-policy paradigm.

AI/ML arXiv cs.AI

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Introduces SEFORA, a public corpus of student essays with instructor feedback, and UniMatch, an evaluation framework for LLM-generated writing feedback.

AI/ML arXiv cs.AI

Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning

Proposes entropy-regularized probabilistic gates to maintain uncertainty in sparse federated optimization, improving generalization in data-scarce federated learning.

Other Hacker News

We Don't Have to Be This Bad at Improving Society

A discussion about the societal shortcomings in improving human conditions and systemic failures.

Other Hacker News

The Fall of the Theorem Economy

An exploration of the 'Theorem Economy' and the shifting nature of mathematical proof and value.

Tech Business/VC Hacker News

Google loses fight over record $4.7B EU antitrust fine

Google loses a legal battle regarding a massive EU antitrust fine of $4.7 billion.

Hardware/Chips Hacker News

My Favorite Keyboards

A user shares their personal preferences and recommendations for various computer keyboards.

AI/ML arXiv cs.AI

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

A research paper proving that GRPO, Dr. GRPO, and DAPO are variations of a single mathematical identity based on group standard deviation.

AI/ML arXiv cs.AI

EVOTS: Evolutionary Transformer Search for Time Series Forecasting

Introduction of EVOTS, an evolutionary neural architecture search framework for optimizing Transformer models for time-series forecasting.

AI/ML arXiv cs.AI

Scaling Up Thermodynamic AI Models

Research on scaling thermodynamic AI models using backpropagation-based algorithms for training on Ising machine hardware.

AI/ML arXiv cs.AI

Play Like Champions: Counterfactual Feedback Generation in Latent Space

A framework for generating counterfactual feedback in latent space to help human players improve at real-time strategy games like StarCraft II.

AI/ML arXiv cs.AI

HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems

HydraCollab is an adaptive collaborative-perception framework for distributed autonomous systems that optimizes bandwidth and accuracy.

AI/ML arXiv cs.AI

SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing

Introduction of SLIM-RL, a risk-budgeted random-masking RL method that improves training efficiency for diffusion LLMs.

AI/ML arXiv cs.AI

Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention

Researchers propose Multi-Scale Layer Attention (MSLA) to improve the recognition of ancient Chinese Oracle Bone Inscriptions by modeling multi-scale and cross-layer feature interactions.