AI/ML arXiv cs.AI

Vero: An Open RL Recipe for General Visual Reasoning

Introduction of Vero, a family of open-weight VLMs and a 600K-sample dataset designed to improve general visual reasoning tasks.

AI/ML arXiv cs.AI

Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent

A system using LLM agents and real-time ontology queries to automate the standardization of legacy biomedical metadata.

Software Engineering arXiv cs.AI

FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning

FM-Agent leverages LLMs to automate compositional reasoning and specification generation for large-scale software systems to find bugs.

AI/ML arXiv cs.AI

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Introduction of DF3DV-1K, a large-scale dataset for benchmarking distractor-free novel view synthesis in radiance fields.

AI/ML arXiv cs.AI

Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

The OCO framework improves out-of-distribution detection by analyzing object co-occurrence patterns in images.

AI/ML arXiv cs.AI

CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation

CADBench is a multimodal benchmark for evaluating AI-assisted generation of editable CAD programs from images and 3D observations.

AI/ML arXiv cs.AI

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

Multi-agent reinforcement learning is used to achieve superhuman, safe, and agile high-speed quadrotor racing.

AI/ML arXiv cs.AI

Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking

Any2Any provides an efficient transfer paradigm for humanoid whole-body tracking models across different robot embodiments.

Other Hacker News

I solved my mystery fatigue with AI

A personal account of using AI to diagnose and solve chronic fatigue.

AI/ML arXiv cs.AI

DeFrame: Debiasing Large Language Models Against Framing Effects

Introduces DeFrame, a method to debias LLMs against framing effects where semantically equivalent prompts produce different fairness outcomes.

AI/ML arXiv cs.AI

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication

Presents LoRDO, a framework for distributed low-rank optimization that reduces communication overhead by 10x during foundation model training.

AI/ML arXiv cs.AI

Flickering Multi-Armed Bandits

Introduces Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability.

AI/ML arXiv cs.AI

Reinforcement-aware Knowledge Distillation for LLM Reasoning

Proposes RL-aware Knowledge Distillation (RLAD) to efficiently distill reasoning capabilities from large LLMs into smaller students using reinforcement learning.

AI/ML arXiv cs.AI

Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking

Presents Latent Gaussian Splatting (LaGS) for 4D panoptic occupancy tracking in dynamic robotic environments.

AI/ML arXiv cs.AI

The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

Reports on the MAMA-MIA Challenge, a benchmark for breast MRI tumor segmentation and treatment response prediction using AI.

AI/ML arXiv cs.AI

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

Introduces ZeSTA, a domain-conditioned training framework for data-efficient personalized speech synthesis using zero-shot TTS augmentation.

AI/ML arXiv cs.AI

Class-Incremental Motion Forecasting

Proposes an end-to-end framework for class-incremental motion forecasting in autonomous vehicles to adapt to new object classes without forgetting.

AI/ML arXiv cs.AI

The Autonomy Tax: Defense Training Breaks LLM Agents

Analyzes the 'Autonomy Tax,' showing how defense training against prompt injections often degrades the competence of multi-step LLM agents.

Other Hacker News

How to feed a dictator

A discussion thread on Hacker News titled 'How to feed a dictator'.

Other Hacker News

A Love Story

A Hacker News post titled 'A Love Story'.