AI/ML arXiv cs.AI

OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation

Introduces DRIVE-CHOREO, a multi-agent world model that uses an LLM to orchestrate multi-view video generation for autonomous driving.

AI/ML arXiv cs.AI

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Presents SR-REAL, a framework using RL to equip spatial VLMs with dual reasoning paths: linguistic deduction and 3D geometric grounding.

AI/ML arXiv cs.AI

Offline Preference-Based Trajectory Evaluation

Proposes a preference-based trajectory evaluation method for agentic systems to reduce ties and improve discriminative power over terminal success metrics.

AI/ML arXiv cs.AI

Reversal Q-Learning

Introduces Reversal Q-Learning (RQL), an off-policy RL algorithm that uses flow-based policies and virtual on-policy trajectories to improve robotic task performance.

Cybersecurity arXiv cs.AI

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

Describes an AI security agent for banking that combines LSTM, statistical monitors, and graph modules to detect fraud and money laundering across transaction and session streams.

AI/ML arXiv cs.AI

Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery

Proposes a geometry-consistent evaluation protocol for foundation model features in multi-view satellite imagery using the RPC framework.

AI/ML arXiv cs.AI

LLM Features Can Hurt GNNs: Concatenation Interference on Homophilous Graph Benchmarks

Finds that concatenating LLM-generated features to GNNs can actually degrade accuracy on certain homophilous benchmarks.

AI/ML arXiv cs.AI

Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics

Presents Visored, a dependent-type-based prover for LLM-generated mathematics that translates natural language proofs into checked Lean files.

Software Engineering arXiv cs.AI

Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations

Analyzes how and why LLMs disagree with human researchers during title-abstract screening for systematic reviews in software engineering.

AI/ML Hacker News

ChatGPT Spontaneously Generates Sexual Violence and Hardcore Snuff Imagery

Reports of ChatGPT spontaneously generating inappropriate and violent imagery.

Other Hacker News

Pink Cosmo berries a hit in their trial season (2023)

A trial of Pink Cosmo berries in 2023.

AI/ML arXiv cs.AI

AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows

Introduction of AIPatient Arena, an EHR-grounded framework for evaluating LLMs in clinical consultation workflows.

AI/ML arXiv cs.AI

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing

Presentation of STATEWITNESS, an activation explainer designed to audit and detect deceptive behavior in reasoning LLMs.

AI/ML arXiv cs.AI

Online LLM Selection via Constrained Bandits with Time-Varying Demand

A new online learning algorithm for selecting LLMs in edge-cloud systems based on constrained bandits with time-varying demand.

AI/ML arXiv cs.AI

MagicSim: A Unified Infrastructure for Executable Embodied Interaction

Introduction of MagicSim, a unified infrastructure for executable embodied interaction that bridges robot learning and simulation.

AI/ML arXiv cs.AI

Geometry-Aware Post-Hoc Uncertainty Quantification in Operator Learning

Proposed REEF-GP, a post-hoc uncertainty quantification framework for neural operators in PDE solving.

AI/ML arXiv cs.AI

Unlocking LLM Code Correction with Iterative Feedback Loops

Investigation into LLM code correction using iterative feedback loops with compiler and testcase feedback.

AI/ML arXiv cs.AI

FoundCause: Causal Discovery with Latent Confounders from Observational Data

Introduction of FoundCause, an amortized causal discovery model that handles latent confounders from observational data.

AI/ML arXiv cs.AI

Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery

Analysis of how routing accuracy degrades in large-scale enterprise agent catalogs and the effectiveness of embedding-based shortlisting.

Other Hacker News

Taxonomy of the Occlupanida (parasitoids on bread bag tags)

A discussion regarding a taxonomy of parasitoids that live on bread bag tags.