All Articles
18346 articles total
OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
Introduces DRIVE-CHOREO, a multi-agent world model that uses an LLM to orchestrate multi-view video generation for autonomous driving.
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
Presents SR-REAL, a framework using RL to equip spatial VLMs with dual reasoning paths: linguistic deduction and 3D geometric grounding.
Offline Preference-Based Trajectory Evaluation
Proposes a preference-based trajectory evaluation method for agentic systems to reduce ties and improve discriminative power over terminal success metrics.
Reversal Q-Learning
Introduces Reversal Q-Learning (RQL), an off-policy RL algorithm that uses flow-based policies and virtual on-policy trajectories to improve robotic task performance.
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts
Describes an AI security agent for banking that combines LSTM, statistical monitors, and graph modules to detect fraud and money laundering across transaction and session streams.
Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery
Proposes a geometry-consistent evaluation protocol for foundation model features in multi-view satellite imagery using the RPC framework.
LLM Features Can Hurt GNNs: Concatenation Interference on Homophilous Graph Benchmarks
Finds that concatenating LLM-generated features to GNNs can actually degrade accuracy on certain homophilous benchmarks.
Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics
Presents Visored, a dependent-type-based prover for LLM-generated mathematics that translates natural language proofs into checked Lean files.
Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations
Analyzes how and why LLMs disagree with human researchers during title-abstract screening for systematic reviews in software engineering.
ChatGPT Spontaneously Generates Sexual Violence and Hardcore Snuff Imagery
Reports of ChatGPT spontaneously generating inappropriate and violent imagery.
Pink Cosmo berries a hit in their trial season (2023)
A trial of Pink Cosmo berries in 2023.
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
Introduction of AIPatient Arena, an EHR-grounded framework for evaluating LLMs in clinical consultation workflows.
Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing
Presentation of STATEWITNESS, an activation explainer designed to audit and detect deceptive behavior in reasoning LLMs.
Online LLM Selection via Constrained Bandits with Time-Varying Demand
A new online learning algorithm for selecting LLMs in edge-cloud systems based on constrained bandits with time-varying demand.
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
Introduction of MagicSim, a unified infrastructure for executable embodied interaction that bridges robot learning and simulation.
Geometry-Aware Post-Hoc Uncertainty Quantification in Operator Learning
Proposed REEF-GP, a post-hoc uncertainty quantification framework for neural operators in PDE solving.
Unlocking LLM Code Correction with Iterative Feedback Loops
Investigation into LLM code correction using iterative feedback loops with compiler and testcase feedback.
FoundCause: Causal Discovery with Latent Confounders from Observational Data
Introduction of FoundCause, an amortized causal discovery model that handles latent confounders from observational data.
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery
Analysis of how routing accuracy degrades in large-scale enterprise agent catalogs and the effectiveness of embedding-based shortlisting.
Taxonomy of the Occlupanida (parasitoids on bread bag tags)
A discussion regarding a taxonomy of parasitoids that live on bread bag tags.