All Articles
16236 articles total
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
A behavioral experiment studying how competitive pressure in an AI race can drive actors toward unsafe development practices.
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Introduces Desktop-Delta Bench (DDB), a diagnostic benchmark for evaluating how computer-use agents understand desktop GUI transitions.
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
A new framework for LLM-based multi-agent systems uses Bayesian Networks and calibrated log-probabilities to quantify and propagate uncertainty in high-stakes actuarial risk modeling.
Distributing Security Controls Through Harness Engineering
The SHarD (Secure Harness Distribution) framework allows security controls like OS sandboxing and tool restriction to be distributed to AI coding agents via a single install command.
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Messier is a unified corpus of nearly 1 million records across 30 benchmarks and 700+ agents, providing a standardized infrastructure for evaluating AI agent capabilities.
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
The Interactive Reward Agent (IRA) utilizes a propose-then-verify framework to evaluate GUI agent success by interacting with environment states beyond mere screenshots.
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Introduces CoRT, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to improve reward allocation without auxiliary scorers.
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Investigates how different adaptation sites in Transformers (early, middle, late layers) affect the learning and generalization of specific objectives like factual association and causal mapping.
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
Proposes OmniDelta, a training-free framework for budget allocation in token compression for Omni-modal LLMs, significantly reducing GPU memory and increasing inference speed.
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Presents DecoEvo, a method for co-evolving solver and rubric-generator skills in text space to avoid bottlenecks in open-ended task optimization.
Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
Introduces Cognivia, an AI-powered CBT copilot for mental healthcare that identifies cognitive distortions and generates rational responses based on authoritative CBT texts.
Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
Study on how LLM-generated recommendation explanations using behavioral framing and social norms can nudge users toward sustainable choices.
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Analyzes how loss invariance in clinical models can lead to concept layers that lack physical scale, emphasizing the need for validation against objective invariance structures.
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
Introduces a joint agent-speculator RL method that teaches agents to predict their next tool call, reducing latency via self-speculation and KV cache reuse.
Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling
Proposes a framework for Distributed Constraint Optimization (DCOP) using online learning and iterative pricing, specifically applied to large-scale satellite scheduling.
HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
Presents HiSkill, a hierarchical skill graph framework that organizes LLM agent trajectories into a graph to bridge high-level skills with executable actions.
SQLite in Production: Optimizing WAL Mode, Concurrency, and VFS Layers
A technical discussion on optimizing SQLite for production environments, focusing on Write-Ahead Logging (WAL) mode, concurrency, and Virtual File System (VFS) layers.
How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model
A study on the trade-offs between LoRA rank, target modules, and quantization (INT8, NF4) for adapting a 60M-parameter T5-small model for text-to-SQL tasks.
Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters
A paper proposing a set-theoretic formalization and algorithm for extracting legal causes from closed chapters of jurisprudence using a method called al-Sabr wa al-Taqsim.
Multi-Sensor Alignment for Weather Simulations
Introduces the Reference Dataset Alignment Method (ReDAM) and Unified-weather-edit to align multi-sensor weather simulations for better 3D object detection in autonomous vehicles.