AI/ML arXiv cs.AI

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

A behavioral experiment studying how competitive pressure in an AI race can drive actors toward unsafe development practices.

AI/ML arXiv cs.AI

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Introduces Desktop-Delta Bench (DDB), a diagnostic benchmark for evaluating how computer-use agents understand desktop GUI transitions.

AI/ML arXiv cs.AI

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

A new framework for LLM-based multi-agent systems uses Bayesian Networks and calibrated log-probabilities to quantify and propagate uncertainty in high-stakes actuarial risk modeling.

Cybersecurity arXiv cs.AI

Distributing Security Controls Through Harness Engineering

The SHarD (Secure Harness Distribution) framework allows security controls like OS sandboxing and tool restriction to be distributed to AI coding agents via a single install command.

AI/ML arXiv cs.AI

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Messier is a unified corpus of nearly 1 million records across 30 benchmarks and 700+ agents, providing a standardized infrastructure for evaluating AI agent capabilities.

AI/ML arXiv cs.AI

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

The Interactive Reward Agent (IRA) utilizes a propose-then-verify framework to evaluate GUI agent success by interacting with environment states beyond mere screenshots.

AI/ML arXiv cs.AI

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Introduces CoRT, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to improve reward allocation without auxiliary scorers.

AI/ML arXiv cs.AI

Localized Adaptation Reveals Distinct Learning Signatures in Transformers

Investigates how different adaptation sites in Transformers (early, middle, late layers) affect the learning and generalization of specific objectives like factual association and causal mapping.

AI/ML arXiv cs.AI

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Proposes OmniDelta, a training-free framework for budget allocation in token compression for Omni-modal LLMs, significantly reducing GPU memory and increasing inference speed.

AI/ML arXiv cs.AI

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Presents DecoEvo, a method for co-evolving solver and rubric-generator skills in text space to avoid bottlenecks in open-ended task optimization.

AI/ML arXiv cs.AI

Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare

Introduces Cognivia, an AI-powered CBT copilot for mental healthcare that identifies cognitive distortions and generates rational responses based on authoritative CBT texts.

AI/ML arXiv cs.AI

Nudging Sustainable Choices through LLM-Generated Recommendation Explanations

Study on how LLM-generated recommendation explanations using behavioral framing and social norms can nudge users toward sustainable choices.

AI/ML arXiv cs.AI

Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

Analyzes how loss invariance in clinical models can lead to concept layers that lack physical scale, emphasizing the need for validation against objective invariance structures.

AI/ML arXiv cs.AI

Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL

Introduces a joint agent-speculator RL method that teaches agents to predict their next tool call, reducing latency via self-speculation and KV cache reuse.

AI/ML arXiv cs.AI

Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

Proposes a framework for Distributed Constraint Optimization (DCOP) using online learning and iterative pricing, specifically applied to large-scale satellite scheduling.

AI/ML arXiv cs.AI

HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs

Presents HiSkill, a hierarchical skill graph framework that organizes LLM agent trajectories into a graph to bridge high-level skills with executable actions.

Software Engineering Hacker News

SQLite in Production: Optimizing WAL Mode, Concurrency, and VFS Layers

A technical discussion on optimizing SQLite for production environments, focusing on Write-Ahead Logging (WAL) mode, concurrency, and Virtual File System (VFS) layers.

AI/ML arXiv cs.AI

How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model

A study on the trade-offs between LoRA rank, target modules, and quantization (INT8, NF4) for adapting a 60M-parameter T5-small model for text-to-SQL tasks.

AI/ML arXiv cs.AI

Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters

A paper proposing a set-theoretic formalization and algorithm for extracting legal causes from closed chapters of jurisprudence using a method called al-Sabr wa al-Taqsim.

AI/ML arXiv cs.AI

Multi-Sensor Alignment for Weather Simulations

Introduces the Reference Dataset Alignment Method (ReDAM) and Unified-weather-edit to align multi-sensor weather simulations for better 3D object detection in autonomous vehicles.