All Articles
16226 articles total
CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
Introduces CARE-MH, a unified framework for the reproducible and comparable evaluation of Large Language Models used in mental health support.
Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course
A study on how undergraduate students use GenAI for object-oriented programming, finding that debugging and explanation seeking are more common than code generation.
Ancient Rome's version of Google Maps: how long to reach the beach
An exploration of how Ancient Romans tracked travel times and distances to locations like the beach.
ChatGPT claims rogue AI attacked more companies
Reports on ChatGPT making claims regarding rogue AI attacks on companies.
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
Proposes AgentToolMO, a 3GPP NRM information model for managing trust in multi-vendor autonomous networks using AI agents.
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
Introduces Penelope, a latent-reasoning framework for Transformers that reduces inference latency by localizing recurrent computation.
dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees
Presents dtControl2+$\\varepsilon$, a tool for creating smaller, more explainable decision trees for Markov decision processes while maintaining optimality.
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
Introduces ClinPRISM, a multimodal LLM reasoning framework designed for question answering over irregular clinical time series data.
Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation
Discusses a solver-guided LLM framework using GRPO to optimize operations research formulation selection for inventory allocation.
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer
Presents CHARM, a multimodal graph foundation model using hierarchical context modeling for zero-shot transfer across graph domains.
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
A behavioral experiment studying how competitive pressure in an AI race can drive actors toward unsafe development practices.
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Introduces Desktop-Delta Bench (DDB), a diagnostic benchmark for evaluating how computer-use agents understand desktop GUI transitions.
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
A new framework for LLM-based multi-agent systems uses Bayesian Networks and calibrated log-probabilities to quantify and propagate uncertainty in high-stakes actuarial risk modeling.
Distributing Security Controls Through Harness Engineering
The SHarD (Secure Harness Distribution) framework allows security controls like OS sandboxing and tool restriction to be distributed to AI coding agents via a single install command.
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Messier is a unified corpus of nearly 1 million records across 30 benchmarks and 700+ agents, providing a standardized infrastructure for evaluating AI agent capabilities.
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
The Interactive Reward Agent (IRA) utilizes a propose-then-verify framework to evaluate GUI agent success by interacting with environment states beyond mere screenshots.
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Introduces CoRT, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to improve reward allocation without auxiliary scorers.
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Investigates how different adaptation sites in Transformers (early, middle, late layers) affect the learning and generalization of specific objectives like factual association and causal mapping.
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
Proposes OmniDelta, a training-free framework for budget allocation in token compression for Omni-modal LLMs, significantly reducing GPU memory and increasing inference speed.
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Presents DecoEvo, a method for co-evolving solver and rubric-generator skills in text space to avoid bottlenecks in open-ended task optimization.