AI/ML arXiv cs.AI

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Introduces CARE-MH, a unified framework for the reproducible and comparable evaluation of Large Language Models used in mental health support.

AI/ML arXiv cs.AI

Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course

A study on how undergraduate students use GenAI for object-oriented programming, finding that debugging and explanation seeking are more common than code generation.

Other Hacker News

Ancient Rome's version of Google Maps: how long to reach the beach

An exploration of how Ancient Romans tracked travel times and distances to locations like the beach.

AI/ML Hacker News

ChatGPT claims rogue AI attacked more companies

Reports on ChatGPT making claims regarding rogue AI attacks on companies.

Cybersecurity arXiv cs.AI

Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks

Proposes AgentToolMO, a 3GPP NRM information model for managing trust in multi-vendor autonomous networks using AI agents.

AI/ML arXiv cs.AI

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Introduces Penelope, a latent-reasoning framework for Transformers that reduces inference latency by localizing recurrent computation.

AI/ML arXiv cs.AI

dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees

Presents dtControl2+$\\varepsilon$, a tool for creating smaller, more explainable decision trees for Markov decision processes while maintaining optimality.

AI/ML arXiv cs.AI

A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series

Introduces ClinPRISM, a multimodal LLM reasoning framework designed for question answering over irregular clinical time series data.

AI/ML arXiv cs.AI

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

Discusses a solver-guided LLM framework using GRPO to optimize operations research formulation selection for inventory allocation.

AI/ML arXiv cs.AI

CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

Presents CHARM, a multimodal graph foundation model using hierarchical context modeling for zero-shot transfer across graph domains.

AI/ML arXiv cs.AI

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

A behavioral experiment studying how competitive pressure in an AI race can drive actors toward unsafe development practices.

AI/ML arXiv cs.AI

Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?

Introduces Desktop-Delta Bench (DDB), a diagnostic benchmark for evaluating how computer-use agents understand desktop GUI transitions.

AI/ML arXiv cs.AI

Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks

A new framework for LLM-based multi-agent systems uses Bayesian Networks and calibrated log-probabilities to quantify and propagate uncertainty in high-stakes actuarial risk modeling.

Cybersecurity arXiv cs.AI

Distributing Security Controls Through Harness Engineering

The SHarD (Secure Harness Distribution) framework allows security controls like OS sandboxing and tool restriction to be distributed to AI coding agents via a single install command.

AI/ML arXiv cs.AI

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation

Messier is a unified corpus of nearly 1 million records across 30 benchmarks and 700+ agents, providing a standardized infrastructure for evaluating AI agent capabilities.

AI/ML arXiv cs.AI

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

The Interactive Reward Agent (IRA) utilizes a propose-then-verify framework to evaluate GUI agent success by interacting with environment states beyond mere screenshots.

AI/ML arXiv cs.AI

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Introduces CoRT, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual replay to improve reward allocation without auxiliary scorers.

AI/ML arXiv cs.AI

Localized Adaptation Reveals Distinct Learning Signatures in Transformers

Investigates how different adaptation sites in Transformers (early, middle, late layers) affect the learning and generalization of specific objectives like factual association and causal mapping.

AI/ML arXiv cs.AI

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Proposes OmniDelta, a training-free framework for budget allocation in token compression for Omni-modal LLMs, significantly reducing GPU memory and increasing inference speed.

AI/ML arXiv cs.AI

DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space

Presents DecoEvo, a method for co-evolving solver and rubric-generator skills in text space to avoid bottlenecks in open-ended task optimization.