AI/ML arXiv cs.AI

Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

The PCMA framework is proposed for multi-objective multi-agent reinforcement learning to coordinate agent-specific preferences and improve team performance.

AI/ML arXiv cs.AI

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

ClinHallu is a benchmark for diagnosing stage-wise hallucinations in medical multimodal LLMs, providing a structured reasoning trace for better diagnosis.

AI/ML arXiv cs.AI

Learning optimal policies from event logs through reinforcement learning: a comparison of deep and MDP-based approaches

A comparison of deep RL and MDP-based approaches for learning optimal behavioral policies from historical event logs in process mining.

AI/ML arXiv cs.AI

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins

ANSR-DT is a neuro-symbolic framework for industrial digital twins that combines CNN-LSTM, Prolog-based reasoning, and PPO for adaptive and explainable decisions.

AI/ML arXiv cs.AI

LLM-Powered AI Agent Systems and Their Applications in Industry

A comprehensive review of the evolution, architectures, and industrial applications of LLM-powered AI agent systems, including current challenges and solutions.

Software Engineering Hacker News

An O(x)Caml book that runs

A discussion or announcement regarding an OCaml book that is interactive or executable.

AI/ML arXiv cs.AI

When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime

A study on 'silent failures' in production LLM agent runtimes, identifying a 'fail-plausible' pattern where LLMs hallucinate narratives instead of reporting errors.

AI/ML arXiv cs.AI

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

Introduction of AudioDER, a deduplicated reasoning-oriented dataset designed to improve the post-training of Large Audio-Language Models (LALMs).

Open Source arXiv cs.AI

Regulating the Machine Contributor: Governance and Policy Alignment in Open Source

An analysis of AI-driven contributions to open-source projects, proposing a governance framework to align autonomous agents with existing community policies.

AI/ML arXiv cs.AI

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

A comparative study of deep learning architectures for behavioral forecasting in mobile health, finding that PatchTST and TimesFM are particularly effective.

AI/ML arXiv cs.AI

Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts

Proposal of AdaCSM, a Mixture-of-Experts (MoE) framework for adaptive deep clustering in clinical survival prediction to improve patient stratification.

Other arXiv cs.AI

Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms

A computational analysis exploring structural correspondences between Beethoven's Moonlight Sonata and various machine learning architectures.

AI/ML arXiv cs.AI

When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks

Research showing that self-improving VLMs can regress when using task-specific verifiers that are inaccurate for the target task, despite decreasing training loss.

AI/ML arXiv cs.AI

From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing

A method for converting self-supervised speech models into Mixture-of-Experts (MoE) architectures to enhance robustness in anti-spoofing detection.

AI/ML arXiv cs.AI

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

Introduction of LEAF-X, an entropy-guided XAI framework that provides more faithful and sparse explanations for Transformer-based ASR models like Whisper.

Hardware/Chips Hacker News

Banned Book Library in a Wi-Fi Smart Light Bulb

A creative project involving the installation of a banned book library within a Wi-Fi smart light bulb.

Tech Business/VC Hacker News

San Francisco Weighs PG&E Takeover Amid Soaring Utility Costs

San Francisco is considering taking over PG&E due to rising utility costs.

Cybersecurity arXiv cs.AI

Securing the Future of IoMT in the Post-Quantum Era: An Edge-Native Federated Learning Approach

A study proposing a Kubernetes-based framework for securing Internet of Medical Things (IoMT) devices using Post-Quantum Cryptography and Federated Learning.

Cybersecurity arXiv cs.AI

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails

Research demonstrating how LLM-based agent guardrails can be targeted by denial-of-service attacks through crafted reasoning loops.

AI/ML arXiv cs.AI

TRACE: Trajectory-Routed Causal Memory for Delayed-Evidence Visuomotor Imitation

Introduction of TRACE, a memory framework for visuomotor imitation in robotics that uses path signatures to handle delayed evidence.