AI/ML arXiv cs.AI

FlexMS: A Unified Public Benchmark for Molecule Tandem Mass Spectrum Prediction

Introduces FlexMS, a standardized public benchmark for predicting molecule tandem mass spectra to ensure fair architectural comparisons in metabolomics.

AI/ML arXiv cs.AI

Generative AI for Managerial Decision-Making under Ambiguity and Sycophancy

A study on how Generative AI handles ambiguity and sycophancy in managerial decision-making, highlighting the need for human oversight.

AI/ML arXiv cs.AI

An Analysis of the Coordination Gap between Joint and Modular Learning for Job Shop Scheduling with Transportation Resources

Analyzes the performance gap between joint and modular reinforcement learning for job shop scheduling with transportation resources in manufacturing.

AI/ML arXiv cs.AI

PRISM: Perception Reasoning Interleaved for Sequential Decision Making

Introduces PRISM, a framework that couples perception and decision-making for embodied agents using a dynamic question-answer pipeline between VLMs and LLMs.

AI/ML arXiv cs.AI

AdaTKG: Adaptive Memory for Temporal Knowledge Graph Reasoning

Presents AdaTKG, a method for temporal knowledge graph reasoning that uses adaptive per-entity memory updated via exponential moving averages.

Tech Business/VC TechCrunch

Sundar Pichai faces boos, walkout at Stanford graduation ceremony over Google’s Israel, ICE ties

Sundar Pichai faced protests and boos during a Stanford graduation ceremony over Google's defense contracts and ties to Israel and ICE.

Other Ars Technica

Key mission for Europe's commercial space enterprise scrubbed again

Isar Aerospace has faced multiple delays in its commercial space missions due to a lack of flight experience.

Cybersecurity arXiv cs.AI

Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications

Researchers demonstrate that audible acoustic frequencies can cause camera vibrations that lead to misclassifications in AI computer vision models like YOLO11.

AI/ML arXiv cs.AI

CottonLeafVision: An Explainable and Robust Deep Learning Framework for Cotton Leaf Disease Classification

CottonLeafVision is a deep learning framework using DenseNet201 and Grad-CAM to accurately classify and detect cotton leaf diseases.

AI/ML arXiv cs.AI

Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit

This research explores the theoretical necessity of generating 'trivial' mathematical statements when using AI and proof assistants to discover valuable mathematics.

AI/ML arXiv cs.AI

Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning

The PCMA framework is proposed for multi-objective multi-agent reinforcement learning to coordinate agent-specific preferences and improve team performance.

AI/ML arXiv cs.AI

ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning

ClinHallu is a benchmark for diagnosing stage-wise hallucinations in medical multimodal LLMs, providing a structured reasoning trace for better diagnosis.

AI/ML arXiv cs.AI

Learning optimal policies from event logs through reinforcement learning: a comparison of deep and MDP-based approaches

A comparison of deep RL and MDP-based approaches for learning optimal behavioral policies from historical event logs in process mining.

AI/ML arXiv cs.AI

ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins

ANSR-DT is a neuro-symbolic framework for industrial digital twins that combines CNN-LSTM, Prolog-based reasoning, and PPO for adaptive and explainable decisions.

AI/ML arXiv cs.AI

LLM-Powered AI Agent Systems and Their Applications in Industry

A comprehensive review of the evolution, architectures, and industrial applications of LLM-powered AI agent systems, including current challenges and solutions.

Software Engineering Hacker News

An O(x)Caml book that runs

A discussion or announcement regarding an OCaml book that is interactive or executable.

AI/ML arXiv cs.AI

When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime

A study on 'silent failures' in production LLM agent runtimes, identifying a 'fail-plausible' pattern where LLMs hallucinate narratives instead of reporting errors.

AI/ML arXiv cs.AI

AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

Introduction of AudioDER, a deduplicated reasoning-oriented dataset designed to improve the post-training of Large Audio-Language Models (LALMs).

Open Source arXiv cs.AI

Regulating the Machine Contributor: Governance and Policy Alignment in Open Source

An analysis of AI-driven contributions to open-source projects, proposing a governance framework to align autonomous agents with existing community policies.

AI/ML arXiv cs.AI

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

A comparative study of deep learning architectures for behavioral forecasting in mobile health, finding that PatchTST and TimesFM are particularly effective.