AI/ML arXiv cs.AI

Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges

A comprehensive survey of Federated Explainable AI (FedXAI), exploring the integration of privacy-preserving learning and model transparency.

AI/ML arXiv cs.AI

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

Presents a risk-based framework for deciding when to invoke expensive LLMs in streaming systems to balance semantic understanding with computational cost.

AI/ML arXiv cs.AI

Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review

A systematic literature review on autonomous UAV route planning for maximizing environmental monitoring coverage under various constraints.

AI/ML arXiv cs.AI

Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes

Documents a critical failure in Claude Code where timed-out process outputs are erroneously recorded as confirmed results in session summaries.

AI/ML arXiv cs.AI

HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models

Introduces the HRO framework for zero-shot object goal navigation, using LLMs to provide hierarchical spatial reasoning for better robot exploration.

Other arXiv cs.AI

When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data

A study on the identifiability of combined loads from stress-intensity profiles using the SIFBench finite-element dataset.

Software Engineering Hacker News

Dense Arena Interning: The Engine of Compiler Performance

A technical discussion on how dense arena interning can drive compiler performance improvements.

Tech Business/VC The Verge

OnePlus officially gives up on the US and Europe

OnePlus has officially announced it will cease product launches in the US and European markets.

AI/ML arXiv cs.AI

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Researchers investigate whether agent-optimization methods can compound performance over time in continual learning scenarios using Terminal-Bench 2.0.

AI/ML arXiv cs.AI

AI-accelerated End-to-End Framework for Rapid Professional Upskilling

A new framework uses AI to accelerate various stages of professional upskilling, from content development to assessment.

AI/ML arXiv cs.AI

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education

Earthquaker-AI integrates RAG and robotics to create an educational framework for earthquake preparedness in primary schools.

AI/ML arXiv cs.AI

Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models

Deep Interaction is a new method for human-AI interaction that allows users to directly edit CoT reasoning steps to steer LLMs efficiently.

Software Engineering arXiv cs.AI

FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents

FixItFlow is an automated system that generates troubleshooting guides for cloud incidents using historical data and LLMs.

AI/ML arXiv cs.AI

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

A new sequential evaluation benchmark, Safe-Psych, assesses how LLMs handle diagnostic uncertainty in psychiatric clinical settings.

AI/ML arXiv cs.AI

Designing Safety-Constrained LLM Systems for Public Health Information Access

A design for a safety-constrained LLM system specifically for public health information access, focusing on maternal and child health.

AI/ML arXiv cs.AI

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Researchers introduce a protocol to measure the utility-risk frontier of dual-use biology assistants by evaluating how user-facing safeguards affect benign utility.

Other Hacker News

What's the story behind the names of Cloudflare's name servers? (2013)

A discussion thread regarding the historical naming conventions used for Cloudflare's name servers.

AI/ML arXiv cs.AI

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

The introduction of STOCKTAKE, a new benchmark for measuring the 'knowing-doing gap' in LLM agents during multi-week decision tasks.

AI/ML arXiv cs.AI

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

The UESF-Bench benchmark and SeekFollow-VLA framework are introduced for unified evaluation of embodied agents' seeking and following capabilities.

AI/ML arXiv cs.AI

Explaining Reinforcement Learning Agents via Inductive Logic Programming

A research paper proposing the use of Inductive Logic Programming (ILP) to extract symbolic representations from Reinforcement Learning policies for better explainability.