All Articles
16960 articles total
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
A comprehensive survey of Federated Explainable AI (FedXAI), exploring the integration of privacy-preserving learning and model transparency.
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
Presents a risk-based framework for deciding when to invoke expensive LLMs in streaming systems to balance semantic understanding with computational cost.
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
A systematic literature review on autonomous UAV route planning for maximizing environmental monitoring coverage under various constraints.
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
Documents a critical failure in Claude Code where timed-out process outputs are erroneously recorded as confirmed results in session summaries.
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
Introduces the HRO framework for zero-shot object goal navigation, using LLMs to provide hierarchical spatial reasoning for better robot exploration.
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
A study on the identifiability of combined loads from stress-intensity profiles using the SIFBench finite-element dataset.
Dense Arena Interning: The Engine of Compiler Performance
A technical discussion on how dense arena interning can drive compiler performance improvements.
OnePlus officially gives up on the US and Europe
OnePlus has officially announced it will cease product launches in the US and European markets.
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Researchers investigate whether agent-optimization methods can compound performance over time in continual learning scenarios using Terminal-Bench 2.0.
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
A new framework uses AI to accelerate various stages of professional upskilling, from content development to assessment.
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
Earthquaker-AI integrates RAG and robotics to create an educational framework for earthquake preparedness in primary schools.
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
Deep Interaction is a new method for human-AI interaction that allows users to directly edit CoT reasoning steps to steer LLMs efficiently.
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
FixItFlow is an automated system that generates troubleshooting guides for cloud incidents using historical data and LLMs.
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
A new sequential evaluation benchmark, Safe-Psych, assesses how LLMs handle diagnostic uncertainty in psychiatric clinical settings.
Designing Safety-Constrained LLM Systems for Public Health Information Access
A design for a safety-constrained LLM system specifically for public health information access, focusing on maternal and child health.
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Researchers introduce a protocol to measure the utility-risk frontier of dual-use biology assistants by evaluating how user-facing safeguards affect benign utility.
What's the story behind the names of Cloudflare's name servers? (2013)
A discussion thread regarding the historical naming conventions used for Cloudflare's name servers.
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
The introduction of STOCKTAKE, a new benchmark for measuring the 'knowing-doing gap' in LLM agents during multi-week decision tasks.
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
The UESF-Bench benchmark and SeekFollow-VLA framework are introduced for unified evaluation of embodied agents' seeking and following capabilities.
Explaining Reinforcement Learning Agents via Inductive Logic Programming
A research paper proposing the use of Inductive Logic Programming (ILP) to extract symbolic representations from Reinforcement Learning policies for better explainability.