First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
This research introduces NOHARM, a comprehensive benchmark and evaluation framework designed to systematically assess the clinical safety of AI-generated medical consultation advice. The work addresses the critical gap in understanding the potential for harm posed by widely deployed Large Language Models (LLMs) and specialized clinical AI tools in healthcare settings. Developed by a large consortium of researchers, including individuals from multiple institutions and potentially affiliated with technology companies given the subject matter, the findings were published on arXiv.
NOHARM comprises 1,100 primary care-to-specialist consultation cases, meticulously annotated by over 12,000 experts, evaluating 4,249 distinct clinical management options across 10 medical specialties. The core contribution lies in its rigorous methodology for quantifying the frequency and severity of potentially harmful errors. This benchmark is intended for AI developers, medical researchers, and clinicians aiming to ensure the safe integration of AI into clinical workflows.
Two primary technical findings emerge. First, direct application of recommendations from 20 notable LLMs and 4 Retrieval-Augmented Generation (RAG) clinical AI tools revealed a significant potential for severe harm in up to 24.6% of cases, with errors of omission being the predominant cause of severe harm. Second, while clinical AI tools generally outperformed generalist LLMs, and multi-agent AI teaming further improved LLM performance, a randomized study of 101 physicians showed that AI-assisted clinicians often overlooked valuable AI-generated advice, resulting in human-AI team performance that was suboptimal compared to their combined potential.
The implications of this work are profound. It establishes a crucial standard for measuring AI safety in medicine, moving beyond general performance metrics to focus on patient well-being. The findings highlight the need for more robust AI safety evaluation pipelines and suggest that current AI integration strategies in healthcare may not fully capitalize on the complementary strengths of humans and AI. Going forward, NOHARM is poised to drive the development of safer, more reliable medical AI systems and inform best practices for human-AI collaboration in clinical decision-making, potentially influencing regulatory approaches and clinical deployment guidelines. This analysis is based on the abstract provided.