About Us
Research Watch
•स्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छ•नेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकट•Frontline Perspectives on Nursing Leadership in Nepal•Protecting the Smallest Lungs from the Hidden Grip of RSV in Kathmandu•The Heavy Burden of Bullying on Student Wellbeing in Nepal•The Emerging Landscape of Thyroid Health in Central Nepal•How a Recent Western Nepal Study is Redefining Anemia Diagnosis•How H. Pylori is Impacting the Health of Karnali’s High-Altitude Communities•Sweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s Youth•How Missing Checklists and Protocols are Costing Lives in Nepal’s ERs•स्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छ•नेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकट•Frontline Perspectives on Nursing Leadership in Nepal•Protecting the Smallest Lungs from the Hidden Grip of RSV in Kathmandu•The Heavy Burden of Bullying on Student Wellbeing in Nepal•The Emerging Landscape of Thyroid Health in Central Nepal•How a Recent Western Nepal Study is Redefining Anemia Diagnosis•How H. Pylori is Impacting the Health of Karnali’s High-Altitude Communities•Sweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s Youth•How Missing Checklists and Protocols are Costing Lives in Nepal’s ERs

Validating LLM judges for automated oversight of patient communication.

Researchers

Zidu Xu, Johnathan Zeng, Shuang Zhou, Zhihong Zhang, Thibault Heintz, Marion Tonneau, Arsalan Yaghoubi, Bingyang Ye, Vikram Goddla, Lisa Lehmann, Yu-Hui Chen, Elad Sharon, David E Kozono, Anna Revette, Julia Maues, Thelma Brown, Paul Catalano, Raymond H Mak, Dimitry Dligach, Danielle S Bitterman

Abstract

LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints. We developed a stakeholder-informed seven-criterion evaluation rubric spanning safety, reliability, and communication quality. Clinician reference ratings showed strong interrater reliability across all criteria. We validated the rubric on the Informed CONsent Benchmark (ICON-Bench) and benchmarked 19 LLM judges across multiple implementation strategies. LLM judges achieved strong clinician agreement for safety screening and factual verification (Spearman <i>&#x3c1;</i> &gt; 0.80) but weaker agreement for communication quality ( <i>&#x3c1;</i> &lt; 0.60). Safety-specialized guard models underperformed general-purpose models. Patient advocates rated communication quality lower than both clinicians and LLM judges. These findings support LLM judges for scalable patient communication oversight while demonstrating the need for recalibration to patient-centered evaluation standards.
Source: PubMed (PMID: 42780134)View Original on PubMed