Validating LLM judges for automated oversight of patient communication.
Researchers
Zidu Xu, Johnathan Zeng, Shuang Zhou, Zhihong Zhang, Thibault Heintz, Marion Tonneau, Arsalan Yaghoubi, Bingyang Ye, Vikram Goddla, Lisa Lehmann, Yu-Hui Chen, Elad Sharon, David E Kozono, Anna Revette, Julia Maues, Thelma Brown, Paul Catalano, Raymond H Mak, Dimitry Dligach, Danielle S Bitterman
Abstract
LLMs are increasingly used to mediate patient communication, yet scalable evaluation of their safety, accuracy, and communication quality remains an open problem. LLM judges have emerged as automated evaluators, but whether they can holistically replicate human expert judgment is unvalidated. Informed consent for clinical trials presents a demanding case for such validation because it requires conveying complex information to lay audiences under ethical and safety constraints. We developed a stakeholder-informed seven-criterion evaluation rubric spanning safety, reliability, and communication quality. Clinician reference ratings showed strong interrater reliability across all criteria. We validated the rubric on the Informed CONsent Benchmark (ICON-Bench) and benchmarked 19 LLM judges across multiple implementation strategies. LLM judges achieved strong clinician agreement for safety screening and factual verification (Spearman <i>ρ</i> > 0.80) but weaker agreement for communication quality ( <i>ρ</i> < 0.60). Safety-specialized guard models underperformed general-purpose models. Patient advocates rated communication quality lower than both clinicians and LLM judges. These findings support LLM judges for scalable patient communication oversight while demonstrating the need for recalibration to patient-centered evaluation standards.Source: PubMed (PMID: 42780134)View Original on PubMed