About Us
Research Watch
स्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERsस्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERs

ChatGPT as a source of surgical information: Evaluation of responses to patient questions on hallux rigidus fusion.

Researchers

Kamil Balaban, Mehmet Batu Ertan, Mahmut Kalem

Abstract

Patients are increasingly turning to online resources and artificial intelligence (AI)-based tools to obtain information about orthopedic conditions and surgical options. Large language models, such as ChatGPT, are becoming prominent in patient education; however, their reliability and readability remain uncertain. This study evaluated the quality and readability of responses generated by ChatGPT-4o and ChatGPT-5 to frequently asked patient questions regarding hallux rigidus fusion surgery. Twenty commonly asked patient questions were compiled and presented to ChatGPT-4o and ChatGPT-5. Readability was assessed using the Flesch-Kincaid Grade Level, Gunning Fog, Coleman-Liau, and Simple Measure of Gobbledygook indices. Quality was evaluated with the DISCERN tool, response accuracy scores, and Journal of the American Medical Association (JAMA) criteria. Interrater agreement was measured using the Intraclass Correlation Coefficient (ICC). ChatGPT-4o generated longer responses (802 vs. 242 words; p<0.001) with slightly higher readability grade levels (10.81 vs. 10.37; p=0.031). Accuracy (2.00 vs. 1.85; p=0.323) and DISCERN scores (49.35 vs. 48.93; p=0.747) showed no significant differences. All responses received a JAMA score of 0 due to the absence of citations, authorship, or transparency indicators. Interrater reliability indicated moderate to good agreement (ICC: 0.68-0.80). ChatGPT-4o and ChatGPT-5 provide generally satisfactory yet non-comprehensive, limited-quality information at a level above tenth-grade regarding hallux rigidus fusion surgery. Although linguistically coherent, responses lack evidence-based detail and individualized guidance. These models may supplement, but cannot replace, expert orthopedic counseling. Ensuring physician oversight and integrating validated, updated clinical content remain essential for safe implementation of AI-generated patient information.
Source: PubMed (PMID: 42633015)View Original on PubMed