About Us
Research Watch
स्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERsस्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERs

Privacy Leakage in Federated Learning in Radiology Reports: Comparative Evaluation of Tokenizer and Batch-Size Privacy Risks.

Researchers

Santhosh Parampottupadam, Andrés Martínez Mora, Dimitrios Bounias, Sinem Sav, Klaus Maier-Hein, Ralf Floca

Abstract

Federated learning (FL) enables multi-institutional model training on clinical text without sharing raw data; however, gradient inversion methods can reconstruct sensitive information from shared model updates. The extent of such privacy leakage in FL applied to radiology reports, and the role of tokenizer design, remains unclear. This study aimed to quantify gradient-based reconstruction of radiology report text in an FL setting and to compare privacy risk across 3 transformer tokenization strategies in a controlled, tokenizer-aware evaluation. Six FL clients trained a GPT-2-style transformer (sequence length 32) on 2 public clinical-text corpora comprising 368,751 diagnostic reports, 98,206 discharge summaries, and 1500 MIMIC-CXR (Medical Information Mart for Intensive Care Chest X-Ray) radiology reports. Models were trained using 3 tokenizers (GPT-2, RadBERT, and LLaMA-2) with batch sizes of 64, 128, and 256. An active malicious-server threat model was assumed, and analytic gradient inversion was applied to recover text. Reconstruction fidelity was measured over 5 runs using exact sentence accuracy, sentence-level bilingual evaluation understudy (S-BLEU), and recall-oriented understudy for gisting evaluation (ROUGE-L). Exact sentence reconstruction ranged from 27% to 75% across tokenizers, datasets, and batch sizes. At batch size 64 on the discharge dataset, accuracy was 64.7% (GPT-2), 70% (RadBERT), and 67.5% (LLaMA-2), decreasing to 27.3%, 28.5%, and 27.5% at batch size 256. S-BLEU declined with increasing batch size (eg, discharge reports from 0.69 to 0.31). Reconstruction fidelity did not differ significantly across tokenizers (all but one of 27 comparisons nonsignificant; none significant after Holm correction), and approximately 75% of clinical concepts were represented in the reconstructed text (a corpus-level upper bound), regardless of tokenizer. Batch size was the dominant factor governing leakage. Under a worst-case malicious server that tampers with the shared model and observes unprotected per-client gradients (no secure aggregation or differential privacy), substantial portions of radiology-report text can be reconstructed, with up to approximately 75% of reconstructed 32-token sequences and 75% of clinical concepts (not direct patient identifiers) recovered from FL gradients. In a controlled ablation holding model architecture fixed, tokenizer choice, including domain-specific tokenizers, did not significantly affect leakage under the evaluated conditions, whereas batch size was the primary determinant, and no tokenizer significantly reduced the risk. Tokenizer selection should therefore not be treated as a privacy safeguard in this setting. Safeguards such as secure aggregation and differential privacy should therefore be evaluated as candidate protections for FL deployments that must satisfy Health Insurance Portability and Accountability Act (HIPAA) and General Data Protection Regulation (GDPR) requirements in radiology natural language processing (NLP); legal compliance additionally depends on organizational safeguards, risk assessment, and governance beyond the scope of this study.
Source: PubMed (PMID: 42727096)View Original on PubMed