About Us
Research Watch
स्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERsस्क्रिनमा देखिने चुरोट: सुर्तीजन्य हानि न्यूनीकरण नीतिमा दक्षिण एसियाले अझै के छुटाइरहेको छनेपालमा पिसाब नलीको संक्रमण र एन्टिबायोटिक प्रतिरोधको बढ्दो संकटFrontline Perspectives on Nursing Leadership in NepalProtecting the Smallest Lungs from the Hidden Grip of RSV in KathmanduThe Heavy Burden of Bullying on Student Wellbeing in NepalThe Emerging Landscape of Thyroid Health in Central NepalHow a Recent Western Nepal Study is Redefining Anemia DiagnosisHow H. Pylori is Impacting the Health of Karnali’s High-Altitude CommunitiesSweet Poison, Bitter Reality: The Unseen Diabetes Epidemic Among Nepal’s YouthHow Missing Checklists and Protocols are Costing Lives in Nepal’s ERs

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Researchers

Peng Wang, Kai Wang

Abstract

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.
Source: PubMed (PMID: 42754603)View Original on PubMed