A hybrid artificial intelligence system ranked previously identified disease-causing variants among its top three candidates in more than 90 percent of retrospectively assessed rare disease cases, according to a study published in npj Genomic Medicine.
The system, called aiDIVA, was developed to reduce the number of variants requiring manual review following exome or genome sequencing. It combines evidence-based scoring, machine learning, phenotype matching, inheritance information, and large language models to generate a ranked list for clinical assessment.
Researchers trained the pathogenicity model using more than 155,000 variants from ClinVar, including approximately 90,000 pathogenic and 65,000 benign variants. The system evaluated genomic features such as population frequency, evolutionary conservation, predicted protein effects, and compatibility with the patient’s reported phenotype.
Separate models assessed dominant and recessive inheritance. The 10 highest-ranked genes from each model were then submitted to GPT-4o, which considered phenotype terms, age, sex, zygosity, and predicted variant effects. A final ensemble model combined the results.
The researchers benchmarked aiDIVA using 3,041 rare disease cases previously solved by genetics specialists at the University Hospital of Tübingen. The cohort included 1,975 exome and 1,066 genome sequencing data sets. Cases involving copy number or other structural variants were excluded.
Among cases containing variants already represented in ClinVar or the Human Gene Mutation Database, aiDIVA placed the causal variant within the top three in 97 percent of cases and at rank one in 88 percent. It also outperformed the tested versions of Exomiser, Lirical, Xrare, and AI-MARRVEL in the study’s comparisons.
To reduce the effect of overlap between model training and evaluation data, the team assessed an independent group of 1,014 cases diagnosed after GPT-4o was trained. The causal variants were not present in ClinVar or the Human Gene Mutation Database at the training cutoff. In this group, aiDIVA placed 93 percent of causal variants within the top three and 96 percent within the top 10.
The researchers also reanalyzed 4,877 cases that had remained unsolved since 2020 or earlier. Updated annotations and aiDIVA scores were used to select 500 cases for expert review. Specialists subsequently classified 45 as newly solved. A version of aiDIVA that excluded direct evidence from clinical variant databases placed 34 of these variants within the top three.
The system may support variant triage and periodic reanalysis as databases and disease associations change. It remains a decision-support tool rather than an autonomous diagnostic system. Final classification still requires expert review, application of American College of Medical Genetics and Genomics criteria, and assessment of segregation, phenotype, and other clinical evidence.
The researchers noted that language-model responses were inconsistent in a small proportion of repeated tests and could cite nonexistent or irrelevant literature. Sending phenotype data to cloud-based models also raises privacy considerations. In addition, aiDIVA does not currently support structural variants, and its performance for in-frame insertions and deletions was lower than for single-nucleotide variants.
