Deep Learning-Based Variant Classification
DOI:
https://doi.org/10.5281/zenodo.19548890Keywords:
variant classification; deep learning; pathogenicity prediction; ClinVar; VCDLI; VUS resolution; protein language model; AlphaMissense; splicing; clinical genomics; ACMG guidelines; conservationAbstract
Clinical genomics identifies thousands of genetic variants per patient, and classifying each variant as pathogenic or benign is the critical bottleneck for diagnostic interpretation, yet deep learning models that predict variant pathogenicity from sequence, conservation, and functional features show variable accuracy across variant types and clinical contexts. We evaluated 214 deep learning variant classification programmes across centres in Spain, Sweden, and Switzerland between 2018 and 2022, spanning five model categories: sequence conservation-based deep classifiers, protein language model effect predictors, structure-informed variant scoring, splicing impact deep models, and multi-evidence ensemble classifiers. A Variant Classification Deep Learning Index (VCDLI) was constructed from five sub-scores - pathogenicity prediction accuracy, variant-of-uncertain-significance resolution rate, clinical-grade calibration, cross-gene generalisation, and interpretability for clinical reporting -- with weights from regression against sustained adoption into clinical genomics pipelines. VCDLI correlated with adoption at r = +0.84 and discriminated adopted from non-adopted models with an AUC of 0.882. Multi-evidence ensemble classifiers scored highest (mean VCDLI 0.824), while structure-informed scoring trailed at 0.598. Only 35.5 percent exceeded the 0.75 threshold. Pathogenicity accuracy carried the largest weight (beta = +0.278), followed by VUS resolution (beta = +0.230).Downloads
Published
2026-08-25
Issue
Section
Articles
How to Cite
Deep Learning-Based Variant Classification. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 2(2), 73-80. https://doi.org/10.5281/zenodo.19548890

