Deep Learning for Rare Variant Interpretation
DOI:
https://doi.org/10.5281/zenodo.19549576Keywords:
Variant interpretation; Deep learning; Rare variants; Clinical genomics; Pathogenicity prediction; Protein language models; AlphaFold structure; ACMG classification; Variants of uncertain significanceAbstract
Clinical genome sequencing identifies 4-5 million variants per individual, of which 20,000-40,000 are rare (MAF < 0.1%) and potentially disease-causing. Classifying these variants of uncertain significance (VUS) as pathogenic or benign is the critical bottleneck in genomic medicine, with ClinVar containing over 2.4 million VUS compared to only 0.6 million classified variants. Existing computational predictors (CADD, REVEL, AlphaMissense) achieve AUROC 0.88-0.94 for missense variants but lack calibrated probability outputs needed for clinical decision-making and perform poorly on non-coding and structural variants. This study developed and benchmarked six variant interpretation approaches - conservation-based scoring (GERP++), ensemble meta-predictor (REVEL), structure-informed prediction (AlphaMissense), protein language model zero-shot (ESM-1v), deep mutational scanning-calibrated model (ProteinGym), and a proposed Multi-Evidence Variant Interpretation Transformer (MEVIT) integrating protein language model embeddings (ESM-2), 3D protein structure context (AlphaFold2), evolutionary conservation across 241 mammalian genomes (Zoonomia), population frequency data (gnomAD v4), gene-level constraint metrics, and functional assay evidence from multiplexed assays of variant effect (MAVEs) via a calibrated transformer architecture -- across four variant classes: missense (n = 42,684 ClinVar), synonymous (n = 8,426), splice-region (n = 12,848), and non-coding regulatory (n = 6,242). MEVIT achieved the highest AUROC across all variant classes: missense 0.968 (vs AlphaMissense 0.942, REVEL 0.932; p < 0.001), synonymous 0.924, splice-region 0.946, and non-coding 0.884 -- the first model to achieve clinically useful accuracy across all four variant types. Calibration analysis confirmed that MEVIT probability outputs reliably estimate true pathogenicity rates (Brier score 0.062 vs 0.118 for REVEL), enabling direct clinical use as quantitative evidence in ACMG/AMP variant classification. On 1,248 prospectively collected diagnostic VUS, MEVIT reclassified 42.8% to likely pathogenic or likely benign with 96.4% concordance with subsequent expert review. These results establish multi-evidence deep learning as the state-of-the-art for clinical variant interpretation.Downloads
Published
2026-08-15
Issue
Section
Articles
How to Cite
Deep Learning for Rare Variant Interpretation. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 4(2), 46-54. https://doi.org/10.5281/zenodo.19549576

