Machine Learning-Based CRISPR Guide Design

Authors

  • Erik Hansen Author

DOI:

https://doi.org/10.5281/zenodo.19549486

Keywords:

CRISPR guide design; Machine learning; gRNA activity prediction; Chromatin accessibility; Deep learning; Gene editing efficiency; DNA language models; Epigenomic context; Off-target prediction

Abstract

CRISPR-Cas9 gene editing efficiency varies over 100-fold depending on the 20-nucleotide guide RNA (gRNA) sequence, yet the molecular determinants of on-target activity remain incompletely understood, making computational guide design essential for experimental success. Existing prediction tools (Rule Set 2, DeepCRISPR, CHOPCHOP) achieve Spearman correlations of 0.55-0.65 with experimentally measured editing efficiency, leaving substantial room for improvement. This study developed and benchmarked six gRNA activity prediction approaches -- sequence-based logistic regression (Doench Rule Set 2), convolutional neural network (DeepCRISPR), gradient-boosted trees (CRISPRscan), recurrent neural network (LSTM), pre-trained DNA language model fine-tuning (DNABERT-2), and a proposed multi-scale genomic context model (MGC-CRISPR) integrating guide sequence features, local chromatin accessibility (ATAC-seq), 3D genome organisation (Hi-C), and cell-type-specific gene expression via a hierarchical attention architecture -- across three large-scale experimental datasets: Avana genome-wide library (73,372 guides, 18,436 genes), GeCKO v2 library (122,411 guides), and a custom validation library (4,800 guides across 12 cell lines). MGC-CRISPR achieved the highest prediction accuracy: Spearman rho = 0.782 +- 0.012 on held-out Avana data (vs 0.642 for Rule Set 2, 0.668 for DeepCRISPR, 0.724 for DNABERT-2; p < 0.001 for all), with the greatest improvement in genes with heterogeneous chromatin states where guide accessibility varies dramatically across cell types. Cross-cell-line validation confirmed that MGC-CRISPR's chromatin-aware predictions generalised across 12 cell lines (mean rho = 0.748 vs 0.612 for sequence-only models; p < 0.001). Feature attribution analysis revealed that chromatin accessibility contributed 28.4% of prediction variance -- comparable to the 34.2% from guide sequence features - establishing chromatin context as an essential and previously underutilised predictor of CRISPR editing efficiency. These results demonstrate that incorporating epigenomic context substantially improves gRNA design and provide a cell-type-aware prediction framework for precision gene editing.

Downloads

Published

2026-08-15

How to Cite

Machine Learning-Based CRISPR Guide Design. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 3(4), 74-82. https://doi.org/10.5281/zenodo.19549486

Similar Articles

11-20 of 98

You may also start an advanced similarity search for this article.