Big Data Architectures for Genomic Databases
DOI:
https://doi.org/10.5281/zenodo.19543762Keywords:
big data; genomic database; data lakehouse; columnar store; Apache Spark; GDAI; variant query; data governance; GDPR; federated query; Hail; cloud-nativeAbstract
Population-scale genomic databases now store petabytes of variant calls, phenotype annotations, and multi-omics measurements for millions of individuals, demanding data architectures that balance query performance, storage efficiency, data governance, and analytical flexibility in ways that traditional relational databases cannot achieve. We evaluated 214 big data architecture programmes for genomic databases across centres in Estonia and France between 2017 and 2021, spanning five architecture categories: columnar analytical stores, distributed file systems with query engines, graph databases for variant-phenotype networks, cloud-native data lakehouse platforms, and federated query systems for multi-site genomic data. A Genomic Database Architecture Index (GDAI) was constructed from five sub-scores -- variant query throughput, storage compression efficiency, analytical workload support, data governance and access control, and interoperability with bioinformatics tools -- with weights from regression against sustained institutional adoption beyond 12 months. GDAI correlated with adoption at r = +0.84 and discriminated adopted from non-adopted architectures with an AUC of 0.882. Cloud-native lakehouse platforms scored highest (mean GDAI 0.824), while graph databases trailed at 0.598. Only 35.5 percent of programmes exceeded the 0.75 threshold. Variant query throughput carried the largest regression weight (beta = +0.278), followed by analytical workload support (beta = +0.230).Downloads
Published
2026-08-25
Issue
Section
Articles
How to Cite
Big Data Architectures for Genomic Databases. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 1(3), 113-120. https://doi.org/10.5281/zenodo.19543762

