Integrative Omics Data Harmonization Frameworks
DOI:
https://doi.org/10.5281/zenodo.19543659Keywords:
multi-omics; data harmonisation; batch correction; deep learning integration; network fusion; Bayesian models; OHMI; knowledge graph; genomics; transcriptomics; proteomics; metabolomicsAbstract
Multi-omics studies generate genomic, transcriptomic, proteomic, metabolomic, and epigenomic datasets that must be harmonised across platforms, batches, and cohorts before integrated analysis can yield biological insight, yet no consensus framework exists for this harmonisation process. We evaluated 218 integrative omics harmonisation programmes across bioinformatics centres in Spain, Estonia, and Sweden between 2016 and 2021, spanning five framework categories: statistical batch-correction pipelines, deep-learning latent-space integration, network-based data fusion, Bayesian hierarchical models, and knowledge-graph-guided alignment. An Omics Harmonisation Maturity Index (OHMI) was constructed from five sub-scores -- cross-platform concordance improvement, biological signal preservation, scalability to biobank-scale cohorts, downstream discovery yield, and reproducibility across independent datasets -- with weights from regression against sustained adoption into active multi-omics research pipelines. OHMI correlated with adoption at r = +0.84 and discriminated adopted from non-adopted frameworks with an AUC of 0.882. Statistical batch-correction pipelines scored highest (mean OHMI 0.824), while Bayesian hierarchical models trailed at 0.596. Only 35.3 percent of programmes exceeded the 0.75 threshold. Cross-platform concordance carried the largest regression weight (beta = +0.278), followed by biological signal preservation (beta = +0.230).Downloads
Published
2026-08-25
Issue
Section
Articles
How to Cite
Integrative Omics Data Harmonization Frameworks. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 1(1), 25-32. https://doi.org/10.5281/zenodo.19543659

