High-Dimensional Data Visualization in Systems Biology
DOI:
https://doi.org/10.5281/zenodo.19549612Keywords:
Dimensionality reduction; Data visualisation; Single-cell analysis; t-SNE; UMAP; Topological data analysis; Systems biology; Multi-omics; Spatial transcriptomicsAbstract
Systems biology generates datasets with thousands to millions of features per sample -- single-cell transcriptomes (20,000 genes x 10^6 cells), spatial transcriptomics (10,000 genes x 10^5 spots), multi-omics profiles (>100,000 combined features), and biological network states (>10^4 nodes) -- yet human visual cognition is limited to 2-3 spatial dimensions. Dimensionality reduction (DR) methods that project high-dimensional data into visualisable embeddings are therefore essential for biological discovery, from identifying cell types in single-cell atlases to detecting patient subtypes in multi-omics cohorts. However, existing DR methods introduce systematic distortions: t-SNE creates artificial clusters and exaggerates local separation, UMAP preserves local but not global structure, and PCA captures only linear variance. This study developed and benchmarked six dimensionality reduction approaches -- PCA, t-SNE, UMAP, diffusion maps, PHATE, and a proposed Hierarchical Topology-Preserving Embedding (HiToP) combining multi-scale neighbourhood graph construction, topology-preserving loss (persistent homology matching), hierarchical zoom levels from global overview to local detail, and uncertainty quantification per embedded point -- across five systems biology visualisation tasks: single-cell atlas visualisation (Human Cell Atlas, 1.2 million cells), spatial transcriptomics tissue mapping (10x Visium, 48 tissue sections), cancer multi-omics subtyping (TCGA, 9,624 tumours), developmental trajectory inference (mouse gastrulation, 116,312 cells), and biological network visualisation (STRING PPI, 18,000 nodes). HiToP achieved the highest composite fidelity score (0.886) across six quantitative metrics: local structure preservation (trustworthiness 0.962), global structure preservation (Spearman r of pairwise distances 0.824), topology preservation (persistent homology Wasserstein distance 0.042), cluster separation accuracy (ARI 0.924 vs ground-truth labels), trajectory continuity (Kendall tau 0.892 for pseudotime), and scalability (1.2 million cells in 8.4 minutes). These results establish topology-preserving hierarchical embedding as the state-of-the-art for systems biology data visualisation.Downloads
Published
2026-08-15
Issue
Section
Articles
How to Cite
High-Dimensional Data Visualization in Systems Biology. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 4(2), 84-92. https://doi.org/10.5281/zenodo.19549612

