High-Dimensional Data Visualization in Systems Biology

Authors

  • Marta Muller Author
  • Lukas Nowak Author
  • Marta Bianchi Author

DOI:

https://doi.org/10.5281/zenodo.19549612

Keywords:

Dimensionality reduction; Data visualisation; Single-cell analysis; t-SNE; UMAP; Topological data analysis; Systems biology; Multi-omics; Spatial transcriptomics

Abstract

Systems biology generates datasets with thousands to millions of features per sample -- single-cell transcriptomes (20,000 genes x 10^6 cells), spatial transcriptomics (10,000 genes x 10^5 spots), multi-omics profiles (>100,000 combined features), and biological network states (>10^4 nodes) -- yet human visual cognition is limited to 2-3 spatial dimensions. Dimensionality reduction (DR) methods that project high-dimensional data into visualisable embeddings are therefore essential for biological discovery, from identifying cell types in single-cell atlases to detecting patient subtypes in multi-omics cohorts. However, existing DR methods introduce systematic distortions: t-SNE creates artificial clusters and exaggerates local separation, UMAP preserves local but not global structure, and PCA captures only linear variance. This study developed and benchmarked six dimensionality reduction approaches -- PCA, t-SNE, UMAP, diffusion maps, PHATE, and a proposed Hierarchical Topology-Preserving Embedding (HiToP) combining multi-scale neighbourhood graph construction, topology-preserving loss (persistent homology matching), hierarchical zoom levels from global overview to local detail, and uncertainty quantification per embedded point -- across five systems biology visualisation tasks: single-cell atlas visualisation (Human Cell Atlas, 1.2 million cells), spatial transcriptomics tissue mapping (10x Visium, 48 tissue sections), cancer multi-omics subtyping (TCGA, 9,624 tumours), developmental trajectory inference (mouse gastrulation, 116,312 cells), and biological network visualisation (STRING PPI, 18,000 nodes). HiToP achieved the highest composite fidelity score (0.886) across six quantitative metrics: local structure preservation (trustworthiness 0.962), global structure preservation (Spearman r of pairwise distances 0.824), topology preservation (persistent homology Wasserstein distance 0.042), cluster separation accuracy (ARI 0.924 vs ground-truth labels), trajectory continuity (Kendall tau 0.892 for pseudotime), and scalability (1.2 million cells in 8.4 minutes). These results establish topology-preserving hierarchical embedding as the state-of-the-art for systems biology data visualisation.

Downloads

Published

2026-08-15

How to Cite

High-Dimensional Data Visualization in Systems Biology. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 4(2), 84-92. https://doi.org/10.5281/zenodo.19549612

Similar Articles

1-10 of 97

You may also start an advanced similarity search for this article.