Integrative Transcriptome-Proteome Modeling

Authors

  • Jonas Lindberg Author
  • Daniel Kovacs Author
  • Marco Popescu Author

DOI:

https://doi.org/10.5281/zenodo.19549602

Keywords:

Transcriptome-proteome correlation; Protein abundance prediction; Post-transcriptional regulation; Translation efficiency; Deep learning; Multi-omics integration; Codon optimality; Proteome imputation; Cancer proteomics

Abstract

The central dogma of molecular biology posits that genetic information flows from DNA to RNA to protein, yet mRNA abundance explains only 40-60% of protein level variance in human cells, with post-transcriptional regulation (translation efficiency, mRNA stability, protein degradation) accounting for the remainder. This mRNA-protein discordance limits the utility of transcriptomic data -- the most widely available omics modality -- for predicting the functional proteome that directly governs cellular phenotype. Computational models that accurately predict protein abundance from transcriptomic and genomic features would extend the value of existing RNA-seq datasets to proteomic-level insight without costly mass spectrometry experiments. This study developed and benchmarked six transcriptome-to-proteome prediction approaches-- direct mRNA correlation (baseline), codon usage-adjusted prediction, ribosome profiling-informed model (Ribo-RPKM), multi-feature gradient-boosted regression (XGBoost on 142 features), protein language model-enhanced prediction (ESM-2 embeddings), and a proposed Deep Transcriptome-Proteome Translator (DeepTPT) combining mRNA expression, codon optimality scores, 5'/3'-UTR regulatory element predictions, miRNA targeting scores, protein half-life estimates from sequence, and protein-protein interaction context via a multi-input transformer architecture -- across matched transcriptome-proteome datasets from TCGA-CPTAC (11 cancer types, n = 1,486 tumours), GTEx (32 normal tissues), and CCLE (375 cell lines). DeepTPT achieved the highest mRNA-to-protein prediction accuracy: median Pearson r = 0.82 across genes (vs direct mRNA r = 0.61, XGBoost r = 0.74, ESM-2 r = 0.76; p < 0.001), explaining 67.2% of protein variance compared to 37.2% for mRNA alone. DeepTPT-imputed proteomes enabled cancer subtype classification with F1 0.884 (vs actual proteomics F1 0.912; 96.9% of measured proteome utility) and drug response prediction with AUROC 0.824 (vs actual 0.846; 97.4%). Feature attribution confirmed that mRNA abundance contributed 58.4% of predicted protein variance, codon optimality 14.2%, UTR regulation 12.8%, protein stability 8.6%, and PPI context 6.0%. These results establish deep transcriptome-proteome translation as a viable computational alternative to mass spectrometry for large-scale proteomic inference

Downloads

Published

2026-08-15

How to Cite

Integrative Transcriptome-Proteome Modeling. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 4(2), 75-83. https://doi.org/10.5281/zenodo.19549602

Similar Articles

1-10 of 92

You may also start an advanced similarity search for this article.