Large-Scale Pretraining Strategies
DOI:
https://doi.org/10.5281/zenodo.19614752Keywords:
large-scale pretraining; self-supervised learning; transfer learning; masked language modelling; contrastive learning; curriculum learning; compute scaling; foundation modelsAbstract
Large-scale pretraining has fundamentally reshaped what neural networks are capable of. By exposing models to massive corpora before task-specific fine-tuning, practitioners have achieved performance gains that smaller, supervised-only approaches simply cannot replicate. Yet the design space for pretraining -- which objectives to use, how much compute to allocate, how to schedule data across modalities -- remains poorly systematised. This study benchmarks eight pretraining strategies across four architectural families (transformer encoders, decoder-only language models, vision-language models, and graph neural networks), evaluating downstream task performance on twelve benchmarks covering natural language understanding, code generation, visual question answering, and molecular property prediction. Experiments used compute budgets ranging from 10^17 to 10^21 FLOPs. We find that masked-language modelling (MLM) and next-token prediction (NTP) remain the strongest single objectives for text, but contrastive pretraining yields a 4.2-6.8 percentage point advantage on cross-modal tasks. Curriculum scheduling - gradually increasing sequence length and data difficulty -- reduces training instability and improves final accuracy by 2.1-3.7% relative across encoders. Compute scaling follows a power-law relationship (R2 = 0.94) between FLOPs and validation loss, with diminishing returns becoming pronounced beyond 10^20 FLOPs for unimodal tasks. A practitioner decision framework mapping task type and compute budget to optimal pretraining strategy is proposed.Downloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Large-Scale Pretraining Strategies. (2026). Bio-QI Journal, 3(1), 1-8. https://doi.org/10.5281/zenodo.19614752

