Distributed AI Model Training

Authors

  • Eva Petrov Author
  • Lukas Costa Author

DOI:

https://doi.org/10.5281/zenodo.19614931

Keywords:

distributed training; data parallelism; model parallelism; DeepSpeed ZeRO; Megatron-LM; FSDP; MFU; tensor parallelism; pipeline parallelism; large language model training

Abstract

raining large-scale AI models requires distributed computing infrastructure that can efficiently orchestrate computation across hundreds to thousands of accelerators. As models have grown from millions to hundreds of billions of parameters, distributed training has evolved from simple data parallelism across a few GPUs to sophisticated combinations of data, model, tensor, and pipeline parallelism across multi-rack supercomputer clusters, with communication efficiency as a first-order design constraint. This study systematically evaluates eight distributed training frameworks and parallelism strategies across four model scales (1B, 7B, 70B, and 400B parameters) on two hardware configurations (A100 cluster and H100 cluster) in terms of training throughput (tokens per second), hardware utilisation (Model FLOPs Utilisation, MFU), communication overhead, and fault tolerance. Frameworks evaluated include PyTorch DDP, DeepSpeed ZeRO (stages 1-3), Megatron-LM, FSDP (Fully Sharded Data Parallel), Colossal-AI, PipelineParallel (GPipe-style), Tensor Parallelism (Megatron-style), and a proposed Adaptive Parallelism Scheduler (APS). APS achieves the highest MFU (52.4%) at 70B scale on A100 by dynamically selecting the optimal parallelism strategy based on model architecture and cluster topology. DeepSpeed ZeRO-3 achieves the best memory efficiency enabling 400B training on 512 A100s. Megatron-LM achieves the highest raw throughput at 7B scale. A parallelism strategy selection framework covering model scale, hardware configuration, and training budget is proposed.

Downloads

Published

2026-08-19

How to Cite

Most read articles by the same author(s)