Diffusion Models in Generative AI

Authors

  • Hugo Nowak Author
  • Anna Popescu Author
  • Daniel Silva Author

DOI:

https://doi.org/10.5281/zenodo.19614633

Keywords:

diffusion models; denoising diffusion; latent diffusion; text-to-image generation; classifier-free guidance; consistency models; image synthesis; generative AI

Abstract

Diffusion models have emerged as the dominant paradigm for high-fidelity generative AI, surpassing GANs on image synthesis benchmarks while offering superior training stability, mode coverage, and controllability. This study presents a controlled comparison of five diffusion model variants -- DDPM (baseline), DDIM (accelerated sampling), latent diffusion (Stable Diffusion architecture), classifier-free guidance, and consistency models -- across four generative tasks: unconditional image generation (CIFAR-10, LSUN-Bedroom 256x256), text-to-image synthesis (MS-COCO, evaluated by FID and CLIP score), image inpainting (Places2), and audio generation (speech synthesis on LJSpeech). All models were trained under standardised compute budgets on identical hardware. A total of 1,680 experiments were conducted. Latent diffusion achieved the best image quality (FID = 4.82 +- 0.24 on LSUN-Bedroom) at 8.4x lower computational cost than pixel-space diffusion by operating in a compressed latent space. Classifier-free guidance improved text-image alignment (CLIP score = 0.314 +- 0.004 vs. 0.286 for unguided) at the cost of reduced diversity. Consistency models achieved 64x faster sampling than DDPM (single-step generation) with only 12.4% FID degradation, enabling real-time applications. DDIM reduced sampling steps from 1000 to 50 with 4.2% FID increase. The quality-speed Pareto frontier was mapped: latent diffusion with 50-step DDIM sampling achieved FID = 5.24 in 0.8 seconds per image on A100, representing the current optimal trade-off for production deployment. A practical diffusion model selection guide mapping generation quality, speed requirements, and controllability needs to recommended configurations is proposed.

Downloads

Published

2026-08-19

How to Cite

Diffusion Models in Generative AI. (2026). Bio-QI  Journal, 2(1), 10-18. https://doi.org/10.5281/zenodo.19614633