Generative Adversarial Networks in Synthetic Data Creation
DOI:
https://doi.org/10.5281/zenodo.19614539Keywords:
generative adversarial networks; synthetic data; data augmentation; diffusion models; tabular data generation; privacy-preserving synthesis; class imbalance; medical image synthesisAbstract
Synthetic data generation has become essential for training machine learning models in domains where real data is scarce, sensitive, or imbalanced. Generative adversarial networks offer a powerful framework for producing realistic synthetic samples, yet the practical question of which GAN variant to use for which data modality -- and whether GANs outperform non-adversarial alternatives -- remains inadequately addressed. This study presents a controlled comparison of six generative methods -- DCGAN, StyleGAN2, CTGAN (tabular), TimeGAN (temporal), a variational autoencoder (VAE) baseline, and diffusion models (DDPM) -- across four synthetic data applications: medical image augmentation (chest X-ray, retinal fundus), tabular data synthesis (census, credit), time series generation (financial, clinical), and class-imbalance oversampling (fraud, rare disease). Quality was assessed on three dimensions: fidelity (statistical similarity to real data via FID, MMD, and column-wise distributional tests), utility (downstream ML performance when training on synthetic data), and privacy (resistance to membership inference attacks measuring whether individual real records can be identified from synthetic output). A total of 2,880 experiments were conducted. StyleGAN2 achieved the best image fidelity (FID = 8.4 +- 1.2 on chest X-rays) and downstream utility (classifier AUC within 1.8% of real-data baseline). CTGAN produced the highest-utility tabular data (downstream AUC within 2.4% of real). Diffusion models matched or exceeded GAN fidelity on images (FID = 7.8 +- 0.9) with superior training stability but 4.2x slower generation. The privacy-utility trade-off was quantified: adding differential privacy to CTGAN (epsilon = 10) reduced membership inference advantage from 12.4% to 3.8% at a utility cost of 4.6 AUC points. A practical selection framework mapping data modality, quality requirements, and privacy constraints to recommended generative configurations is proposedDownloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Generative Adversarial Networks in Synthetic Data Creation. (2026). Bio-QI Journal, 1(2), 75-83. https://doi.org/10.5281/zenodo.19614539

