Generative AI for Code Synthesis
DOI:
https://doi.org/10.5281/zenodo.19614759Keywords:
code synthesis; generative AI; large language models; HumanEval; bug repair; code translation; test-case generation; software engineeringAbstract
Generative AI has moved code synthesis from an academic curiosity to a mainstream software development tool in under three years. Yet systematic evidence on where these models succeed, where they fail, and how their outputs should be validated in production remains sparse. This study evaluates seven generative AI systems for code synthesis-- GPT-4, Claude 3 Opus, Gemini 1.5 Pro, CodeLlama-34B, DeepSeek-Coder-33B, WizardCoder-34B, and StarCoder2-15B -- across four synthesis tasks: function-level generation, bug repair, code translation, and test-case synthesis. Evaluation used HumanEval, MBPP, SWE-bench, and a novel multi-language translation suite covering Python, Java, C++, Rust, and TypeScript. We find that frontier models (GPT-4, Claude 3 Opus, Gemini 1.5 Pro) achieve HumanEval pass@1 scores of 82.4-86.8%, compared to 58.6-67.2% for open-weight models. Bug repair on SWE-bench reveals a sharper gap: frontier models resolve 18.4-22.6% of issues vs. 8.2-12.4% for open-weight models. Code translation accuracy degrades substantially for Rust targets (14.2-28.6% reduction vs. Python targets), revealing systematic weakness in low-resource language generation. Test-case synthesis achieves branch coverage of 61.4-74.8% -- competitive with manual testing in constrained settings. A structured validation framework for production deployment of AI-generated code is proposed, covering static analysis, adversarial test generation, and human review thresholdsDownloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Generative AI for Code Synthesis. (2026). Bio-QI Journal, 3(1), 17-24. https://doi.org/10.5281/zenodo.19614759

