Generative AI for Code Synthesis

Authors

  • Sofia Kovacs Author
  • Oscar Kovacs Author
  • Daniel Kovacs Author

DOI:

https://doi.org/10.5281/zenodo.19614759

Keywords:

code synthesis; generative AI; large language models; HumanEval; bug repair; code translation; test-case generation; software engineering

Abstract

Generative AI has moved code synthesis from an academic curiosity to a mainstream software development tool in under three years. Yet systematic evidence on where these models succeed, where they fail, and how their outputs should be validated in production remains sparse. This study evaluates seven generative AI systems for code synthesis-- GPT-4, Claude 3 Opus, Gemini 1.5 Pro, CodeLlama-34B, DeepSeek-Coder-33B, WizardCoder-34B, and StarCoder2-15B -- across four synthesis tasks: function-level generation, bug repair, code translation, and test-case synthesis. Evaluation used HumanEval, MBPP, SWE-bench, and a novel multi-language translation suite covering Python, Java, C++, Rust, and TypeScript. We find that frontier models (GPT-4, Claude 3 Opus, Gemini 1.5 Pro) achieve HumanEval pass@1 scores of 82.4-86.8%, compared to 58.6-67.2% for open-weight models. Bug repair on SWE-bench reveals a sharper gap: frontier models resolve 18.4-22.6% of issues vs. 8.2-12.4% for open-weight models. Code translation accuracy degrades substantially for Rust targets (14.2-28.6% reduction vs. Python targets), revealing systematic weakness in low-resource language generation. Test-case synthesis achieves branch coverage of 61.4-74.8% -- competitive with manual testing in constrained settings. A structured validation framework for production deployment of AI-generated code is proposed, covering static analysis, adversarial test generation, and human review thresholds

Downloads

Published

2026-08-19

How to Cite