AI Security and Adversarial Attack Mitigation

Authors

  • Clara Kovacs Author
  • Helena Popescu Author

DOI:

https://doi.org/10.5281/zenodo.19614599

Keywords:

adversarial attacks; adversarial training; certified robustness; randomised smoothing; adversarial patches; AI security; robust deep learning; defence-in-depth

Abstract

Deep learning models are vulnerable to adversarial attacks -- carefully crafted input perturbations that cause confident misclassification while remaining imperceptible to humans -- posing serious security risks in safety-critical deployments including autonomous driving, medical diagnosis, and biometric authentication. This study presents a systematic evaluation of six adversarial defence strategies -- adversarial training (PGD-AT), certified robust training (randomised smoothing), input preprocessing (JPEG compression, feature squeezing), adversarial detection (MagNet), model ensembling with diversity, and vision transformer architectures as inherently robust alternatives -- against five attack families: gradient-based (PGD, C&W;), transfer-based (black-box), query-based (Square Attack), patch-based (adversarial patches), and natural corruptions (ImageNet-C). Evaluations were conducted across three security-sensitive domains: traffic sign recognition (GTSRB), face verification (LFW), and medical image classification (ISIC skin lesion). A total of 2,160 attack-defence experiments were conducted under standardised threat models. PGD adversarial training achieved the highest robust accuracy under L-infinity attacks (48.6 +- 1.4% at epsilon = 8/255 on CIFAR-10) but imposed a 7.2% clean accuracy cost. Randomised smoothing provided the only certified robustness guarantee (certified radius = 0.5 at 72.4% accuracy on CIFAR-10) but was computationally expensive at inference (100x slowdown from Monte Carlo sampling). Vision transformers showed 14.2% higher natural robustness than CNNs of equivalent size without any adversarial training, but remained vulnerable to adaptive attacks. No single defence provided comprehensive protection across all attack types, confirming that defence-in-depth combining multiple strategies is essential. A practical security deployment guide mapping threat model, accuracy tolerance, and computational budget to recommended defence configurations is proposed.

Downloads

Published

2026-08-19

How to Cite

AI Security and Adversarial Attack Mitigation. (2026). Bio-QI  Journal, 1(3), 156-164. https://doi.org/10.5281/zenodo.19614599

Similar Articles

1-10 of 98

You may also start an advanced similarity search for this article.