Vision Transformers in Image Recognition

Authors

  • Andreas Muller Author
  • Laura Hansen Author
  • Laura Ivanov Author

DOI:

https://doi.org/10.5281/zenodo.19614669

Keywords:

vision transformer; image classification; Swin Transformer; ConvNeXt; robustness; fine-grained recognition; attention mechanism; transfer learning

Abstract

Vision transformers have challenged the decade-long dominance of convolutional neural networks in image recognition, demonstrating that the self-attention mechanism can capture visual features as effectively as convolutions when trained with sufficient data and appropriate regularisation. This study presents a controlled comparison of six vision architectures-- standard ViT, DeiT (data-efficient ViT with distillation), Swin Transformer (hierarchical shifted windows), ConvNeXt (modernised CNN), hybrid CoAtNet (convolution + attention), and EfficientNetV2 (CNN baseline) -- across four image recognition tasks: ImageNet-1K classification, fine-grained recognition (CUB-200 birds, Stanford Cars), out-of-distribution robustness (ImageNet-C, ImageNet-R, ImageNet-Sketch), and transfer learning to small datasets (CIFAR-100, Oxford Flowers). All models were evaluated at three computational scales (small ~22M, base ~86M, large ~300M parameters) with matched training recipes to isolate architectural effects. A total of 2,520 experiments were conducted. Swin Transformer achieved the highest ImageNet-1K accuracy at all scales (84.5 +- 0.2% at base), outperforming ViT (81.8%), DeiT (83.1%), ConvNeXt (84.1%), CoAtNet (84.3%), and EfficientNetV2 (83.6%). On out-of-distribution robustness, ViT variants showed 18.4% smaller effective robustness gap than CNNs of equivalent accuracy, confirming that attention-based architectures develop more shape-biased representations. ConvNeXt matched Swin's ImageNet accuracy while retaining CNN deployment simplicity (no shifting window logic), suggesting that architectural innovations from transformers can be back-ported to CNNs. On fine-grained recognition, attention-based models outperformed CNNs by 2.4-4.8 points, with attention maps revealing learned part-based decomposition of objects. A practical architecture selection guide mapping dataset size, robustness requirements, and deployment constraints to recommended configurations is proposed.

Downloads

Published

2026-08-19

How to Cite

Vision Transformers in Image Recognition. (2026). Bio-QI  Journal, 2(2), 55-62. https://doi.org/10.5281/zenodo.19614669