Papers › Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective

Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective

7 Feb 2023arXiv:2302.03751archive 2025-07-28

Haoran Zhu, Boyuan Chen, Carter Yang

Vision transformer (ViT) is an attention neural network architecture that is shown to be effective for computer vision tasks. However, compared to ResNet-18 with a similar number of parameters, ViT has a significantly lower evaluation accuracy when trained on small datasets. To facilitate studies in related fields, we provide a visual intuition to help understand why it is the case. We first compare the performance of the two models and confirm that ViT has less accuracy than ResNet-18 when trained on small datasets. We then interpret the results by showing attention map visualization for ViT and feature map visualization for ResNet-18. The difference is further analyzed through a representation similarity perspective. We conclude that the representation of ViT trained on small datasets is hugely different from ViT trained on large datasets, which may be the reason why the performance drops a lot on small datasets.

PaperPDFCode

Code

boyuanjackchen/miniproject2_vistrans officialmentioned in papermentioned on GitHubpytorch report
kentaroy47/vision-transformers-cifar10 officialmentioned in papermentioned on GitHubpytorchMIT report
Ugenteraan/Vanilla-ViT mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections