Methods › Computer Vision › Image Models

Image Models

33 methods 3,998 papers tagged archive 2025-07-28

The archive attaches this collection's text per method and the copies differ: 3 distinct texts across 33 of the 33 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.

Text 1, carried by 26 of 33 methods:

Image Models are methods that build representations of images for downstream tasks such as classification and object detection. The most popular subcategory are convolutional neural networks. Below you can find a continuously updated list of image models.

Text 2, carried by 6 of 33 methods:

Vision Transformers are Transformer-like models applied to visual tasks. They stem from the work of ViT which directly applied a Transformer architecture on non-overlapping medium-sized image patches for image classification. Below you can find a continually updating list of vision transformers.

According to [1], ViT type models can be further categorized into uniform scale ViTs, multi-scale ViT, hybrid ViTs with convolutions, and self-supervised ViTs. The methods listed below provide a comprehensive overview of ViT models applied to a range of vision tasks.

[1] Transformers in Vision: A Survey

Text 3, carried by 1 of 33 methods:

Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.

Methods

All 33 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

Vision Transformer – 2,144
Interpretability – 1,322
EfficientNet – 195
MLP-Mixer – 96
DeiT Data-efficient Image Transformer – 93
WideResNet – 58
DPT Dense Prediction Transformer – 25
Res2Net – 25
MetaFormer – 20
ResNeSt – 13
CvT Convolutional Vision Transformer – 12
TNT Transformer in Transformer – 12
SANet Self-Attention Network – 11
Bottleneck Transformer – 10
ResMLP Residual Multi-Layer Perceptrons – 10
PoolFormer – 9
gMLP – 7
IRN Invertible Rescaling Network – 6
PVTv2 Pyramid Vision Transformer v2 – 4
ConViT – 3
CrossViT – 3
LR-Net – 3
DeepViT – 2
HaloNet – 2
IICNet – 2
ProxylessNet-CPU – 2
ProxylessNet-GPU – 2
ProxylessNet-Mobile – 2
AdAct Adaptive Activation – 1
DeepSIM – 1
Medical Image Deblurring – 1
MetaNeXt – 1
PCNN Ranking Probable-Class Nearest-Neighbor Ranking – 1