Methods › Computer Vision › Image Models
Image Models
The archive attaches this collection's text per method and the copies differ: 3 distinct texts across 33 of the 33 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.
Text 1, carried by 26 of 33 methods:
Image Models are methods that build representations of images for downstream tasks such as classification and object detection. The most popular subcategory are convolutional neural networks. Below you can find a continuously updated list of image models.
Text 2, carried by 6 of 33 methods:
Vision Transformers are Transformer-like models applied to visual tasks. They stem from the work of ViT which directly applied a Transformer architecture on non-overlapping medium-sized image patches for image classification. Below you can find a continually updating list of vision transformers.
According to [1], ViT type models can be further categorized into uniform scale ViTs, multi-scale ViT, hybrid ViTs with convolutions, and self-supervised ViTs. The methods listed below provide a comprehensive overview of ViT models applied to a range of vision tasks.
Text 3, carried by 1 of 33 methods:
Transformers are a type of neural network architecture that have several properties that make them effective for modeling data with long-range dependencies. They generally feature a combination of multi-headed attention mechanisms, residual connections, layer normalization, feedforward connections, and positional embeddings.
Methods
All 33 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Vision Transformer | – | 2,144 |
| Interpretability | – | 1,322 |
| EfficientNet | – | 195 |
| MLP-Mixer | – | 96 |
| DeiT Data-efficient Image Transformer | – | 93 |
| WideResNet | – | 58 |
| DPT Dense Prediction Transformer | – | 25 |
| Res2Net | – | 25 |
| MetaFormer | – | 20 |
| ResNeSt | – | 13 |
| CvT Convolutional Vision Transformer | – | 12 |
| TNT Transformer in Transformer | – | 12 |
| SANet Self-Attention Network | – | 11 |
| Bottleneck Transformer | – | 10 |
| ResMLP Residual Multi-Layer Perceptrons | – | 10 |
| PoolFormer | – | 9 |
| gMLP | – | 7 |
| IRN Invertible Rescaling Network | – | 6 |
| PVTv2 Pyramid Vision Transformer v2 | – | 4 |
| ConViT | – | 3 |
| CrossViT | – | 3 |
| LR-Net | – | 3 |
| DeepViT | – | 2 |
| HaloNet | – | 2 |
| IICNet | – | 2 |
| ProxylessNet-CPU | – | 2 |
| ProxylessNet-GPU | – | 2 |
| ProxylessNet-Mobile | – | 2 |
| AdAct Adaptive Activation | – | 1 |
| DeepSIM | – | 1 |
| Medical Image Deblurring | – | 1 |
| MetaNeXt | – | 1 |
| PCNN Ranking Probable-Class Nearest-Neighbor Ranking | – | 1 |