Methods › Computer Vision › Multi-Modal Methods
Multi-Modal Methods
Vision Transformers are Transformer-like models applied to visual tasks. They stem from the work of ViT which directly applied a Transformer architecture on non-overlapping medium-sized image patches for image classification. Below you can find a continually updating list of vision transformers.
According to [1], ViT type models can be further categorized into uniform scale ViTs, multi-scale ViT, hybrid ViTs with convolutions, and self-supervised ViTs. The methods listed below provide a comprehensive overview of ViT models applied to a range of vision tasks.
Methods
All 11 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| GLIDE Guided Language to Image Diffusion for Generation and Editing | – | 28 |
| EmbraceNet EmbraceNet: A robust deep learning architecture for multimodal classification | – | 4 |
| MAVL Multiscale Attention ViT with Late fusion | – | 4 |
| UNIMO | – | 4 |
| VATT | – | 3 |
| Vokenization | – | 3 |
| CTAL | – | 2 |
| SyCoCa Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment | – | 2 |
| AVSlowFast Audiovisual SlowFast Network | – | 1 |
| DSiRe Dataset Size Recovery | – | 1 |
| PO3D-VQA Parts, Poses, and Occlusions in 3D Visual Question Answering | – | 1 |