Methods › Computer Vision › Multi-Modal Methods

Multi-Modal Methods

11 methods 53 papers tagged archive 2025-07-28

Vision Transformers are Transformer-like models applied to visual tasks. They stem from the work of ViT which directly applied a Transformer architecture on non-overlapping medium-sized image patches for image classification. Below you can find a continually updating list of vision transformers.

According to [1], ViT type models can be further categorized into uniform scale ViTs, multi-scale ViT, hybrid ViTs with convolutions, and self-supervised ViTs. The methods listed below provide a comprehensive overview of ViT models applied to a range of vision tasks.

[1] Transformers in Vision: A Survey

Methods

All 11 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

GLIDE Guided Language to Image Diffusion for Generation and Editing – 28
EmbraceNet EmbraceNet: A robust deep learning architecture for multimodal classification – 4
MAVL Multiscale Attention ViT with Late fusion – 4
UNIMO – 4
VATT – 3
Vokenization – 3
CTAL – 2
SyCoCa Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment – 2
AVSlowFast Audiovisual SlowFast Network – 1
DSiRe Dataset Size Recovery – 1
PO3D-VQA Parts, Poses, and Occlusions in 3D Visual Question Answering – 1