Methods › Computer Vision › Vision and Language Pre-Trained Models › ALIGN

ALIGN

5,527 papers tagged archive 2025-07-28

Introduced by Chao Jia et al. in Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss (formulated as normalized softmax) that pushes the embeddings of the matched image-text pair together and pushing those of non-matched image-text pair apart. The model learns to align visual and language representations of the image and text pairs using the contrastive loss. The representations can be used for vision-only or vision-language task transfer. Without any fine-tuning, ALIGN powers zero-shot visual classification and cross-modal search including image-to-text search, text-to image search and even search with joint image+text queries.

PaperSource

Papers archive 2025-07-28

30 shown of 5,527, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 1,380 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling409
Language Modeling320
Contrastive Learning263
Retrieval239
Large Language Model218
Domain Adaptation211
Image Generation185
Semantic Segmentation176
Representation Learning167
Question Answering161
Object Detection149
object-detection146
Object140
Decision Making137
Diversity132
Transfer Learning125
Segmentation116
Text Generation108
Recommendation Systems106
reinforcement-learning104

Usage over time archive 2025-07-28

Papers per year tagged with ALIGN: 2021 to 2025, peak 2,346 2,346 0 2021: 24 papers 2021 2022: 477 papers 2022 2023: 1227 papers 2023 2024: 2346 papers 2024 2025: 1453 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (5,527 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections