{"url":"/method/align","slug":"align","name":"ALIGN","full_name":"ALIGN","full_name_withheld":false,"description_markdown":"In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss (formulated as normalized softmax) that pushes the embeddings of the matched image-text pair together and pushing those of non-matched image-text pair apart. The model learns to align visual and language representations of the image and text pairs using the contrastive loss. The representations can be used for vision-only or vision-language task transfer. Without any fine-tuning, ALIGN powers zero-shot visual classification and cross-modal search including image-to-text search, text-to image search and even search with joint image+text queries.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision","paper":"/paper/scaling-up-visual-and-vision-language","first_author":"Chao Jia","n_authors":10,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/scaling-up-visual-and-vision-language"},"source":{"url":"https://arxiv.org/abs/2102.05918v2","title":"Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":5527,"archive_num_papers":5527,"papers_newest_first":[{"paper":null,"title":"SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation","date":"2025-07-16","arxiv_id":"2507.12027","n_code_links":0,"syntology":null},{"paper":null,"title":"Task-Oriented Human Grasp Synthesis via Context- and Task-Aware Diffusers","date":"2025-07-15","arxiv_id":"2507.11287","n_code_links":0,"syntology":null},{"paper":null,"title":"Toward Improving fNIRS Classification: A Study on Activation Functions in Deep Neural Architectures","date":"2025-07-15","arxiv_id":"2507.11436","n_code_links":0,"syntology":null},{"paper":null,"title":"Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning","date":"2025-07-14","arxiv_id":"2507.10348","n_code_links":0,"syntology":null},{"paper":"/paper/scooter-a-human-evaluation-framework-for","title":"SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples","date":"2025-07-10","arxiv_id":"2507.07776","n_code_links":1,"syntology":null},{"paper":null,"title":"Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset","date":"2025-07-09","arxiv_id":"2507.07339","n_code_links":0,"syntology":null},{"paper":null,"title":"Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey","date":"2025-07-09","arxiv_id":"2507.07148","n_code_links":0,"syntology":null},{"paper":"/paper/investalign-overcoming-data-scarcity-in","title":"InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior","date":"2025-07-09","arxiv_id":"2507.06528","n_code_links":1,"syntology":null},{"paper":null,"title":"ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion","date":"2025-07-08","arxiv_id":"2507.05624","n_code_links":0,"syntology":null},{"paper":"/paper/langmamba-a-language-driven-mamba-framework","title":"LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models","date":"2025-07-08","arxiv_id":"2507.06140","n_code_links":1,"syntology":null},{"paper":"/paper/mcam-multimodal-causal-analysis-model-for-ego","title":"MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding","date":"2025-07-08","arxiv_id":"2507.06072","n_code_links":1,"syntology":{"ran":0,"of":3,"unverified":3,"pointer_only":3}},{"paper":"/paper/scoreadv-score-based-targeted-generation-of","title":"ScoreAdv: Score-based Targeted Generation of Natural Adversarial Examples via Diffusion Models","date":"2025-07-08","arxiv_id":"2507.06078","n_code_links":1,"syntology":null},{"paper":null,"title":"Vers un cadre ontologique pour la gestion des comp{é}tences : {à} des fins de formation, de recrutement, de m{é}tier, ou de recherches associ{é}es","date":"2025-07-08","arxiv_id":"2507.05767","n_code_links":0,"syntology":null},{"paper":null,"title":"Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning","date":"2025-07-07","arxiv_id":"2507.05418","n_code_links":0,"syntology":null},{"paper":"/paper/neural-driven-image-editing","title":"Neural-Driven Image Editing","date":"2025-07-07","arxiv_id":"2507.05397","n_code_links":1,"syntology":null},{"paper":null,"title":"CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step","date":"2025-07-06","arxiv_id":"2507.04451","n_code_links":0,"syntology":null},{"paper":null,"title":"Rectifying Adversarial Sample with Low Entropy Prior for Test-Time Defense","date":"2025-07-04","arxiv_id":"2507.03427","n_code_links":0,"syntology":null},{"paper":null,"title":"Adopting a human developmental visual diet yields robust, shape-based AI vision","date":"2025-07-03","arxiv_id":"2507.03168","n_code_links":0,"syntology":null},{"paper":null,"title":"De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks","date":"2025-07-03","arxiv_id":"2507.02606","n_code_links":0,"syntology":null},{"paper":"/paper/hita-holistic-tokenizer-for-autoregressive","title":"Hita: Holistic Tokenizer for Autoregressive Image Generation","date":"2025-07-03","arxiv_id":"2507.02358","n_code_links":0,"syntology":{"ran":4,"of":11,"unverified":7,"pointer_only":0}},{"paper":"/paper/ld-rps-zero-shot-unified-image-restoration","title":"LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling","date":"2025-07-01","arxiv_id":"2507.00790","n_code_links":1,"syntology":{"ran":3,"of":8,"unverified":5,"pointer_only":8}},{"paper":null,"title":"Large Language Models Don't Make Sense of Word Problems. A Scoping Review from a Mathematics Education Perspective","date":"2025-06-30","arxiv_id":"2506.24006","n_code_links":0,"syntology":null},{"paper":null,"title":"Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging","date":"2025-06-29","arxiv_id":"2506.23266","n_code_links":0,"syntology":null},{"paper":"/paper/selecting-and-merging-towards-adaptable-and","title":"Selecting and Merging: Towards Adaptable and Scalable Named Entity Recognition with Large Language Models","date":"2025-06-28","arxiv_id":"2506.22813","n_code_links":1,"syntology":null},{"paper":"/paper/class-agnostic-region-of-interest-matching-in","title":"Class-Agnostic Region-of-Interest Matching in Document Images","date":"2025-06-26","arxiv_id":"2506.21055","n_code_links":1,"syntology":null},{"paper":null,"title":"Deception Detection in Dyadic Exchanges Using Multimodal Machine Learning: A Study on a Swedish Cohort","date":"2025-06-26","arxiv_id":"2506.21429","n_code_links":0,"syntology":null},{"paper":null,"title":"Elucidating and Endowing the Diffusion Training Paradigm for General Image Restoration","date":"2025-06-26","arxiv_id":"2506.21722","n_code_links":0,"syntology":null},{"paper":"/paper/enhancing-homophily-heterophily-separation","title":"Enhancing Homophily-Heterophily Separation: Relation-Aware Learning in Heterogeneous Graphs","date":"2025-06-26","arxiv_id":"2506.20980","n_code_links":1,"syntology":null},{"paper":null,"title":"Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments","date":"2025-06-26","arxiv_id":"2506.21497","n_code_links":0,"syntology":null},{"paper":"/paper/hierarchical-sub-action-tree-for-continuous","title":"Hierarchical Sub-action Tree for Continuous Sign Language Recognition","date":"2025-06-26","arxiv_id":"2506.20947","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":409},{"task":"/task/language-modeling","name":"Language Modeling","papers":320},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":263},{"task":"/task/retrieval","name":"Retrieval","papers":239},{"task":"/task/large-language-model","name":"Large Language Model","papers":218},{"task":"/task/domain-adaptation","name":"Domain Adaptation","papers":211},{"task":"/task/image-generation","name":"Image Generation","papers":185},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":176},{"task":"/task/representation-learning","name":"Representation Learning","papers":167},{"task":"/task/question-answering","name":"Question Answering","papers":161},{"task":"/task/object-detection","name":"Object Detection","papers":149},{"task":"/task/object-detection-1","name":"object-detection","papers":146},{"task":"/task/object","name":"Object","papers":140},{"task":"/task/decision-making","name":"Decision Making","papers":137},{"task":"/task/diversity","name":"Diversity","papers":132},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":125},{"task":"/task/segmentation","name":"Segmentation","papers":116},{"task":"/task/text-generation","name":"Text Generation","papers":108},{"task":"/task/recommendation-systems","name":"Recommendation Systems","papers":106},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":104}],"tasks_shown":20,"n_tasks":1380,"usage_by_year":[{"year":"2021","papers":24},{"year":"2022","papers":477},{"year":"2023","papers":1227},{"year":"2024","papers":2346},{"year":"2025","papers":1453}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/align"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}