Methods › Computer Vision › Vision and Language Pre-Trained Models › ALIGN
ALIGN
Introduced by Chao Jia et al. in Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss (formulated as normalized softmax) that pushes the embeddings of the matched image-text pair together and pushing those of non-matched image-text pair apart. The model learns to align visual and language representations of the image and text pairs using the contrastive loss. The representations can be used for vision-only or vision-language task transfer. Without any fine-tuning, ALIGN powers zero-shot visual classification and cross-modal search including image-to-text search, text-to image search and even search with joint image+text queries.
Papers archive 2025-07-28
30 shown of 5,527, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
SGLoc: Semantic Localization System for Camera Pose Estimation from 3D Gaussian Splatting Representation 16 Jul 2025 · 0 repositories · arXiv:2507.12027
-
Task-Oriented Human Grasp Synthesis via Context- and Task-Aware Diffusers 15 Jul 2025 · 0 repositories · arXiv:2507.11287
-
Toward Improving fNIRS Classification: A Study on Activation Functions in Deep Neural Architectures 15 Jul 2025 · 0 repositories · arXiv:2507.11436
-
Feature Distillation is the Better Choice for Model-Heterogeneous Federated Learning 14 Jul 2025 · 0 repositories · arXiv:2507.10348
-
SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples 10 Jul 2025 · 1 repository · arXiv:2507.07776
-
Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset 9 Jul 2025 · 0 repositories · arXiv:2507.07339
-
Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey 9 Jul 2025 · 0 repositories · arXiv:2507.07148
-
InvestAlign: Overcoming Data Scarcity in Aligning Large Language Models with Investor Decision-Making Processes under Herd Behavior 9 Jul 2025 · 1 repository · arXiv:2507.06528
-
ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion 8 Jul 2025 · 0 repositories · arXiv:2507.05624
-
LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models 8 Jul 2025 · 1 repository · arXiv:2507.06140
-
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding 8 Jul 2025 · 1 repository · arXiv:2507.06072Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)
-
ScoreAdv: Score-based Targeted Generation of Natural Adversarial Examples via Diffusion Models 8 Jul 2025 · 1 repository · arXiv:2507.06078
-
Vers un cadre ontologique pour la gestion des comp{é}tences : {à} des fins de formation, de recrutement, de m{é}tier, ou de recherches associ{é}es 8 Jul 2025 · 0 repositories · arXiv:2507.05767
-
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning 7 Jul 2025 · 0 repositories · arXiv:2507.05418
-
Neural-Driven Image Editing 7 Jul 2025 · 1 repository · arXiv:2507.05397
-
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step 6 Jul 2025 · 0 repositories · arXiv:2507.04451
-
Rectifying Adversarial Sample with Low Entropy Prior for Test-Time Defense 4 Jul 2025 · 0 repositories · arXiv:2507.03427
-
Adopting a human developmental visual diet yields robust, shape-based AI vision 3 Jul 2025 · 0 repositories · arXiv:2507.03168
-
De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks 3 Jul 2025 · 0 repositories · arXiv:2507.02606
-
Hita: Holistic Tokenizer for Autoregressive Image Generation 3 Jul 2025 · 0 repositories · arXiv:2507.02358Syntology ran 4 of 11 samples · 7 unverified
-
LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling 1 Jul 2025 · 1 repository · arXiv:2507.00790Syntology ran 3 of 8 samples · 5 unverified · 8 pointer-only (licence)
-
Large Language Models Don't Make Sense of Word Problems. A Scoping Review from a Mathematics Education Perspective 30 Jun 2025 · 0 repositories · arXiv:2506.24006
-
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging 29 Jun 2025 · 0 repositories · arXiv:2506.23266
-
Selecting and Merging: Towards Adaptable and Scalable Named Entity Recognition with Large Language Models 28 Jun 2025 · 1 repository · arXiv:2506.22813
-
Class-Agnostic Region-of-Interest Matching in Document Images 26 Jun 2025 · 1 repository · arXiv:2506.21055
-
Deception Detection in Dyadic Exchanges Using Multimodal Machine Learning: A Study on a Swedish Cohort 26 Jun 2025 · 0 repositories · arXiv:2506.21429
-
Elucidating and Endowing the Diffusion Training Paradigm for General Image Restoration 26 Jun 2025 · 0 repositories · arXiv:2506.21722
-
Enhancing Homophily-Heterophily Separation: Relation-Aware Learning in Heterogeneous Graphs 26 Jun 2025 · 1 repository · arXiv:2506.20980
-
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments 26 Jun 2025 · 0 repositories · arXiv:2506.21497
-
Hierarchical Sub-action Tree for Continuous Sign Language Recognition 26 Jun 2025 · 1 repository · arXiv:2506.20947
Tasks archive 2025-07-28
20 shown of 1,380 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Language Modelling | 409 |
| Language Modeling | 320 |
| Contrastive Learning | 263 |
| Retrieval | 239 |
| Large Language Model | 218 |
| Domain Adaptation | 211 |
| Image Generation | 185 |
| Semantic Segmentation | 176 |
| Representation Learning | 167 |
| Question Answering | 161 |
| Object Detection | 149 |
| object-detection | 146 |
| Object | 140 |
| Decision Making | 137 |
| Diversity | 132 |
| Transfer Learning | 125 |
| Segmentation | 116 |
| Text Generation | 108 |
| Recommendation Systems | 106 |
| reinforcement-learning | 104 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections