Methods › Computer Vision › Vision and Language Pre-Trained Models › ALBEF
ALBEF
Introduced by Junnan Li et al. in Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
ALBEF introduces a contrastive loss to align the image and text representations before fusing them through cross-modal attention. This enables more grounded vision and language representation learning. ALBEF also doesn't require bounding box annotations. The model consists of an image encode, a text encoder, and a multimodal encoder. The image-text contrastive loss helps to align the unimodal representations of an image-text pair before fusion. The image-text matching loss and a masked language modeling loss are applied to learn multimodal interactions between image and text. In addition, momentum distillation is used to generate pseudo-targets. This improves learning with noisy data.
Papers archive 2025-07-28
17 shown of 17, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses 11 Dec 2024 · 0 repositories · arXiv:2412.08110
-
Nearest Neighbor Normalization Improves Multimodal Retrieval 31 Oct 2024 · 1 repository · arXiv:2410.24114
-
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM 29 Apr 2024 · 0 repositories · arXiv:2404.19128
-
Learning from Synthetic Data for Visual Grounding 20 Mar 2024 · 0 repositories · arXiv:2403.13804
-
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction 16 Mar 2024 · 0 repositories · arXiv:2403.10883
-
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival 16 Mar 2024 · 0 repositories · arXiv:2403.10887
-
Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models 26 Jul 2023 · 1 repository · arXiv:2307.14061Syntology ran 7 of 12 samples · 5 unverified
-
RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search 23 May 2023 · 1 repository · arXiv:2305.13653Syntology ran 3 of 7 samples · 4 unverified
-
MultiModal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision Language Models 16 Mar 2023 · 1 repository · arXiv:2303.12734
-
Is Multimodal Vision Supervision Beneficial to Language? 10 Feb 2023 · 1 repository · arXiv:2302.05016
-
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks 15 Dec 2022 · 1 repository · arXiv:2212.08158Syntology ran 1 of 10 samples · 9 unverified
-
Leveraging per Image-Token Consistency for Vision-Language Pre-training 20 Nov 2022 · 0 repositories · arXiv:2211.15398
-
GRIT-VLP: Grouped Mini-batch Sampling for Efficient Vision and Language Pre-training 8 Aug 2022 · 1 repository · arXiv:2208.04060Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal Understanding 5 Aug 2022 · 1 repository · arXiv:2208.03030
-
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations 1 Jul 2022 · 1 repository · arXiv:2207.00221
-
MixGen: A New Multi-Modal Data Augmentation 16 Jun 2022 · 1 repository · arXiv:2206.08358Syntology ran 0 of 2 samples · 2 unverified
-
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation 16 Jul 2021 · 6 repositories · arXiv:2107.07651Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 38 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections