{"url":"/method/albef","slug":"albef","name":"ALBEF","full_name":"ALBEF","full_name_withheld":false,"description_markdown":"ALBEF introduces a contrastive loss to align the image and text representations before fusing them through cross-modal attention. This enables more grounded vision and language representation learning. ALBEF also doesn't require bounding box annotations. The model consists of an image encode, a text encoder, and a multimodal encoder. The image-text contrastive loss helps to align the unimodal representations of an image-text pair before fusion. The image-text matching loss and a masked language modeling loss are applied to learn multimodal interactions between image and text. In addition, momentum distillation is used to generate pseudo-targets. This improves learning with noisy data.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","paper":"/paper/align-before-fuse-vision-and-language","first_author":"Junnan Li","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/align-before-fuse-vision-and-language"},"source":{"url":"https://arxiv.org/abs/2107.07651v2","title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":17,"archive_num_papers":17,"papers_newest_first":[{"paper":null,"title":"Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses","date":"2024-12-11","arxiv_id":"2412.08110","n_code_links":0,"syntology":null},{"paper":"/paper/nearest-neighbor-normalization-improves","title":"Nearest Neighbor Normalization Improves Multimodal Retrieval","date":"2024-10-31","arxiv_id":"2410.24114","n_code_links":1,"syntology":null},{"paper":null,"title":"Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM","date":"2024-04-29","arxiv_id":"2404.19128","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning from Synthetic Data for Visual Grounding","date":"2024-03-20","arxiv_id":"2403.13804","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction","date":"2024-03-16","arxiv_id":"2403.10883","n_code_links":0,"syntology":null},{"paper":null,"title":"LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival","date":"2024-03-16","arxiv_id":"2403.10887","n_code_links":0,"syntology":null},{"paper":"/paper/set-level-guidance-attack-boosting","title":"Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models","date":"2023-07-26","arxiv_id":"2307.14061","n_code_links":1,"syntology":{"ran":7,"of":12,"unverified":5,"pointer_only":0}},{"paper":"/paper/rasa-relation-and-sensitivity-aware","title":"RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search","date":"2023-05-23","arxiv_id":"2305.13653","n_code_links":1,"syntology":{"ran":3,"of":7,"unverified":4,"pointer_only":0}},{"paper":"/paper/multimodal-bias-introducing-a-framework-for","title":"MultiModal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision Language Models","date":"2023-03-16","arxiv_id":"2303.12734","n_code_links":1,"syntology":null},{"paper":"/paper/is-multi-modal-vision-supervision-beneficial","title":"Is Multimodal Vision Supervision Beneficial to Language?","date":"2023-02-10","arxiv_id":"2302.05016","n_code_links":1,"syntology":null},{"paper":"/paper/mm-shap-a-performance-agnostic-metric-for","title":"MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks","date":"2022-12-15","arxiv_id":"2212.08158","n_code_links":1,"syntology":{"ran":1,"of":10,"unverified":9,"pointer_only":0}},{"paper":null,"title":"Leveraging per Image-Token Consistency for Vision-Language Pre-training","date":"2022-11-20","arxiv_id":"2211.15398","n_code_links":0,"syntology":null},{"paper":"/paper/grit-vlp-grouped-mini-batch-sampling-for","title":"GRIT-VLP: Grouped Mini-batch Sampling for Efficient Vision and Language Pre-training","date":"2022-08-08","arxiv_id":"2208.04060","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":"/paper/chiqa-a-large-scale-image-based-real-world","title":"ChiQA: A Large Scale Image-based Real-World Question Answering Dataset for Multi-Modal Understanding","date":"2022-08-05","arxiv_id":"2208.03030","n_code_links":1,"syntology":null},{"paper":"/paper/vl-checklist-evaluating-pre-trained-vision","title":"VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations","date":"2022-07-01","arxiv_id":"2207.00221","n_code_links":1,"syntology":null},{"paper":"/paper/mixgen-a-new-multi-modal-data-augmentation","title":"MixGen: A New Multi-Modal Data Augmentation","date":"2022-06-16","arxiv_id":"2206.08358","n_code_links":1,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":"/paper/align-before-fuse-vision-and-language","title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","date":"2021-07-16","arxiv_id":"2107.07651","n_code_links":6,"syntology":{"ran":3,"of":5,"unverified":2,"pointer_only":3}}],"papers_shown":17,"tasks":[{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":7},{"task":"/task/retrieval","name":"Retrieval","papers":7},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":7},{"task":"/task/question-answering","name":"Question Answering","papers":5},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":5},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":4},{"task":"/task/image-retrieval","name":"Image Retrieval","papers":3},{"task":"/task/language-modelling","name":"Language Modelling","papers":3},{"task":"/task/visual-grounding","name":"Visual Grounding","papers":3},{"task":"/task/cross-modal-retrieval","name":"Cross-Modal Retrieval","papers":2},{"task":"/task/image-text-matching","name":"Image-text matching","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/masked-language-modeling","name":"Masked Language Modeling","papers":2},{"task":"/task/representation-learning","name":"Representation Learning","papers":2},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":2},{"task":"/task/adversarial-robustness","name":"Adversarial Robustness","papers":1},{"task":"/task/caption-generation","name":"Caption Generation","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/grounded-language-learning","name":"Grounded language learning","papers":1}],"tasks_shown":20,"n_tasks":38,"usage_by_year":[{"year":"2021","papers":1},{"year":"2022","papers":6},{"year":"2023","papers":4},{"year":"2024","papers":6}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/albef"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}