{"url":"/method/blip","slug":"blip","name":"BLIP","full_name":"BLIP: Bootstrapping Language-Image Pre-training","full_name_withheld":false,"description_markdown":"Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at https://github.com/salesforce/BLIP.","description_state":"present","introduced_year":null,"introduced_by":{"title":"BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation","paper":"/paper/blip-bootstrapping-language-image-pre","first_author":"Junnan Li","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/blip-bootstrapping-language-image-pre"},"source":{"url":"https://arxiv.org/abs/2201.12086v2","title":"BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":93,"archive_num_papers":93,"papers_newest_first":[{"paper":"/paper/text-visual-semantic-constrained-ai-generated","title":"Text-Visual Semantic Constrained AI-Generated Image Quality Assessment","date":"2025-07-14","arxiv_id":"2507.10432","n_code_links":1,"syntology":null},{"paper":"/paper/vistext-mosquito-a-multimodal-dataset-and","title":"VisText-Mosquito: A Multimodal Dataset and Benchmark for AI-Based Mosquito Breeding Site Detection and Reasoning","date":"2025-06-17","arxiv_id":"2506.14629","n_code_links":1,"syntology":null},{"paper":null,"title":"Fusing Cross-modal and Uni-modal Representations: A Kronecker Product Approach","date":"2025-06-10","arxiv_id":"2506.08645","n_code_links":0,"syntology":null},{"paper":null,"title":"A Narrative Review on Large AI Models in Lung Cancer Screening, Diagnosis, and Treatment Planning","date":"2025-06-08","arxiv_id":"2506.07236","n_code_links":0,"syntology":null},{"paper":null,"title":"When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification","date":"2025-05-22","arxiv_id":"2505.16149","n_code_links":0,"syntology":null},{"paper":null,"title":"MedBLIP: Fine-tuning BLIP for Medical Image Captioning","date":"2025-05-20","arxiv_id":"2505.14726","n_code_links":0,"syntology":null},{"paper":"/paper/from-complexity-to-clarity-transforming-chest","title":"From Complexity to Clarity: Transforming Chest X-ray Reports with Chained Prompting (Student Abstract)","date":"2025-04-11","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/from-complexity-to-clarity-transforming-chest-1","title":"From Complexity to Clarity: Transforming Chest X-ray Reports with Chained Prompting (Student Abstract) Authors","date":"2025-04-11","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Learning Sparse Disentangled Representations for Multimodal Exclusion Retrieval","date":"2025-04-04","arxiv_id":"2504.03184","n_code_links":0,"syntology":null},{"paper":null,"title":"OMR-Diffusion:Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Intent Understanding","date":"2025-03-22","arxiv_id":"2503.17660","n_code_links":0,"syntology":null},{"paper":null,"title":"TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation","date":"2025-03-22","arxiv_id":"2503.17669","n_code_links":0,"syntology":null},{"paper":null,"title":"Are Large Language Models Good Data Preprocessors?","date":"2025-02-24","arxiv_id":"2502.16790","n_code_links":0,"syntology":null},{"paper":null,"title":"NanoVLMs: How small can we go and still make coherent Vision Language Models?","date":"2025-02-11","arxiv_id":"2502.07838","n_code_links":0,"syntology":null},{"paper":null,"title":"An Evaluation Framework for Product Images Background Inpainting based on Human Feedback and Product Consistency","date":"2024-12-23","arxiv_id":"2412.17504","n_code_links":0,"syntology":null},{"paper":null,"title":"Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses","date":"2024-12-11","arxiv_id":"2412.08110","n_code_links":0,"syntology":null},{"paper":"/paper/attacks-on-multimodal-models","title":"Attacks on multimodal models","date":"2024-12-02","arxiv_id":"2412.01725","n_code_links":1,"syntology":null},{"paper":null,"title":"Understanding the World's Museums through Vision-Language Reasoning","date":"2024-12-02","arxiv_id":"2412.01370","n_code_links":0,"syntology":null},{"paper":"/paper/nearest-neighbor-normalization-improves","title":"Nearest Neighbor Normalization Improves Multimodal Retrieval","date":"2024-10-31","arxiv_id":"2410.24114","n_code_links":1,"syntology":null},{"paper":null,"title":"Technical Report for Soccernet 2023 -- Dense Video Captioning","date":"2024-10-31","arxiv_id":"2411.00882","n_code_links":0,"syntology":null},{"paper":null,"title":"EfficientEQA: An Efficient Approach for Open Vocabulary Embodied Question Answering","date":"2024-10-26","arxiv_id":"2410.20263","n_code_links":0,"syntology":null},{"paper":null,"title":"Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing","date":"2024-10-23","arxiv_id":"2410.18267","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Zero-Shot Camera Trap Image Categorization","date":"2024-10-16","arxiv_id":"2410.12769","n_code_links":0,"syntology":null},{"paper":null,"title":"AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models","date":"2024-10-07","arxiv_id":"2410.05346","n_code_links":0,"syntology":null},{"paper":null,"title":"VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks","date":"2024-10-07","arxiv_id":"2410.05160","n_code_links":0,"syntology":null},{"paper":"/paper/multiclimate-multimodal-stance-detection-on","title":"MultiClimate: Multimodal Stance Detection on Climate Change Videos","date":"2024-09-26","arxiv_id":"2409.18346","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":null,"title":"NeIn: Telling What You Don't Want","date":"2024-09-09","arxiv_id":"2409.06481","n_code_links":0,"syntology":null},{"paper":"/paper/evaluation-and-comparison-of-visual-language","title":"Evaluation and Comparison of Visual Language Models for Transportation Engineering Problems","date":"2024-09-03","arxiv_id":"2409.02278","n_code_links":1,"syntology":null},{"paper":null,"title":"Medical Report Generation Is A Multi-label Classification Problem","date":"2024-08-30","arxiv_id":"2409.00250","n_code_links":0,"syntology":null},{"paper":null,"title":"Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities","date":"2024-08-13","arxiv_id":"2408.06721","n_code_links":0,"syntology":null},{"paper":null,"title":"MOSAIC: Multimodal Multistakeholder-aware Visual Art Recommendation","date":"2024-07-31","arxiv_id":"2407.21758","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/retrieval","name":"Retrieval","papers":22},{"task":"/task/image-captioning","name":"Image Captioning","papers":19},{"task":"/task/question-answering","name":"Question Answering","papers":14},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":13},{"task":"/task/language-modelling","name":"Language Modelling","papers":12},{"task":"/task/image-generation","name":"Image Generation","papers":9},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":9},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":9},{"task":"/task/decoder","name":"Decoder","papers":7},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":7},{"task":"/task/language-modeling","name":"Language Modeling","papers":7},{"task":"/task/large-language-model","name":"Large Language Model","papers":6},{"task":"/task/attribute","name":"Attribute","papers":5},{"task":"/task/cross-modal-retrieval","name":"Cross-Modal Retrieval","papers":5},{"task":"/task/image-retrieval","name":"Image Retrieval","papers":5},{"task":"/task/object-detection","name":"Object Detection","papers":5},{"task":"/task/object-detection-1","name":"object-detection","papers":5},{"task":"/task/diversity","name":"Diversity","papers":4},{"task":"/task/image-classification","name":"Image Classification","papers":4},{"task":"/task/image-text-matching","name":"Image-text matching","papers":4}],"tasks_shown":20,"n_tasks":123,"usage_by_year":[{"year":"2022","papers":7},{"year":"2023","papers":31},{"year":"2024","papers":42},{"year":"2025","papers":13}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/blip"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}