Methods › Computer Vision › Vision and Language Pre-Trained Models › BLIP
BLIP: Bootstrapping Language-Image Pre-training
BLIP
Introduced by Junnan Li et al. in BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at https://github.com/salesforce/BLIP.
Papers archive 2025-07-28
30 shown of 93, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Text-Visual Semantic Constrained AI-Generated Image Quality Assessment 14 Jul 2025 · 1 repository · arXiv:2507.10432
-
VisText-Mosquito: A Multimodal Dataset and Benchmark for AI-Based Mosquito Breeding Site Detection and Reasoning 17 Jun 2025 · 1 repository · arXiv:2506.14629
-
Fusing Cross-modal and Uni-modal Representations: A Kronecker Product Approach 10 Jun 2025 · 0 repositories · arXiv:2506.08645
-
A Narrative Review on Large AI Models in Lung Cancer Screening, Diagnosis, and Treatment Planning 8 Jun 2025 · 0 repositories · arXiv:2506.07236
-
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification 22 May 2025 · 0 repositories · arXiv:2505.16149
-
MedBLIP: Fine-tuning BLIP for Medical Image Captioning 20 May 2025 · 0 repositories · arXiv:2505.14726
-
From Complexity to Clarity: Transforming Chest X-ray Reports with Chained Prompting (Student Abstract) 11 Apr 2025 · 1 repository
-
From Complexity to Clarity: Transforming Chest X-ray Reports with Chained Prompting (Student Abstract) Authors 11 Apr 2025 · 1 repository
-
Learning Sparse Disentangled Representations for Multimodal Exclusion Retrieval 4 Apr 2025 · 0 repositories · arXiv:2504.03184
-
OMR-Diffusion:Optimizing Multi-Round Enhanced Training in Diffusion Models for Improved Intent Understanding 22 Mar 2025 · 0 repositories · arXiv:2503.17660
-
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation 22 Mar 2025 · 0 repositories · arXiv:2503.17669
-
Are Large Language Models Good Data Preprocessors? 24 Feb 2025 · 0 repositories · arXiv:2502.16790
-
NanoVLMs: How small can we go and still make coherent Vision Language Models? 11 Feb 2025 · 0 repositories · arXiv:2502.07838
-
An Evaluation Framework for Product Images Background Inpainting based on Human Feedback and Product Consistency 23 Dec 2024 · 0 repositories · arXiv:2412.17504
-
Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses 11 Dec 2024 · 0 repositories · arXiv:2412.08110
-
Attacks on multimodal models 2 Dec 2024 · 1 repository · arXiv:2412.01725
-
Understanding the World's Museums through Vision-Language Reasoning 2 Dec 2024 · 0 repositories · arXiv:2412.01370
-
Nearest Neighbor Normalization Improves Multimodal Retrieval 31 Oct 2024 · 1 repository · arXiv:2410.24114
-
Technical Report for Soccernet 2023 -- Dense Video Captioning 31 Oct 2024 · 0 repositories · arXiv:2411.00882
-
EfficientEQA: An Efficient Approach for Open Vocabulary Embodied Question Answering 26 Oct 2024 · 0 repositories · arXiv:2410.20263
-
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing 23 Oct 2024 · 0 repositories · arXiv:2410.18267
-
Towards Zero-Shot Camera Trap Image Categorization 16 Oct 2024 · 0 repositories · arXiv:2410.12769
-
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models 7 Oct 2024 · 0 repositories · arXiv:2410.05346
-
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 7 Oct 2024 · 0 repositories · arXiv:2410.05160
-
MultiClimate: Multimodal Stance Detection on Climate Change Videos 26 Sep 2024 · 1 repository · arXiv:2409.18346Syntology ran 1 of 1 samples · 0 unverified
-
NeIn: Telling What You Don't Want 9 Sep 2024 · 0 repositories · arXiv:2409.06481
-
Evaluation and Comparison of Visual Language Models for Transportation Engineering Problems 3 Sep 2024 · 1 repository · arXiv:2409.02278
-
Medical Report Generation Is A Multi-label Classification Problem 30 Aug 2024 · 0 repositories · arXiv:2409.00250
-
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities 13 Aug 2024 · 0 repositories · arXiv:2408.06721
-
MOSAIC: Multimodal Multistakeholder-aware Visual Art Recommendation 31 Jul 2024 · 0 repositories · arXiv:2407.21758
Tasks archive 2025-07-28
20 shown of 123 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections