{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bootstrapping-vision-language-learning-with-1","title":"Bootstrapping Vision-Language Learning with Decoupled Language Pre-training","arxiv_id":"2307.07063","date":"2023-07-13","proceeding":"NeurIPS 2023 11","authors":["Yiren Jian","Chongyang Gao","Soroush Vosoughi"],"abstract":"We present a novel methodology aimed at optimizing the application of frozen large language models (LLMs) for resource-intensive vision-language (VL) pre-training. The current paradigm uses visual features as prompts to guide language models, with a focus on determining the most relevant visual features for corresponding text. Our approach diverges by concentrating on the language component, specifically identifying the optimal prompts to align with visual features. We introduce the Prompt-Transformer (P-Former), a model that predicts these ideal prompts, which is trained exclusively on linguistic data, bypassing the need for image-text pairings. This strategy subtly bifurcates the end-to-end VL training process into an additional, separate stage. Our experiments reveal that our framework significantly enhances the performance of a robust image-to-text baseline (BLIP-2), and effectively narrows the performance gap between models trained with either 4M or 129M image-text pairs. Importantly, our framework is modality-agnostic and flexible in terms of architectural design, as validated by its successful application in a video learning task using varied base modules. The code will be made available at https://github.com/yiren-jian/BLIText.","url_abs":"https://arxiv.org/abs/2307.07063v4","url_pdf":"https://arxiv.org/pdf/2307.07063v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bootstrapping-vision-language-learning-with-1","repo_url":"https://github.com/yiren-jian/blitext","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"image-to-text","task_name":"Image to text"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"base","method_name":"BASE"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.07063","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.07063"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/yiren-jian/BLIText","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yiren-jian/blitext","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"summary":{"ran":7,"unverified":2},"by_repo_kind":{"official":{"samples":9,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"56f02812d66a0d11","entry":"generate_caption","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/caption.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/caption.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"56f02812d66a0d11"}},{"code_sha256_prefix":"92f2e16ee0d3a24d","entry":"get_concat_v","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"92f2e16ee0d3a24d"}},{"code_sha256_prefix":"13a43a7d815937e6","entry":"read_img","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/calculate_coco_features.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/calculate_coco_features.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"13a43a7d815937e6"}},{"code_sha256_prefix":"926bb38f979a61ef","entry":"resize_img","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/utils.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"926bb38f979a61ef"}},{"code_sha256_prefix":"f809f2ecf34c1c99","entry":"resize_img_w","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"f809f2ecf34c1c99"}},{"code_sha256_prefix":"3d4bbb5ca53fd24a","entry":"sample_dataset","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/dataset_browser.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/dataset_browser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3d4bbb5ca53fd24a"}},{"code_sha256_prefix":"cb33571427334815","entry":"tile","repo":"yiren-jian/BLIText","repo_kind":"official","path":"lavis/models/base_model.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/lavis/models/base_model.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"cb33571427334815"}},{"code_sha256_prefix":"0ec9fc2025c16f65","entry":"all_gather_with_grad","repo":"yiren-jian/BLIText","repo_kind":"official","path":"lavis/models/base_model.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/lavis/models/base_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"0ec9fc2025c16f65"}},{"code_sha256_prefix":"3e12ea8c09674aa8","entry":"compute_gradcam_batch","repo":"yiren-jian/BLIText","repo_kind":"official","path":"app/multimodal_search.py","file_url":"https://github.com/yiren-jian/BLIText/blob/HEAD/app/multimodal_search.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3e12ea8c09674aa8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}