{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/training-vision-language-models-with-less","title":"Training Vision-Language Models with Less Bimodal Supervision","arxiv_id":"2211.00262","date":"2022-11-01","proceeding":null,"authors":["Elad Segal","Ben Bogin","Jonathan Berant"],"abstract":"Standard practice in pretraining multimodal models, such as vision-language models, is to rely on pairs of aligned inputs from both modalities, for example, aligned image-text pairs. However, such pairs can be difficult to obtain in low-resource settings and for some modality pairs (e.g., structured tables and images). In this work, we investigate the extent to which we can reduce the reliance on such parallel data, which we term \\emph{bimodal supervision}, and use models that are pretrained on each modality independently. We experiment with a high-performing vision-language model, and analyze the effect of bimodal supervision on three vision-language tasks. We find that on simpler tasks, such as VQAv2 and GQA, one can eliminate bimodal supervision completely, suffering only a minor loss in performance. Conversely, for NLVR2, which requires more complex reasoning, training without bimodal supervision leads to random performance. Nevertheless, using only 5\\% of the bimodal data (142K images along with their captions), or leveraging weak supervision in the form of a list of machine-generated labels for each image, leads to only a moderate degradation compared to using 3M image-text pairs: 74\\%$\\rightarrow$$\\sim$70\\%. Our code is available at https://github.com/eladsegal/less-bimodal-sup.","url_abs":"https://arxiv.org/abs/2211.00262v1","url_pdf":"https://arxiv.org/pdf/2211.00262v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"training-vision-language-models-with-less","repo_url":"https://github.com/eladsegal/less-bimodal-sup","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2211.00262","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2211.00262"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eladsegal/less-bimodal-sup","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"cb025115712dfc97","entry":"force_eval","repo":"eladsegal/less-bimodal-sup","repo_kind":"official","path":"src/callbacks/freezer.py","file_url":"https://github.com/eladsegal/less-bimodal-sup/blob/HEAD/src/callbacks/freezer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cb025115712dfc97"}},{"code_sha256_prefix":"d715edbedc06dfd2","entry":"identity_func","repo":"eladsegal/less-bimodal-sup","repo_kind":"official","path":"src/evaluators/hf_evaluator.py","file_url":"https://github.com/eladsegal/less-bimodal-sup/blob/HEAD/src/evaluators/hf_evaluator.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d715edbedc06dfd2"}},{"code_sha256_prefix":"225dafec388f7a2f","entry":"unflatten","repo":"eladsegal/less-bimodal-sup","repo_kind":"official","path":"utils/jsonnet.py","file_url":"https://github.com/eladsegal/less-bimodal-sup/blob/HEAD/utils/jsonnet.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"225dafec388f7a2f"}},{"code_sha256_prefix":"d467682ffde173a4","entry":"with_fallback","repo":"eladsegal/less-bimodal-sup","repo_kind":"official","path":"utils/jsonnet.py","file_url":"https://github.com/eladsegal/less-bimodal-sup/blob/HEAD/utils/jsonnet.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d467682ffde173a4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}