{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/two-effects-one-trigger-on-the-modality-gap","title":"Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models","arxiv_id":"2404.07983","date":"2024-04-11","proceeding":null,"authors":["Simon Schrodi","David T. Hoffmann","Max Argus","Volker Fischer","Thomas Brox"],"abstract":"Contrastive vision-language models (VLMs), like CLIP, have gained popularity for their versatile applicability to various downstream tasks. Despite their successes in some tasks, like zero-shot object recognition, they perform surprisingly poor on other tasks, like attribute recognition. Previous work has attributed these challenges to the modality gap, a separation of image and text in the shared representation space, and to a bias towards objects over other factors, such as attributes. In this analysis paper, we investigate both phenomena thoroughly. We evaluated off-the-shelf VLMs and while the gap's influence on performance is typically overshadowed by other factors, we find indications that closing the gap indeed leads to improvements. Moreover, we find that, contrary to intuition, only few embedding dimensions drive the gap and that the embedding spaces are differently organized. To allow for a clean study of object bias, we introduce a definition and a corresponding measure of it. Equipped with this tool, we find that object bias does not lead to worse performance on other concepts, such as attributes per se. However, why do both phenomena, modality gap and object bias, emerge in the first place? To answer this fundamental question and uncover some of the inner workings of contrastive VLMs, we conducted experiments that allowed us to control the amount of shared information between the modalities. These experiments revealed that the driving factor behind both the modality gap and the object bias, is an information imbalance between images and captions, and unveiled an intriguing connection between the modality gap and entropy of the logits.","url_abs":"https://arxiv.org/abs/2404.07983v3","url_pdf":"https://arxiv.org/pdf/2404.07983v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"two-effects-one-trigger-on-the-modality-gap","repo_url":"https://github.com/lmb-freiburg/two-effects-one-trigger","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.07983","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.07983"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lmb-freiburg/two-effects-one-trigger","reach":null}],"summary":{"ran_draft_wrong":4,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"08e7bba3e22dab87","entry":"compute_accuracy","repo":"lmb-freiburg/two-effects-one-trigger","repo_kind":"official","path":"analysis/gap_precompute.py","file_url":"https://github.com/lmb-freiburg/two-effects-one-trigger/blob/HEAD/analysis/gap_precompute.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"08e7bba3e22dab87"}},{"code_sha256_prefix":"b2a030fd39fe6690","entry":"compute_l2m","repo":"lmb-freiburg/two-effects-one-trigger","repo_kind":"official","path":"analysis/gap_precompute.py","file_url":"https://github.com/lmb-freiburg/two-effects-one-trigger/blob/HEAD/analysis/gap_precompute.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b2a030fd39fe6690"}},{"code_sha256_prefix":"be5986bcfffaea8b","entry":"encode_clip","repo":"lmb-freiburg/two-effects-one-trigger","repo_kind":"official","path":"analysis/object_bias_precompute.py","file_url":"https://github.com/lmb-freiburg/two-effects-one-trigger/blob/HEAD/analysis/object_bias_precompute.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"be5986bcfffaea8b"}},{"code_sha256_prefix":"11e8a5d9bccd29e9","entry":"recall_at_k","repo":"lmb-freiburg/two-effects-one-trigger","repo_kind":"official","path":"analysis/gap_precompute.py","file_url":"https://github.com/lmb-freiburg/two-effects-one-trigger/blob/HEAD/analysis/gap_precompute.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"11e8a5d9bccd29e9"}},{"code_sha256_prefix":"a82147ecdc2a56ca","entry":"inter_intra_cossine_sims","repo":"lmb-freiburg/two-effects-one-trigger","repo_kind":"official","path":"analysis/object_bias_precompute.py","file_url":"https://github.com/lmb-freiburg/two-effects-one-trigger/blob/HEAD/analysis/object_bias_precompute.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a82147ecdc2a56ca"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}