{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/do-vision-language-pretrained-models-learn","title":"Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?","arxiv_id":"2203.17271","date":"2022-03-31","proceeding":null,"authors":["Tian Yun","Usha Bhalla","Ellie Pavlick","Chen Sun"],"abstract":"Vision-language (VL) pretrained models have achieved impressive performance on multimodal reasoning and zero-shot recognition tasks. Many of these VL models are pretrained on unlabeled image and caption pairs from the internet. In this paper, we study whether representations of primitive concepts--such as colors, shapes, or the attributes of object parts--emerge automatically within these pretrained VL models. We propose a two-step framework, Compositional Concept Mapping (CompMap), to investigate this. CompMap first asks a VL model to generate primitive concept activations with text prompts, and then learns to construct a composition model that maps the primitive concept activations (e.g. the likelihood of black tail or red wing) to composite concepts (e.g. a red-winged blackbird). We show that a composition model can be reliably learn from ground truth primitive concepts. We thus hypothesize that if primitive concepts indeed emerge in a VL pretrained model, its primitive concept activations can be used to learn a composition model similar to the one designed by experts. We propose a quantitative metric to measure the degree of similarity, and refer to the metric as the interpretability metric. We also measure the classification accuracy when using the primitive concept activations and the learned composition model to predict the composite concepts, and refer to it as the usefulness metric. Our study reveals that state-of-the-art VL pretrained models learn primitive concepts that are highly useful for fine-grained visual recognition on the CUB dataset, and compositional generalization tasks on the MIT-States dataset. However, we observe that the learned composition models have low interpretability in our qualitative analyses. Our results reveal the limitations of existing VL models, and the necessity of pretraining objectives that encourage the acquisition of primitive concepts.","url_abs":"https://arxiv.org/abs/2203.17271v3","url_pdf":"https://arxiv.org/pdf/2203.17271v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"do-vision-language-pretrained-models-learn","repo_url":"https://github.com/tttyuntian/vlm_primitive_concepts","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"fine-grained-visual-recognition","task_name":"Fine-Grained Visual Recognition"},{"task_slug":"multimodal-reasoning","task_name":"Multimodal Reasoning"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2203.17271","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2203.17271"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tttyuntian/vlm_primitive_concepts","reach":null}],"summary":{"ran":3,"ran_honours":1},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8c9945f5c6ad1124","entry":"get_logger","repo":"tttyuntian/vlm_primitive_concepts","repo_kind":"official","path":"vlm_concept/cub_200_2011/precompute_features.py","file_url":"https://github.com/tttyuntian/vlm_primitive_concepts/blob/HEAD/vlm_concept/cub_200_2011/precompute_features.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8c9945f5c6ad1124"}},{"code_sha256_prefix":"93bd857d933286f9","entry":"get_logger","repo":"tttyuntian/vlm_primitive_concepts","repo_kind":"official","path":"vlm_concept/mit_states/train_retrieval_model.py","file_url":"https://github.com/tttyuntian/vlm_primitive_concepts/blob/HEAD/vlm_concept/mit_states/train_retrieval_model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"93bd857d933286f9"}},{"code_sha256_prefix":"102caf3e5c00141a","entry":"get_output_path","repo":"tttyuntian/vlm_primitive_concepts","repo_kind":"official","path":"vlm_concept/mit_states/train_retrieval_model.py","file_url":"https://github.com/tttyuntian/vlm_primitive_concepts/blob/HEAD/vlm_concept/mit_states/train_retrieval_model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"102caf3e5c00141a"}},{"code_sha256_prefix":"43205e727681d033","entry":"get_seen_unseen_indices","repo":"tttyuntian/vlm_primitive_concepts","repo_kind":"official","path":"vlm_concept/mit_states/train_retrieval_model.py","file_url":"https://github.com/tttyuntian/vlm_primitive_concepts/blob/HEAD/vlm_concept/mit_states/train_retrieval_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"43205e727681d033"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}