{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/parts-of-speech-grounded-subspaces-in-vision","title":"Parts of Speech-Grounded Subspaces in Vision-Language Models","arxiv_id":"2305.14053","date":"2023-05-23","proceeding":null,"authors":["James Oldfield","Christos Tzelepis","Yannis Panagakis","Mihalis A. Nicolaou","Ioannis Patras"],"abstract":"Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP's joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What's more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists' painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists' styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification.","url_abs":"https://arxiv.org/abs/2305.14053v2","url_pdf":"https://arxiv.org/pdf/2305.14053v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"parts-of-speech-grounded-subspaces-in-vision","repo_url":"https://github.com/james-oldfield/pos-subspaces","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"parts-of-speech-grounded-subspaces-in-vision","repo_url":"https://github.com/james-oldfield/mmoe","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"pos","task_name":"POS"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"},{"task_slug":null,"task_name":"zero-shot-classification"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2305.14053","atlas_url":"https://app.syntology.ai/?focus=2305.14053","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.14053"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/james-oldfield/pos-subspaces","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/james-oldfield/mmoe","reach":{"status":"ok"}}],"summary":{"ran_draft_wrong":4,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"f01e8789a46a6be0","entry":"calculate_intrinstic_mean","repo":"james-oldfield/pos-subspaces","repo_kind":"official","path":"model.py","file_url":"https://github.com/james-oldfield/pos-subspaces/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f01e8789a46a6be0"}},{"code_sha256_prefix":"dd606b182570e676","entry":"exponential_map","repo":"james-oldfield/pos-subspaces","repo_kind":"official","path":"model.py","file_url":"https://github.com/james-oldfield/pos-subspaces/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dd606b182570e676"}},{"code_sha256_prefix":"ae88d1abe44272a7","entry":"logarithmic_map","repo":"james-oldfield/pos-subspaces","repo_kind":"official","path":"model.py","file_url":"https://github.com/james-oldfield/pos-subspaces/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ae88d1abe44272a7"}},{"code_sha256_prefix":"ffb9f2582dc472b1","entry":"orthogonal_projection","repo":"james-oldfield/pos-subspaces","repo_kind":"official","path":"model.py","file_url":"https://github.com/james-oldfield/pos-subspaces/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ffb9f2582dc472b1"}},{"code_sha256_prefix":"b240f14d5c12664f","entry":"Model","repo":"james-oldfield/pos-subspaces","repo_kind":"official","path":"model.py","file_url":"https://github.com/james-oldfield/pos-subspaces/blob/HEAD/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b240f14d5c12664f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}