{"url":"/task/compositional-zero-shot-learning","name":"Compositional Zero-Shot Learning","slug":"compositional-zero-shot-learning","description_markdown":"**Compositional Zero-Shot Learning (CZSL)** is a computer vision task in which the goal is to recognize unseen compositions fromed from seen state and object during training. The key challenge in CZSL is the inherent entanglement between the state and object within the context of an image. Some example benchmarks for this task are MIT-states, UT-Zappos, and C-GQA. Models are usually evaluated with the Accuracy for both seen and unseen compositions, as well as their Harmonic Mean(HM).\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Heosuab](https://hellopotatoworld.tistory.com/24) )</span>","categories":[{"name":"Computer Vision","url":"/area/computer-vision"},{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":65,"papers_with_code":31,"benchmarks":4,"benchmark_tables_in_archive":4,"benchmark_tables_shown":4,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":6,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/compositional-zero-shot-learning-on-mit-2","slug":"compositional-zero-shot-learning-on-mit-2","dataset":"MIT-States","dataset_url":"/dataset/mit-states","rows_in_archive":2,"metrics":["AUC","Attribute accuracy","Object accuracy","Seen accuracy","Top-1 accuracy %","Top-2 accuracy %","Top-3 accuracy %","Unseen accuracy","best HM"],"first_row_in_archive_order":{"model":"CANet","paper_title":"Learning Conditional Attributes for Compositional Zero-Shot Learning","paper_url":"/paper/learning-conditional-attributes-for-1","paper_date":"2023-05-29","arxiv_id":"2305.17940","code_links":[{"title":"wqshmzh/canet-czsl","url":"https://github.com/wqshmzh/canet-czsl"}],"syntology":{"n":9,"n_ran":6,"n_unverified":3,"n_pointer_only":9}}},{"leaderboard":"/sota/compositional-zero-shot-learning-on-mit-3","slug":"compositional-zero-shot-learning-on-mit-3","dataset":"MIT-States, generalized split","dataset_url":"/dataset/mit-states","rows_in_archive":2,"metrics":["H-Mean","Seen accuracy","Test AUC top 1","Test AUC top 2","Test AUC top 3","Unseen accuracy","Val AUC top 1","Val AUC top 2","Val AUC top 3"],"first_row_in_archive_order":{"model":"CAILA","paper_title":"CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning","paper_url":"/paper/caila-concept-aware-intra-layer-adapters-for","paper_date":"2023-05-26","arxiv_id":"2305.16681","code_links":[{"title":"zhaohengz/llamp","url":"https://github.com/zhaohengz/llamp"},{"title":"zhaohengz/caila","url":"https://github.com/zhaohengz/caila"}],"syntology":null}},{"leaderboard":"/sota/compositional-zero-shot-learning-on-ut","slug":"compositional-zero-shot-learning-on-ut","dataset":"UT Zappos50K","dataset_url":"/dataset/ut-zappos50k","rows_in_archive":1,"metrics":["AUC","Attribute accuracy","Object accuracy","Seen accuracy","Unseen accuracy","best HM"],"first_row_in_archive_order":{"model":"CANet","paper_title":"Learning Conditional Attributes for Compositional Zero-Shot Learning","paper_url":"/paper/learning-conditional-attributes-for-1","paper_date":"2023-05-29","arxiv_id":"2305.17940","code_links":[{"title":"wqshmzh/canet-czsl","url":"https://github.com/wqshmzh/canet-czsl"}],"syntology":{"n":9,"n_ran":6,"n_unverified":3,"n_pointer_only":9}}},{"leaderboard":"/sota/compositional-zero-shot-learning-on-ut-zappos-1","slug":"compositional-zero-shot-learning-on-ut-zappos-1","dataset":"UT-Zappos","dataset_url":null,"rows_in_archive":1,"metrics":["Top-1 accuracy %","Top-2 accuracy %","Top-3 accuracy %"],"first_row_in_archive_order":{"model":"SymNet","paper_title":"Symmetry and Group in Attribute-Object Compositions","paper_url":"/paper/symmetry-and-group-in-attribute-object","paper_date":"2020-04-01","arxiv_id":"2004.00587","code_links":[{"title":"DirtyHarryLYL/SymNet","url":"https://github.com/DirtyHarryLYL/SymNet"}],"syntology":null}}],"datasets":[{"url":"/dataset/mit-states","name":"MIT-States","full_name":"","num_papers_in_archive":91},{"url":"/dataset/c-gqa","name":"C-GQA","full_name":"Compositional GQA","num_papers_in_archive":42},{"url":"/dataset/ut-zappos50k","name":"UT Zappos50K","full_name":"","num_papers_in_archive":32},{"url":"/dataset/reascan","name":"ReaSCAN","full_name":"ReaSCAN: Compositional Reasoning in Language Grounding","num_papers_in_archive":10},{"url":"/dataset/ao-clevr","name":"AO-CLEVr","full_name":null,"num_papers_in_archive":6},{"url":"/dataset/ut-zappos50k-1","name":"UT-Zappos50K","full_name":"","num_papers_in_archive":0}],"subtasks":[],"parent_tasks":[{"url":"/task/zero-shot-learning","name":"Zero-Shot Learning"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":31,"tagged_in_all":65,"items":[{"url":"/paper/caila-concept-aware-intra-layer-adapters-for","title":"CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning","date":"2023-05-26","arxiv_id":"2305.16681","repositories_listed":2,"syntology":null},{"url":"/paper/learning-graph-embeddings-for-open-world","title":"Learning Graph Embeddings for Open World Compositional Zero-Shot Learning","date":"2021-05-03","arxiv_id":"2105.01017","repositories_listed":2,"syntology":null},{"url":"/paper/open-world-compositional-zero-shot-learning","title":"Open World Compositional Zero-Shot Learning","date":"2021-01-29","arxiv_id":"2101.12609","repositories_listed":2,"syntology":null},{"url":"/paper/msci-addressing-clip-s-inherent-limitations","title":"MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning","date":"2025-05-15","arxiv_id":"2505.10289","repositories_listed":1,"syntology":{"n":28,"n_ran":23,"n_unverified":5,"n_pointer_only":28}},{"url":"/paper/learning-clustering-based-prototypes-for","title":"Learning Clustering-based Prototypes for Compositional Zero-shot Learning","date":"2025-02-10","arxiv_id":"2502.06501","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/unified-framework-for-open-world","title":"Unified Framework for Open-World Compositional Zero-shot Learning","date":"2024-12-05","arxiv_id":"2412.04083","repositories_listed":1,"syntology":null},{"url":"/paper/leveraging-mllm-embeddings-and-attribute","title":"Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning","date":"2024-11-18","arxiv_id":"2411.12584","repositories_listed":1,"syntology":null},{"url":"/paper/attention-based-simple-primitives-for-open","title":"Attention Based Simple Primitives for Open World Compositional Zero-Shot Learning","date":"2024-07-18","arxiv_id":"2407.13715","repositories_listed":1,"syntology":null},{"url":"/paper/contextual-interaction-via-primitive-based","title":"Contextual Interaction via Primitive-based Adversarial Training For Compositional Zero-shot Learning","date":"2024-06-21","arxiv_id":"2406.14962","repositories_listed":1,"syntology":null},{"url":"/paper/synthesize-diagnose-and-optimize-towards-fine","title":"Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding","date":"2023-11-30","arxiv_id":"2312.00081","repositories_listed":1,"syntology":{"n":11,"n_ran":4,"n_unverified":7,"n_pointer_only":11}},{"url":"/paper/gipcol-graph-injected-soft-prompting-for","title":"GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning","date":"2023-11-09","arxiv_id":"2311.05729","repositories_listed":1,"syntology":{"n":12,"n_ran":7,"n_unverified":5,"n_pointer_only":12}},{"url":"/paper/hierarchical-visual-primitive-experts-for","title":"Hierarchical Visual Primitive Experts for Compositional Zero-Shot Learning","date":"2023-08-08","arxiv_id":"2308.04016","repositories_listed":1,"syntology":{"n":12,"n_ran":11,"n_unverified":1,"n_pointer_only":12}},{"url":"/paper/learning-conditional-attributes-for-1","title":"Learning Conditional Attributes for Compositional Zero-Shot Learning","date":"2023-05-29","arxiv_id":"2305.17940","repositories_listed":1,"syntology":{"n":9,"n_ran":6,"n_unverified":3,"n_pointer_only":9}},{"url":"/paper/prompting-language-informed-distribution-for","title":"Prompting Language-Informed Distribution for Compositional Zero-Shot Learning","date":"2023-05-23","arxiv_id":"2305.14428","repositories_listed":1,"syntology":{"n":8,"n_ran":3,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/learning-attention-as-disentangler-for","title":"Learning Attention as Disentangler for Compositional Zero-shot Learning","date":"2023-03-27","arxiv_id":"2303.15111","repositories_listed":1,"syntology":{"n":11,"n_ran":5,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/troika-multi-path-cross-modal-traction-for","title":"Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning","date":"2023-03-27","arxiv_id":"2303.15230","repositories_listed":1,"syntology":null},{"url":"/paper/decomposed-soft-prompt-guided-fusion","title":"Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot Learning","date":"2022-11-19","arxiv_id":"2211.10681","repositories_listed":1,"syntology":{"n":4,"n_ran":1,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/reference-limited-compositional-zero-shot","title":"Reference-Limited Compositional Zero-Shot Learning","date":"2022-08-22","arxiv_id":"2208.10046","repositories_listed":1,"syntology":null},{"url":"/paper/siamese-contrastive-embedding-network-for-1","title":"Siamese Contrastive Embedding Network for Compositional Zero-Shot Learning","date":"2022-06-29","arxiv_id":"2206.14475","repositories_listed":1,"syntology":{"n":7,"n_ran":3,"n_unverified":4,"n_pointer_only":7}},{"url":"/paper/learning-invariant-visual-representations-for","title":"Learning Invariant Visual Representations for Compositional Zero-Shot Learning","date":"2022-06-01","arxiv_id":"2206.00415","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/disentangling-visual-embeddings-for","title":"Disentangling Visual Embeddings for Attributes and Objects","date":"2022-05-17","arxiv_id":"2205.08536","repositories_listed":1,"syntology":null},{"url":"/paper/kg-sp-knowledge-guided-simple-primitives-for","title":"KG-SP: Knowledge Guided Simple Primitives for Open World Compositional Zero-Shot Learning","date":"2022-05-13","arxiv_id":"2205.06784","repositories_listed":1,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/learning-to-compose-soft-prompts-for","title":"Learning to Compose Soft Prompts for Compositional Zero-Shot Learning","date":"2022-04-07","arxiv_id":"2204.03574","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/batchformer-learning-to-explore-sample","title":"BatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning","date":"2022-03-03","arxiv_id":"2203.01522","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/learning-single-multi-attribute-of-object","title":"Learning Single/Multi-Attribute of Object with Symmetry and Group","date":"2021-10-09","arxiv_id":"2110.04603","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/relation-aware-compositional-zero-shot","title":"Relation-aware Compositional Zero-shot Learning for Attribute-Object Pair Recognition","date":"2021-08-10","arxiv_id":"2108.04603","repositories_listed":1,"syntology":null},{"url":"/paper/independent-prototype-propagation-for-zero","title":"Independent Prototype Propagation for Zero-Shot Compositionality","date":"2021-06-01","arxiv_id":"2106.00305","repositories_listed":1,"syntology":null},{"url":"/paper/learning-graph-embeddings-for-compositional","title":"Learning Graph Embeddings for Compositional Zero-shot Learning","date":"2021-02-03","arxiv_id":"2102.01987","repositories_listed":1,"syntology":{"n":11,"n_ran":6,"n_unverified":5,"n_pointer_only":11}},{"url":"/paper/a-causal-view-of-compositional-zero-shot","title":"A causal view of compositional zero-shot recognition","date":"2020-06-25","arxiv_id":"2006.14610","repositories_listed":1,"syntology":null},{"url":"/paper/symmetry-and-group-in-attribute-object","title":"Symmetry and Group in Attribute-Object Compositions","date":"2020-04-01","arxiv_id":"2004.00587","repositories_listed":1,"syntology":null}],"syntology_records":16,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}