{"url":"/sota/compositional-zero-shot-learning-on-mit-3","task":{"name":"Compositional Zero-Shot Learning","url":"/task/compositional-zero-shot-learning","note":null},"dataset":{"name":"MIT-States, generalized split","url":"/dataset/mit-states"},"category":"Computer Vision","categories":["Computer Vision","Methodology"],"category_note":null,"description":"**Compositional Zero-Shot Learning (CZSL)** is a computer vision task in which the goal is to recognize unseen compositions fromed from seen state and object during training. The key challenge in CZSL is the inherent entanglement between the state and object within the context of an image. Some example benchmarks for this task are MIT-states, UT-Zappos, and C-GQA. Models are usually evaluated with the Accuracy for both seen and unseen compositions, as well as their Harmonic Mean(HM).\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Heosuab](https://hellopotatoworld.tistory.com/24) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["H-Mean","Seen accuracy","Test AUC top 1","Test AUC top 2","Test AUC top 3","Unseen accuracy","Val AUC top 1","Val AUC top 2","Val AUC top 3"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"H-Mean":null,"Seen accuracy":"higher","Test AUC top 1":"higher","Test AUC top 2":"higher","Test AUC top 3":"higher","Unseen accuracy":"higher","Val AUC top 1":"higher","Val AUC top 2":"higher","Val AUC top 3":"higher"}},"counts":{"rows":2,"rows_with_code":2,"rows_with_paper_page":2,"rows_dated":2,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"CAILA","metrics":{"H-Mean":"39.9","Seen accuracy":"51.0","Test AUC top 1":"23.4","Test AUC top 2":"-","Test AUC top 3":"-","Unseen accuracy":"53.9","Val AUC top 1":"-","Val AUC top 2":"-","Val AUC top 3":"-"},"uses_additional_data":false,"paper_date":"2023-05-26","paper":"/paper/caila-concept-aware-intra-layer-adapters-for","paper_url":"https://arxiv.org/abs/2305.16681v2","paper_title":"CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning","code":"https://github.com/zhaohengz/llamp","n_code_links":2,"syntology":null},{"rank_in_archive_order":2,"model":"SymNet","metrics":{"H-Mean":"16.1","Seen accuracy":"24.4","Test AUC top 1":"3.0","Test AUC top 2":"7.6","Test AUC top 3":"12.3","Unseen accuracy":"25.2","Val AUC top 1":"4.3","Val AUC top 2":"9.8","Val AUC top 3":"14.8"},"uses_additional_data":false,"paper_date":"2020-04-01","paper":"/paper/symmetry-and-group-in-attribute-object","paper_url":"https://arxiv.org/abs/2004.00587v1","paper_title":"Symmetry and Group in Attribute-Object Compositions","code":"https://github.com/DirtyHarryLYL/SymNet","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}