Browse State-of-the-Art › Multi-label zero-shot learning
Multi-label zero-shot learning
15 papers with code · 3 benchmarks · 2 datasets archive 2025-07-28
The goal of multi-label classification task is to predict a set of labels in an image. As an extension of zero-shot learning (ZSL), multi-label zero-shot learning (ML-ZSL) is developed to identify multiple seen and unseen labels in an image.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| NUS-WIDE (10 rows) | MKT(CLIP) | Open-Vocabulary Multi-Label Classification via Multi-Modal... | code | Syntology ran 3 of 7 samples · 4 unverified | Compare |
| Open Images V4 (8 rows) | MKT(IN-1K) | Open-Vocabulary Multi-Label Classification via Multi-Modal... | code | Syntology ran 3 of 7 samples · 4 unverified | Compare |
| ImageNet-1k to MSCOCO (1 row) | ADDS | Open Vocabulary Multi-Label Classification with Dual-Modal Decoder... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
15 shown of 15 papers with code (27 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Mar 2015 2 repositories listedAttributes act as intermediate representations that enable parameter sharing between classes, a must when training data is scarce.
-
19 Dec 2013 2 repositories listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)In other cases the semantic embedding space is established by an independent natural language processing task, and then the image transformation into that space is learned in a second stage.
-
21 Jun 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Our method achieves an absolute increase of 3.
-
10 May 2024 1 repository listedThe task of medical image recognition is notably complicated by the presence of varied and multiple pathological indications, presenting a unique challenge in multi-label classification with unseen labels.
-
5 Apr 2024 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedWe leverage the graph structure of the unlabeled data and introduce ZLaP, a method based on label propagation (LP) that utilizes geodesic distances for classification.
-
5 Jul 2022 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedSpecifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model.
-
25 Nov 2021 1 repository listed Syntology ran 2 of 5 samples · 3 unverifiedIn this paper, we introduce ML-Decoder, a new attention-based classification head.
-
20 Aug 2021 1 repository listed Syntology ran 7 of 7 samples · 0 unverifiedWe note that the best existing multi-label ZSL method takes a shared approach towards attending to region features with a common set of attention maps for all the classes.
-
19 Aug 2021 1 repository listedCLIP (Contrastive Language-Image Pre-training) is a very recent multi-modal model that jointly learns representations of images and texts.
-
12 May 2021 1 repository listedWe argue that using a single embedding vector to represent an image, as commonly practiced, is not sufficient to rank both relevant seen and unseen labels accurately.
-
27 Jan 2021 1 repository listedNevertheless, computing reliable attention maps for unseen classes during inference in a multi-label setting is still a challenge.
-
1 Jan 2021 1 repository listedWe study the problem of multi-label zero-shot recognition in which labels are in the form of human-object interactions (combinations of actions on objects), each image may contain multiple interactions and some…
-
1 Jun 2020 1 repository listedTherefore, instead of generating attentions for unseen labels which have unknown behaviors and could focus on irrelevant regions due to the lack of any training sample, we let the unseen labels select among a set of…
-
5 Jul 2019 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Audio-based music classification and tagging is typically based on categorical supervised learning with a fixed set of labels.
-
17 Nov 2017 1 repository listedIn this paper, we propose a novel deep learning architecture for multi-label zero-shot learning (ML-ZSL), which is able to predict multiple unseen class labels for each input instance.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections