Browse State-of-the-Art › Zero-Shot Action Recognition
Zero-Shot Action Recognition
40 papers with code · 7 benchmarks · 6 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| UCF101 (35 rows) | OTI(ViT-L/14) | Orthogonal Temporal Interpolation for Zero-Shot Video Recognition | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| HMDB51 (29 rows) | MOV (ViT-L/14) | Multimodal Open-Vocabulary Video Classification via Pre-Trained... | — | — | Compare |
| Kinetics (20 rows) | TC-CLIP | Leveraging Temporal Contextualization for Video Action Recognition | code | Syntology ran 2 of 4 samples · 2 unverified | Compare |
| Olympics (9 rows) | SPOT | Synthetic Sample Selection for Generalized Zero-Shot Learning | — | — | Compare |
| ActivityNet (5 rows) | BIKE | Bidirectional Cross-Modal Knowledge Exploration for Video... | code | — | Compare |
| Charades (4 rows) | MSQNet | Actor-agnostic Multi-label Action Recognition with Multi-modal Query | code | — | Compare |
| THUMOS' 14 (1 row) | MSQNet | Actor-agnostic Multi-label Action Recognition with Multi-modal Query | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 40 papers with code (83 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
3 Oct 2023 6 repositories listed Syntology ran 7 of 14 samples · 7 unverifiedWe thus propose VIDAL-10M with Video, Infrared, Depth, Audio and their corresponding Language, naming as VIDAL-10M.
-
31 Dec 2022 5 repositories listedIn this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the…
-
4 Jul 2022 5 repositories listedIn this study, we focus on transferring knowledge for video classification tasks.
-
27 Mar 2023 4 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedOur approach incorporates new techniques for representation learning, optimization, and augmentation, enabling EVA-CLIP to achieve superior performance compared to previous CLIP models with the same number of parameters…
-
15 Nov 2016 4 repositories listedIn this paper we argue that the key to make deep ZSL models succeed is to choose the right embedding space.
-
15 Apr 2024 2 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)To be specific, we introduce Temporal Contextualization (TC), a layer-wise temporal information infusion mechanism for videos, which 1) extracts core information from each frame, 2) connects relevant information across…
-
4 Aug 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Extensive experiments demonstrate that our approach is effective and can be generalized to different video recognition scenarios.
-
24 Mar 2022 2 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedLarge-scale pretrained image-text models have shown incredible zero-shot performance in a handful of tasks, including video ones such as action recognition and text-to-video retrieval.
-
13 Jan 2022 2 repositories listed Syntology ran 13 of 24 samples · 11 unverified · 6 pointer-only (licence)As an additional benefit, our method achieves competitive results with much shorter pre-training videos on single-modality downstream tasks, e.
-
17 Sep 2021 2 repositories listed Syntology ran 5 of 9 samples · 4 unverifiedMoreover, to handle the deficiency of label texts and make use of tremendous web data, we propose a new paradigm based on this multimodal learning framework for action recognition, which we dub "pre-train, prompt and…
-
30 Mar 2015 2 repositories listedAttributes act as intermediate representations that enable parameter sharing between classes, a must when training data is scarce.
-
30 Sep 2014 2 repositories listedImage classification has advanced significantly in recent years with the availability of large-scale image sets.
-
13 Dec 2024 1 repository listedZero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding.
-
27 Nov 2024 1 repository listedSpecifically, we obtain relative gains of 3.
-
16 Nov 2024 1 repository listedIn zero-shot skeleton-based action recognition, aligning skeleton features with the text features of action labels is essential for accurately predicting unseen actions.
-
19 Jun 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedWhile remarkable progress has been made on supervised skeleton-based action recognition, the challenge of zero-shot recognition remains relatively unexplored.
-
13 Dec 2023 1 repository listed Syntology ran 9 of 13 samples · 4 unverifiedRecent advancements in large-scale pre-training of visual-language models on paired image-text data have demonstrated impressive generalization capabilities for zero-shot tasks.
-
30 Nov 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedDue to the resource-intensive nature of training vision-language models on expansive video data, a majority of studies have centered on adapting pre-trained image-language models to the video domain.
-
29 Sep 2023 1 repository listedThe textual narratives forge connections between seen and unseen classes, overcoming the bottleneck of labeled data that has long impeded advancements in this exciting domain.
-
14 Aug 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We propose a model called OTI for ZSVR by employing orthogonal temporal interpolation and the matching loss based on VLMs.
-
20 Jul 2023 1 repository listedExisting action recognition methods are typically actor-specific due to the intrinsic topological and apparent differences among the actors.
-
13 Jul 2023 1 repository listedSpecifically, we utilize a multi-scale approach to generate video-related descriptions.
-
6 Apr 2023 1 repository listed Syntology ran 5 of 9 samples · 4 unverifiedThrough this prompting scheme, we can achieve state-of-the-art zero-shot performance on Kinetics-600, HMDB51 and UCF101 while remaining competitive in the supervised setting.
-
15 Mar 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary.
-
24 Sep 2022 1 repository listedThis work introduces a new ZSAR method based on the relationships of actions-objects and actions-descriptive sentences.
-
17 May 2022 1 repository listedOur goal in this paper is the adaptation of image-text models for long video retrieval.
-
26 Apr 2022 1 repository listedDominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast global video and text representations,…
-
29 Mar 2022 1 repository listed Syntology ran 2 of 5 samples · 3 unverified · 5 pointer-only (licence)Further, we synthesize features of unseen classes by proposing a class generator that interpolates and extrapolates the features of seen classes.
-
28 Mar 2022 1 repository listedHowever, due to the complexity of actions, it remains challenging to transfer knowledge learned from source to target action domains.
-
10 Mar 2022 1 repository listedWhile video action recognition has been an active area of research for several years, zero-shot action recognition has only recently started gaining traction.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections