Browse State-of-the-Art › Few Shot Action Recognition
Few Shot Action Recognition
31 papers with code · 5 benchmarks · 6 datasets archive 2025-07-28
Few-shot (FS) action recognition is a challenging com- puter vision problem, where the task is to classify an unlabelled query video into one of the action categories in the support set having limited samples per action class.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Kinetics-100 (8 rows) | Name Tuning | Few-Shot Classification of Interactive Activities of Daily Living... | code | — | Compare |
| HMDB51 (7 rows) | STRM | Spatio-temporal Relation Modeling for Few-shot Action Recognition | code | — | Compare |
| UCF101 (7 rows) | STRM | Spatio-temporal Relation Modeling for Few-shot Action Recognition | code | — | Compare |
| Something-Something-100 (5 rows) | HyRSM | Hybrid Relation Guided Set Matching for Few-shot Action Recognition | code | — | Compare |
| MOMA-LRG (4 rows) | Name Tuning | Few-Shot Classification of Interactive Activities of Daily Living... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 31 papers with code (76 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
15 Jan 2021 2 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We propose a novel approach to few-shot action recognition, finding temporally-corresponding frame tuples between the query and videos in the support set.
-
15 Jan 2021 2 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We propose a novel approach to few-shot action recognition, finding temporally-corresponding frame tuples between the query and videos in the support set.
-
15 Dec 2019 2 repositories listedNext, by decomposing and learning the temporal changes in visual relationships that result in an action, we demonstrate the utility of a hierarchical event decomposition by enabling few-shot action recognition,…
-
15 Dec 2019 2 repositories listedNext, by decomposing and learning the temporal changes in visual relationships that result in an action, we demonstrate the utility of a hierarchical event decomposition by enabling few-shot action recognition,…
-
9 May 2025 1 repository listedLarge-scale pre-trained models have achieved remarkable success in language and image tasks, leading an increasing number of studies to explore the application of pre-trained image models, such as CLIP, in the domain of…
-
9 May 2025 1 repository listedLarge-scale pre-trained models have achieved remarkable success in language and image tasks, leading an increasing number of studies to explore the application of pre-trained image models, such as CLIP, in the domain of…
-
8 Apr 2025 1 repository listed Syntology ran 4 of 8 samples · 4 unverified · 8 pointer-only (licence)Few-Shot Action Recognition (FSAR) aims to train a model with only a few labeled video instances.
-
8 Apr 2025 1 repository listed Syntology ran 4 of 8 samples · 4 unverified · 8 pointer-only (licence)Few-Shot Action Recognition (FSAR) aims to train a model with only a few labeled video instances.
-
28 Nov 2024 1 repository listedTo effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware…
-
28 Nov 2024 1 repository listedTo effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware…
-
23 Jul 2024 1 repository listedHigh frame-rate (HFR) videos of action recognition improve fine-grained expression while reducing the spatio-temporal relation and motion information density.
-
23 Jul 2024 1 repository listedHigh frame-rate (HFR) videos of action recognition improve fine-grained expression while reducing the spatio-temporal relation and motion information density.
-
3 Jun 2024 1 repository listedUnderstanding Activities of Daily Living (ADLs) is a crucial step for different applications including assistive robots, smart homes, and healthcare.
-
3 Dec 2023 1 repository listedIn particular, we devise the anisotropic Deformable Spatio-Temporal Attention module as the core component of D²ST-Adapter, which can be tailored with anisotropic sampling densities along spatial and temporal domains to…
-
3 Dec 2023 1 repository listedIn particular, we devise the anisotropic Deformable Spatio-Temporal Attention module as the core component of D²ST-Adapter, which can be tailored with anisotropic sampling densities along spatial and temporal domains to…
-
7 Sep 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)To address this issue, in this work, we propose a novel cross-domain few-shot video action recognition method that leverages self-supervised learning and curriculum learning to balance the information from the source…
-
7 Sep 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)To address this issue, in this work, we propose a novel cross-domain few-shot video action recognition method that leverages self-supervised learning and curriculum learning to balance the information from the source…
-
18 Aug 2023 1 repository listedClass prototype construction and matching are core aspects of few-shot action recognition.
-
18 Aug 2023 1 repository listedClass prototype construction and matching are core aspects of few-shot action recognition.
-
5 Jul 2023 1 repository listedThe second module (MLT) focuses on the Multiple-level feature of the support prototype and query sample to mine more information for the alignment, which operates on different level features.
-
5 Jul 2023 1 repository listedThe second module (MLT) focuses on the Multiple-level feature of the support prototype and query sample to mine more information for the alignment, which operates on different level features.
-
3 Apr 2023 1 repository listedTo address these issues, we develop a Motion-augmented Long-short Contrastive Learning (MoLo) method that contains two crucial components, including a long-short contrastive objective and a motion autodecoder.
-
3 Apr 2023 1 repository listedTo address these issues, we develop a Motion-augmented Long-short Contrastive Learning (MoLo) method that contains two crucial components, including a long-short contrastive objective and a motion autodecoder.
-
15 Mar 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary.
-
15 Mar 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)We adapt a VL model for zero-shot and few-shot action recognition using a collection of unlabeled videos and an unpaired action dictionary.
-
6 Mar 2023 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedLearning from large-scale contrastive language-image pre-training like CLIP has shown remarkable success in a wide range of downstream tasks recently, but it is still under-explored on the challenging few-shot action…
-
6 Mar 2023 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedLearning from large-scale contrastive language-image pre-training like CLIP has shown remarkable success in a wide range of downstream tasks recently, but it is still under-explored on the challenging few-shot action…
-
9 Jan 2023 1 repository listedTo be specific, HyRSM++ consists of two key components, a hybrid relation module and a temporal set matching metric.
-
9 Jan 2023 1 repository listedTo be specific, HyRSM++ consists of two key components, a hybrid relation module and a temporal set matching metric.
-
28 Dec 2022 1 repository listed Syntology ran 3 of 5 samples · 2 unverifiedFor long videos, given a paragraph of description where the sentences describe different segments of the video, by matching all sentence-clip pairs, the paragraph and the full video are aligned implicitly.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections