Browse State-of-the-Art › Action Recognition
Action Recognition
1,058 papers with code · 56 benchmarks · 115 datasets archive 2025-07-28
Action Recognition is a computer vision task that involves recognizing human actions in videos or images. The goal is to classify and categorize the actions being performed in the video or image into a predefined set of action classes.
In the video domain, it is an open question whether training an action classification network on a sufficiently large dataset, will give a similar boost in performance when applied to a different temporal task or dataset. The challenges of building video datasets has meant that most popular benchmarks for action recognition are small, having on the order of 10k videos.
Please note some benchmarks may be located in the Action Classification or Video Classification tasks, e.g. Kinetics-400.
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
56 leaderboard tables shown for this task, 56 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 56 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
115 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 115 until expanded.
Subtasks archive 2025-07-28
15 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 1,058 papers with code (2,759 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 May 2019 144 repositories listed Syntology ran 171 of 302 samples · 131 unverified · 112 pointer-only (licence)Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
22 May 2017 34 repositories listed Syntology ran 16 of 28 samples · 12 unverified · 7 pointer-only (licence)The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale…
-
21 Nov 2017 32 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time.
-
2 Dec 2014 29 repositories listedWe propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset.
-
27 Nov 2017 26 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 3 pointer-only (licence)The purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels.
-
23 Jan 2018 24 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedDynamics of human body skeletons convey significant information for human action recognition.
-
30 Nov 2017 24 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition.
-
30 Oct 2017 24 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 3 pointer-only (licence)Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems.
-
20 Nov 2014 24 repositories listed Syntology ran 12 of 32 samples · 20 unverified · 27 pointer-only (licence)We propose a novel paradigm for evaluating image descriptions that uses human consensus.
-
2 Aug 2016 22 repositories listed Syntology ran 2 of 24 samples · 22 unverified · 3 pointer-only (licence)The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network.
-
9 Feb 2021 16 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
-
24 Jun 2021 15 repositories listed Syntology ran 7 of 32 samples · 25 unverifiedThe vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks.
-
23 Jul 2019 15 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedTo address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending…
-
10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present SlowFast networks for video recognition.
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
16 Feb 2015 12 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)We further evaluate the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.
-
8 May 2017 11 repositories listedFurthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
-
29 Mar 2021 10 repositories listed Syntology ran 14 of 21 samples · 7 unverified · 1 pointer-only (licence)We present pure-transformer based models for video classification, drawing upon the recent success of such models in image classification.
-
23 Mar 2022 9 repositories listed Syntology ran 9 of 13 samples · 4 unverified · 12 pointer-only (licence)Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets.
-
2 Dec 2021 9 repositories listedIn this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection.
-
30 Nov 2018 9 repositories listed Syntology ran 8 of 15 samples · 7 unverified · 4 pointer-only (licence)In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be…
-
23 May 2017 9 repositories listedThe AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time, resulting in 1.
-
22 Apr 2021 8 repositories listed Syntology ran 13 of 26 samples · 13 unverified · 5 pointer-only (licence)We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale…
-
2 Jan 2021 8 repositories listedConvolutional Neural Networks (CNNs) use pooling to decrease the size of activation maps.
-
5 Nov 2018 8 repositories listedIn this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatial temporal network (StNet) architecture for both local and global spatial-temporal modeling in videos.
-
23 Jun 2020 7 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedThis paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
-
4 Apr 2019 7 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIt is natural to ask: 1) if group convolution can help to alleviate the high computational cost of video classification networks; 2) what factors matter the most in 3D group convolutional networks; and 3) what are good…
-
14 Jan 2018 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Over the past decade, multivariate time series classification has received great attention.
-
27 Sep 2016 7 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedDespite the size of the dataset, some of our models train to convergence in less than a day on a single machine using TensorFlow.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections