Browse State-of-the-Art › Action Recognition In Videos
Action Recognition In Videos
71 papers with code · 17 benchmarks · 17 datasets archive 2025-07-28
Action Recognition in Videos is a task in computer vision and pattern recognition where the goal is to identify and categorize human actions performed in a video sequence. The task involves analyzing the spatiotemporal dynamics of the actions and mapping them to a predefined set of action classes, such as running, jumping, or swimming.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
17 leaderboard tables shown for this task, 17 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 17 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
17 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 71 papers with code (124 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
2 Dec 2014 29 repositories listedWe propose a simple, yet effective approach for spatiotemporal feature learning using deep 3-dimensional convolutional networks (3D ConvNets) trained on a large scale supervised video dataset.
-
2 Aug 2016 22 repositories listed Syntology ran 2 of 24 samples · 22 unverified · 3 pointer-only (licence)The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network.
-
10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present SlowFast networks for video recognition.
-
8 May 2017 11 repositories listedFurthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
-
27 Sep 2016 7 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedDespite the size of the dataset, some of our models train to convergence in less than a day on a single machine using TensorFlow.
-
9 Jun 2014 7 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art.
-
3 Dec 2012 7 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedTo the best of our knowledge, UCF101 is currently the most challenging dataset of actions due to its large number of classes, large number of clips and also unconstrained nature of such clips.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
-
15 Nov 2019 5 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 2 pointer-only (licence)YOWO is a single-stage architecture with two branches to extract temporal and spatial information concurrently and predict bounding boxes and action probabilities directly from video clips in one evaluation.
-
2 Oct 2018 5 repositories listedOur representation flow layer is a fully-differentiable layer designed to capture the `flow' of any representation channel within a convolutional neural network for action recognition.
-
22 Nov 2017 5 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species.
-
8 Jul 2015 5 repositories listedHowever, for action recognition in videos, the improvement of deep convolutional networks is not so evident.
-
1 Jun 2023 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedModern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance.
-
11 Apr 2024 3 repositories listedThese spatial features then undergo intermediate temporal modeling facilitated by the Mamba block before progressing to the encoder section, which comprises vanilla upsampling Shift S-GCN blocks.
-
25 Nov 2019 3 repositories listedWe propose a new STAckable Recurrent cell (STAR) for recurrent neural networks (RNNs), which has fewer parameters than widely used LSTM and GRU while being more robust against vanishing or exploding gradients.
-
29 May 2019 3 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedConsider end-to-end training of a multi-modal vs.
-
22 Mar 2017 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe address the problem of activity detection in continuous, untrimmed video streams.
-
17 Feb 2023 2 repositories listedTo extend our approach to video, we integrate ConvNets with state-of-the-art temporal methods such as Transformer and Recurrent Neural Networks.
-
22 Nov 2021 2 repositories listedComputer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications.
-
17 Sep 2021 2 repositories listed Syntology ran 5 of 9 samples · 4 unverifiedMoreover, to handle the deficiency of label texts and make use of tremendous web data, we propose a new paradigm based on this multimodal learning framework for action recognition, which we dub "pre-train, prompt and…
-
29 Mar 2021 2 repositories listedWe design a trainable Motion Band-Pass Module (MBPM) for separating busy information from quiet information in raw video data.
-
6 Aug 2020 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)With the proposed Inter-Intra Contrastive (IIC) framework, we can train spatio-temporal convolutional networks to learn video representations.
-
13 Jul 2020 2 repositories listedMost current action recognition methods heavily rely on appearance information by taking an RGB sequence of entire image regions as input.
-
20 May 2019 2 repositories listedIn particular, it can effectively learn representations for videos by mixing appearance and long-range motion with an RGB-only input.
-
4 Apr 2019 2 repositories listedRecently, convolutional neural networks with 3D kernels (3D CNNs) have been very popular in computer vision community as a result of their superior ability of extracting spatio-temporal features within video frames…
-
26 Feb 2018 2 repositories listed Syntology ran 0 of 25 samples · 25 unverifiedAction recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature.
-
22 Apr 2016 2 repositories listedRecent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information.
-
12 Nov 2015 2 repositories listedWe propose a soft attention based model for the task of action recognition in videos.
-
19 Mar 2025 1 repository listedIn this paper, we propose BHaRNet (Body-Hand action Recognition Network), a novel framework that augments a typical body-expert model with a hand-expert model.
-
10 Aug 2024 1 repository listedIn this work, we present an efficient pose-driven attention-guided multimodal network (EPAM-Net) for action recognition in videos.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections