Browse State-of-the-Art › Temporal Action Localization
Temporal Action Localization
493 papers with code · 14 benchmarks · 42 datasets archive 2025-07-28
Temporal Action Localization aims to detect activities in the video stream and output beginning and end timestamps. It is closely related to Temporal Action Proposal Generation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
14 leaderboard tables shown for this task, 14 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 14 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
42 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 42 until expanded.
Subtasks archive 2025-07-28
8 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 493 papers with code (1,477 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jan 2018 24 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedDynamics of human body skeletons convey significant information for human action recognition.
-
30 Nov 2017 24 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition.
-
30 Oct 2017 24 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 3 pointer-only (licence)Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems.
-
2 Aug 2016 22 repositories listed Syntology ran 2 of 24 samples · 22 unverified · 3 pointer-only (licence)The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network.
-
8 Jun 2018 17 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion…
-
23 Jul 2019 15 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedTo address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending…
-
16 Feb 2015 12 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)We further evaluate the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.
-
8 May 2017 11 repositories listedFurthermore, based on the temporal segment networks, we won the video classification track at the ActivityNet challenge 2016 among 24 teams, which demonstrates the effectiveness of TSN and the proposed good practices.
-
30 Nov 2018 9 repositories listed Syntology ran 8 of 15 samples · 7 unverified · 4 pointer-only (licence)In this work, we propose a new approach for reasoning globally in which a set of features are globally aggregated over the coordinate space and then projected to an interaction space where relational reasoning can be…
-
23 May 2017 9 repositories listedThe AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time, resulting in 1.
-
5 Nov 2018 8 repositories listedIn this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatial temporal network (StNet) architecture for both local and global spatial-temporal modeling in videos.
-
26 Nov 2019 7 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedIn this work, we propose a graph convolutional network (GCN) model to adaptively incorporate multi-level semantic context into video features and cast temporal action detection as a sub-graph localization problem.
-
14 Jan 2018 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Over the past decade, multivariate time series classification has received great attention.
-
9 Jun 2014 7 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)Our architecture is trained and evaluated on the standard video actions benchmarks of UCF-101 and HMDB-51, where it is competitive with the state of the art.
-
3 Dec 2012 7 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedTo the best of our knowledge, UCF101 is currently the most challenging dataset of actions due to its large number of classes, large number of clips and also unconstrained nature of such clips.
-
14 Nov 2022 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedWe launch EVA, a vision-centric foundation model to explore the limits of visual representation at scale using only publicly accessible data.
-
17 Apr 2018 6 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)Skeleton-based human action recognition has recently drawn increasing attentions with the availability of large-scale skeleton datasets.
-
19 Nov 2015 6 repositories listedWe propose an approach to learn spatio-temporal features in videos from intermediate visual representations we call "percepts" using Gated-Recurrent-Unit Recurrent Networks (GRUs).
-
21 Dec 2021 5 repositories listedWeakly-supervised temporal action localization (WTAL) in untrimmed videos has emerged as a practical but challenging task since only video-level labels are available.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
-
8 Apr 2019 5 repositories listedCan performance on the task of action quality assessment (AQA) be improved by exploiting a description of the action and its quality?
-
2 Oct 2018 5 repositories listedOur representation flow layer is a fully-differentiable layer designed to capture the `flow' of any representation channel within a convolutional neural network for action recognition.
-
23 May 2017 5 repositories listedShot boundary detection (SBD) is an important component of many video analysis tasks, such as action recognition, video indexing, summarization and editing.
-
8 Jul 2015 5 repositories listedHowever, for action recognition in videos, the improvement of deep convolutional networks is not so evident.
-
27 Feb 2015 5 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIn this context, we propose an approach that successfully takes into account both the local and global temporal structure of videos to produce descriptions.
-
1 Jul 2021 4 repositories listedDeep neural networks based purely on attention have been successful across several domains, relying on minimal architectural priors from the designer.
-
20 May 2018 4 repositories listedIn addition, the second-order information (the lengths and directions of bones) of the skeleton data, which is naturally more informative and discriminative for action recognition, is rarely investigated in existing…
-
9 Jan 2018 4 repositories listedWe present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds.
-
12 Dec 2017 4 repositories listedSecond, we show the power of hallucinated flow for recognition, successfully transferring the learned motion into a standard two-stream network for activity recognition.
-
30 Mar 2017 4 repositories listedWe demonstrate that using both RNNs (using LSTMs) and Temporal-ConvNets on spatiotemporal feature matrices are able to exploit spatiotemporal dynamics to improve the overall performance.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections