Browse State-of-the-Art › Action Detection
Action Detection
277 papers with code · 11 benchmarks · 34 datasets archive 2025-07-28
Action Detection aims to find both where and when an action occurs within a video clip and classify what the action is taking place. Typically results are given in the form of action tublets, which are action bounding boxes linked across time in the video. This is related to temporal localization, which seeks to identify the start and end frame of an action, and action recognition, which seeks only to classify which action is taking place and typically assumes a trimmed video.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
11 leaderboard tables shown for this task, 11 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 11 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
34 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 34 until expanded.
Subtasks archive 2025-07-28
8 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 277 papers with code (817 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Sep 2015 161 repositories listed Syntology ran 158 of 306 samples · 148 unverified · 163 pointer-only (licence)We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain.
-
8 Jun 2018 17 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Temporal action proposal generation is an important yet challenging problem, since temporal proposals with rich action content are indispensable for analysing real-world videos with long duration and high proportion…
-
23 Jul 2019 15 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedTo address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending…
-
10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present SlowFast networks for video recognition.
-
23 May 2017 9 repositories listedThe AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time, resulting in 1.
-
23 Jun 2020 7 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedThis paper introduces the pipeline to extend the largest dataset in egocentric vision, EPIC-KITCHENS.
-
10 Apr 2022 6 repositories listedIn this paper, we present the challenge setup and assessment of the state-of-the-art deep learning methods proposed by the participants during the challenge.
-
20 Apr 2017 6 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedDetecting actions in untrimmed videos is an important yet challenging task.
-
15 Nov 2019 5 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 2 pointer-only (licence)YOWO is a single-stage architecture with two branches to extract temporal and spatial information concurrently and predict bounding boxes and action probabilities directly from video clips in one evaluation.
-
20 Nov 2020 4 repositories listedWith the advancement in computer vision deep learning, systems now are able to analyze an unprecedented amount of rich visual information from videos to enable applications such as autonomous driving, socially-aware…
-
13 Apr 2019 4 repositories listedTo address these and promote the activity understanding, we build a large-scale Human Activity Knowledge Engine (HAKE) based on the human body part states.
-
21 Nov 2018 4 repositories listedFine-grained action detection is an important task with numerous applications in robotics and human-computer interaction.
-
17 Sep 2024 3 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedOur resulting model is the first real-time full-duplex spoken large language model, with a theoretical latency of 160ms, 200ms in practice, and is available at https://github.
-
14 May 2024 3 repositories listedIn this paper, we propose to squeeze the time axis of a video sequence into the channel dimension and present a lightweight video recognition network, term as \textit{SqueezeTime}, for mobile video understanding.
-
11 Sep 2023 3 repositories listedTemporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video.
-
12 Nov 2022 3 repositories listedEnd-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once.
-
23 Feb 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We also report the performance on the ROAD tasks of Slowfast and YOLOv5 detectors, as well as that of the winners of the ICCV2021 ROAD challenge, which highlight the challenges faced by situation awareness in autonomous…
-
20 Jul 2020 3 repositories listedIn this work, we first empirically find the recognition accuracy is highly correlated with the bounding box size of an actor, and thus higher resolution of actors contributes to better performance.
-
14 Jun 2020 3 repositories listed Syntology ran 1 of 27 samples · 26 unverifiedWe propose to explicitly model the Actor-Context-Actor Relation, which is the relation between two actors based on their interactions with the context.
-
2 Dec 2019 3 repositories listed Syntology ran 2 of 10 samples · 8 unverified · 3 pointer-only (licence)We empirically demonstrate a general and robust grid schedule that yields a significant out-of-the-box training speedup without a loss in accuracy for different models (I3D, non-local, SlowFast), datasets (Kinetics,…
-
4 Nov 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe introduce pyannote.
-
9 Jun 2019 3 repositories listedIn the end, a posteriori SNR weighted energy difference is applied to the extended pitch segments of the denoised speech signal for detecting voice activity.
-
9 Apr 2018 3 repositories listedIn this paper, we introduce a challenging new dataset, MLB-YouTube, designed for fine-grained activity detection.
-
22 Mar 2017 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe address the problem of activity detection in continuous, untrimmed video streams.
-
28 Nov 2016 3 repositories listedWe propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection.
-
29 Aug 2016 3 repositories listedThis thesis explore different approaches using Convolutional and Recurrent Neural Networks to classify and temporally localize activities on videos, furthermore an implementation to achieve it has been proposed.
-
10 Dec 2024 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn this work, we focus on semi-supervised learning for video action detection.
-
31 May 2024 2 repositories listedTo address these challenges, we propose a novel end-to-end skeleton-based model called Skeleton-OOD, which is committed to improving the effectiveness of OOD tasks while ensuring the accuracy of ID recognition.
-
28 Nov 2023 2 repositories listedIn this paper, we reduce the memory consumption for end-to-end training, and manage to scale up the TAD backbone to 1 billion parameters and the input video to 1, 536 frames, leading to significant detection performance.
-
17 Jul 2023 2 repositories listedWe introduce "ivrit.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections