Browse State-of-the-Art › Action Localization
Action Localization
169 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Action Localization is finding the spatial and temporal co ordinates for an action in a video. An action localization model will identify which frame an action start and ends in video and return the x,y coordinates of an action. Further the co ordinates will change when the object performing action undergoes a displacement.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 169 papers with code (369 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 May 2017 9 repositories listedThe AVA dataset densely annotates 80 atomic visual actions in 430 15-minute video clips, where actions are localized in space and time, resulting in 1.
-
21 Dec 2021 5 repositories listedWeakly-supervised temporal action localization (WTAL) in untrimmed videos has emerged as a practical but challenging task since only video-level labels are available.
-
15 Nov 2019 5 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 2 pointer-only (licence)YOWO is a single-stage architecture with two branches to extract temporal and spatial information concurrently and predict bounding boxes and action probabilities directly from video clips in one evaluation.
-
10 Jul 2020 4 repositories listedRecognition of surgical activity is an essential component to develop context-aware decision support for the operating room.
-
13 Dec 2019 4 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedAnnotating videos is cumbersome, expensive and not scalable.
-
7 Jun 2019 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we propose instead to learn such embeddings from video data with readily available natural language annotations in the form of automatically transcribed narrations.
-
11 Sep 2023 3 repositories listedTemporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video.
-
7 Dec 2020 3 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedRecent data augmentation strategies have been reported to address the overfitting problems in static image classifiers.
-
16 Jun 2020 3 repositories listedThis technical report introduces our winning solution to the spatio-temporal action localization track, AVA-Kinetics Crossover, in ActivityNet Challenge 2020.
-
14 Jun 2020 3 repositories listed Syntology ran 1 of 27 samples · 26 unverifiedWe propose to explicitly model the Actor-Context-Actor Relation, which is the relation between two actors based on their interactions with the context.
-
14 Dec 2017 3 repositories listedWe propose a weakly supervised temporal action localization algorithm on untrimmed videos using convolutional neural networks.
-
13 Apr 2017 3 repositories listedWe propose `Hide-and-Seek', a weakly-supervised framework that aims to improve object localization in images and action localization in videos.
-
13 Nov 2023 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)We first benchmark MM-Navigator on our collected iOS screen dataset.
-
19 Jun 2023 2 repositories listedTo fill this gap, we introduce the FHA-Kitchens (Fine-Grained Hand Actions in Kitchen Scenes) dataset, providing both coarse- and fine-grained hand action categories along with localization annotations.
-
16 Nov 2022 2 repositories listedThis report describes our submission to the Ego4D Moment Queries Challenge 2022.
-
22 Jun 2022 2 repositories listedAccordingly, we first exclude these surely non-existent categories by a complementary learning loss.
-
20 May 2022 2 repositories listedTo tackle this issue, we make an early effort to study temporal action localization from the perspective of multi-modality feature learning, based on the observation that different actions exhibit specific preferences…
-
9 Dec 2021 2 repositories listedModern self-supervised learning algorithms typically enforce persistency of instance representations across views.
-
28 Sep 2021 2 repositories listedWe present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks.
-
27 Jul 2021 2 repositories listedIn this work, we argue that the features extracted from the pretrained extractor, e.
-
7 Apr 2021 2 repositories listedTraditional methods mainly focus on foreground and background frames separation with only a single attention branch and class activation sequence.
-
12 Jun 2020 2 repositories listedExperimental results show that our uncertainty modeling is effective at alleviating the interference of background frames and brings a large performance gain without bells and whistles.
-
8 Dec 2019 2 repositories listedIn this report, we introduce the Winner method for HACS Temporal Action Localization Challenge 2019.
-
22 Nov 2019 2 repositories listedThis formulation does not fully model the problem in that background frames are forced to be misclassified as action classes to predict video-level labels accurately.
-
26 Aug 2019 2 repositories listedIn this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which…
-
6 Nov 2018 2 repositories listedOur approach only needs to modify the input image and can work with any network to improve its performance.
-
5 Apr 2018 2 repositories listedSecond, we propose an actor-based attention mechanism that enables the localization of the actions from action class labels and actor proposals and is end-to-end trainable.
-
26 Dec 2017 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThis paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos.
-
4 May 2017 2 repositories listedWe propose the ACtion Tubelet detector (ACT-detector) that takes as input a sequence of frames and outputs tubelets, i.
-
4 Jun 2025 1 repository listedLocating human-object interaction (HOI) actions within video serves as the foundation for multiple downstream tasks, such as human behavior analysis and human-robot skill transfer.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections