Browse State-of-the-Art › Weakly Supervised Action Localization
Weakly Supervised Action Localization
35 papers with code · 9 benchmarks · 5 datasets archive 2025-07-28
In this task, the training data consists of videos with a list of activities in them without any temporal boundary annotations. However, while testing, given a video, the algorithm should recognize the activities in the video and also provide the start and end time.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
9 leaderboard tables shown for this task, 9 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (55 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Jul 2020 4 repositories listedRecognition of surgical activity is an essential component to develop context-aware decision support for the operating room.
-
7 Dec 2020 3 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedRecent data augmentation strategies have been reported to address the overfitting problems in static image classifiers.
-
14 Dec 2017 3 repositories listedWe propose a weakly supervised temporal action localization algorithm on untrimmed videos using convolutional neural networks.
-
13 Apr 2017 3 repositories listedWe propose `Hide-and-Seek', a weakly-supervised framework that aims to improve object localization in images and action localization in videos.
-
22 Jun 2022 2 repositories listedAccordingly, we first exclude these surely non-existent categories by a complementary learning loss.
-
27 Jul 2021 2 repositories listedIn this work, we argue that the features extracted from the pretrained extractor, e.
-
7 Apr 2021 2 repositories listedTraditional methods mainly focus on foreground and background frames separation with only a single attention branch and class activation sequence.
-
12 Jun 2020 2 repositories listedExperimental results show that our uncertainty modeling is effective at alleviating the interference of background frames and brings a large performance gain without bells and whistles.
-
22 Nov 2019 2 repositories listedThis formulation does not fully model the problem in that background frames are forced to be misclassified as action classes to predict video-level labels accurately.
-
5 Apr 2018 2 repositories listedSecond, we propose an actor-based attention mechanism that enables the localization of the actions from action class labels and actor proposals and is end-to-end trainable.
-
9 Mar 2017 2 repositories listedWe exploit the learned models for action recognition (WSR) and detection (WSD) on the untrimmed video datasets of THUMOS14 and ActivityNet.
-
24 Nov 2024 1 repository listedThe AAL branch uses pseudo labels to learn class-agnostic action information.
-
15 Apr 2024 1 repository listedOur method seeks to suppress false positive backgrounds without introducing the background category.
-
21 Dec 2023 1 repository listedIt comprises two core components: a snippet clustering component that groups the snippets into multiple latent clusters and a cluster classification component that further classifies the cluster as foreground or…
-
22 Oct 2023 1 repository listedTo alleviate this issue, we proposed a novel learning pattern in our training stage, which maximizes the probability of action union of surrounding timestamps for unlabeled frames.
-
24 Aug 2023 1 repository listedFor snippet-level learning, we introduce an online-updated memory to store reliable snippet prototypes for each class.
-
26 Jun 2023 1 repository listedIn principle, the two branches are supposed to produce the same actionness activation.
-
29 May 2023 1 repository listedWeakly-supervised temporal action localization aims to localize and recognize actions in untrimmed videos with only video-level category labels during training.
-
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions24 Jul 2022 1 repository listedAction understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences.
-
1 May 2022 1 repository listedC³BN consists of two key ingredients: a micro data augmentation strategy that increases the diversity in-between adjacent snippets by convex combination of adjacent snippets, and a macro-micro consistency regularization…
-
31 Mar 2022 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedWe target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training.
-
11 Aug 2021 1 repository listedTo learn completeness from the obtained sequence, we introduce two novel losses that contrast action instances with background ones in terms of action score and feature similarity, respectively.
-
24 May 2021 1 repository listedTemporal action localization (TAL) is an important and challenging problem in video understanding.
-
6 May 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)A common approach is to train a frame-level classifier where frames with the highest class probability are selected to make a video-level prediction.
-
30 Mar 2021 1 repository listedIn this paper, we argue that learning by comparing helps identify these hard snippets and we propose to utilize snippet Contrastive learning to Localize Actions, CoLA for short.
-
11 Mar 2021 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)To demonstrate the effectiveness of timestamp supervision, we propose an approach to train a segmentation model using only timestamps annotations.
-
D2-Net: Weakly-Supervised Action Localization via Discriminative Embeddings and Denoised Activations11 Dec 2020 1 repository listedThe proposed formulation comprises a discriminative and a denoising loss term for enhancing temporal action localization.
-
13 Jul 2020 1 repository listedTwo triplets of the feature space are considered in our approach: one triplet is used to learn discriminative features for each activity class, and the other one is used to distinguish the features where no activity…
-
31 Mar 2020 1 repository listedWeakly-supervised action localization requires training a model to localize the action segments in the video given only video level action label.
-
27 Mar 2020 1 repository listedBy maximizing the conditional probability with respect to the attention, the action and non-action frames are well separated.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections