Papers › Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action...

Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach

21 Dec 2023ICCV 2023 1arXiv:2312.14138archive 2025-07-28

Qinying Liu, Zilei Wang, Shenghai Rong, Junjie Li, Yixin Zhang

Weakly-supervised temporal action localization aims to localize action instances in videos with only video-level action labels. Existing methods mainly embrace a localization-by-classification pipeline that optimizes the snippet-level prediction with a video classification loss. However, this formulation suffers from the discrepancy between classification and detection, resulting in inaccurate separation of foreground and background (F\&B) snippets. To alleviate this problem, we propose to explore the underlying structure among the snippets by resorting to unsupervised snippet clustering, rather than heavily relying on the video classification loss. Specifically, we propose a novel clustering-based F\&B separation algorithm. It comprises two core components: a snippet clustering component that groups the snippets into multiple latent clusters and a cluster classification component that further classifies the cluster as foreground or background. As there are no ground-truth labels to train these two components, we introduce a unified self-labeling mechanism based on optimal transport to produce high-quality pseudo-labels that match several plausible prior distributions. This ensures that the cluster assignments of the snippets can be accurately associated with their F\&B labels, thereby boosting the F\&B separation. We evaluate our method on three benchmarks: THUMOS14, ActivityNet v1.2 and v1.3. Our method achieves promising performance on all three benchmarks while being significantly more lightweight than previous methods. Code is available at https://github.com/Qinying-Liu/CASE

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

qinying-liu/case officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action LocalizationClassificationClusteringTemporal Action LocalizationVideo ClassificationWeakly Supervised Action LocalizationWeakly-supervised Temporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Weakly Supervised Action Localization ActivityNet-1.2 CASE Mean mAP 27.9 #4 of 19 Archive leaderboard report
Weakly Supervised Action Localization ActivityNet-1.2 CASE mAP@0.5 43.8 #4 of 19 Archive leaderboard report
Weakly Supervised Action Localization ActivityNet-1.3 CASE mAP@0.5 43.2 #5 of 17 Archive leaderboard report
Weakly Supervised Action Localization ActivityNet-1.3 CASE mAP@0.5:0.95 26.8 #5 of 17 Archive leaderboard report
Weakly Supervised Action Localization THUMOS 2014 CASE + Zhou et al. mAP@0.1:0.7 49.2 #5 of 30 Archive leaderboard report
Weakly Supervised Action Localization THUMOS 2014 CASE mAP@0.1:0.5 57.1 #10 of 30 Archive leaderboard report
Weakly Supervised Action Localization THUMOS 2014 CASE mAP@0.1:0.7 46.2 #10 of 30 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections