Browse State-of-the-Art › Video Action Detection
Video Action Detection
19 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
19 shown of 19 papers with code (32 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jul 2020 3 repositories listedIn this work, we first empirically find the recognition accuracy is highly correlated with the bounding box size of an actor, and thus higher resolution of actors contributes to better performance.
-
10 Dec 2024 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn this work, we focus on semi-supervised learning for video action detection.
-
17 Apr 2023 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Our EVAD consists of two specialized designs for video action detection.
-
20 Jul 2022 2 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 1 pointer-only (licence)We introduce the task of spotting temporally precise, fine-grained events in video (detecting the precise moment in time events occur).
-
16 Apr 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We propose the Asynchronous Interaction Aggregation network (AIA) that leverages different interactions to boost action detection.
-
30 Dec 2018 2 repositories listedWhile observing complex events with multiple actors, humans do not assess each actor separately, but infer from the context.
-
4 Apr 2025 1 repository listedIn this work, we focus on scaling open-vocabulary action detection.
-
18 Dec 2024 1 repository listedVideo Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts.
-
25 Oct 2024 1 repository listed Syntology ran 0 of 6 samples · 6 unverified · 6 pointer-only (licence)This paper explores the impact of occlusions in video action detection.
-
25 Jun 2024 1 repository listedThe proliferation of complex deep learning (DL) models has revolutionized various applications, including computer vision-based solutions, prompting their integration into real-time systems.
-
14 Dec 2023 1 repository listedAddressing this gap, our paper introduces an innovative knowledge distillation framework, with the generative model for training a lightweight student model.
-
12 Dec 2023 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedFirst, we demonstrate its effectiveness on video action detection where the proposed approach outperforms prior works in semi-supervised and weakly-supervised learning along with several baseline approaches in both…
-
9 Apr 2022 1 repository listedVideo action detection (spatio-temporal action localization) is usually the starting point for human-centric intelligent analysis of videos nowadays.
-
8 Mar 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we focus on semi-supervised learning for video action detection which utilizes both labeled as well as unlabeled data.
-
13 Jun 2021 1 repository listedThis technical report analyzes an egocentric video action detection method we used in the 2021 EPIC-KITCHENS-100 competition hosted in CVPR2021 Workshop.
-
2 Apr 2021 1 repository listed Syntology ran 5 of 9 samples · 4 unverifiedWe propose TubeR: a simple solution for spatio-temporal video action detection.
-
9 Dec 2019 1 repository listedAction Detection is a complex task that aims to detect and classify human actions in video clips.
-
19 Apr 2019 1 repository listedIn this paper, we propose Spatio-TEmporal Progressive (STEP) action detector---a progressive learning framework for spatio-temporal action detection in videos.
-
30 Mar 2017 1 repository listedA video is first divided into equal length clips and for each clip a set of tube proposals are generated next based on 3D Convolutional Network (ConvNet) features.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections