Browse State-of-the-Art › Video Recognition
Video Recognition
168 papers with code · 0 benchmarks · 10 datasets archive 2025-07-28
Video Recognition is a process of obtaining, processing, and analysing data that it receives from a visual source, specifically video.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 168 papers with code (307 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution10 Apr 2019 28 repositories listed Syntology ran 14 of 34 samples · 20 unverified · 8 pointer-only (licence)Similarly, the output feature maps of a convolution layer can also be seen as a mixture of information at different frequencies.
-
24 Jun 2021 15 repositories listed Syntology ran 7 of 32 samples · 25 unverifiedThe vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks.
-
10 Dec 2018 15 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present SlowFast networks for video recognition.
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
21 Jun 2021 11 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks.
-
10 Apr 2020 10 repositories listed Syntology ran 2 of 22 samples · 20 unverifiedTherefore, in the present paper, we conduct exploration study in order to improve spatiotemporal 3D CNNs as follows: (i) Recently proposed large-scale video datasets help improve spatiotemporal 3D CNNs in terms of video…
-
2 Dec 2021 9 repositories listedIn this paper, we study Multiscale Vision Transformers (MViTv2) as a unified architecture for image and video classification, as well as object detection.
-
22 Apr 2021 8 repositories listed Syntology ran 13 of 26 samples · 13 unverified · 5 pointer-only (licence)We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale…
-
9 Apr 2020 8 repositories listed Syntology ran 2 of 15 samples · 13 unverifiedThis paper presents X3D, a family of efficient video networks that progressively expand a tiny 2D image classification architecture along multiple network axes, in space, time, width and depth.
-
25 Mar 2019 7 repositories listedBatch Normalization (BN) has become an out-of-box technique to improve deep network training.
-
17 Nov 2014 7 repositories listedModels based on deep convolutional networks have dominated recent image interpretation tasks; we investigate whether models which are also recurrent, or "temporally deep", are effective for tasks involving sequences,…
-
31 Dec 2022 5 repositories listedIn this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the…
-
4 Jul 2022 5 repositories listedIn this study, we focus on transferring knowledge for video classification tasks.
-
3 Sep 2021 5 repositories listedA recent work from Bello shows that training and scaling strategies may be more significant than model architectures for visual recognition.
-
1 Jun 2023 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedModern hierarchical vision transformers have added several vision-specific components in the pursuit of supervised classification performance.
-
27 Sep 2021 4 repositories listedSecondly, TSM has high efficiency; it achieves a high frame rate of 74fps and 29fps for online video recognition on Jetson Nano and Galaxy Note8.
-
14 May 2024 3 repositories listedIn this paper, we propose to squeeze the time axis of a video sequence into the channel dimension and present a lightweight video recognition network, term as \textit{SqueezeTime}, for mobile video understanding.
-
13 Jul 2023 3 repositories listed Syntology ran 17 of 32 samples · 15 unverified · 21 pointer-only (licence)Video transformer designs are based on self-attention that can model global context at a high computational cost.
-
21 Mar 2021 3 repositories listed Syntology ran 8 of 13 samples · 5 unverifiedWe present Mobile Video Networks (MoViNets), a family of computation and memory efficient video networks that can operate on streaming video for online inference.
-
13 Dec 2020 3 repositories listedExisting state-of-the-art methods have achieved excellent accuracy regardless of the complexity meanwhile efficient spatiotemporal modeling solutions are slightly inferior in performance.
-
20 Jun 2020 3 repositories listedThis work introduces pyramidal convolution (PyConv), which is capable of processing the input at multiple filter scales.
-
29 Mar 2020 3 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedThen a joint-training strategy is proposed to deal with the domain gaps between multiple data sources and formats in webly-supervised learning.
-
23 Jan 2020 3 repositories listedWe present Audiovisual SlowFast Networks, an architecture for integrated audiovisual perception.
-
23 Nov 2016 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Yet, it is non-trivial to transfer the state-of-the-art image recognition networks to videos as per-frame evaluation is too slow and unaffordable.
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
2 Oct 2023 2 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn this paper, we present a new adaptation paradigm (ZeroI2V) to transfer the image transformers to video recognition tasks (i.
-
18 Jul 2023 2 repositories listed Syntology ran 3 of 6 samples · 3 unverifiedWe conduct comprehensive ablation studies on the instantiation of ATMs and demonstrate that this module provides powerful temporal modeling capability at a low computational cost.
-
26 Mar 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedTo fix this issue, we propose a general framework, named Frame Flexible Network (FFN), which not only enables the model to be evaluated at different frames to adjust its computation, but also reduces the memory costs of…
-
6 Aug 2022 2 repositories listed Syntology ran 5 of 10 samples · 5 unverified · 10 pointer-only (licence)Video recognition has been dominated by the end-to-end learning paradigm -- first initializing a video recognition model with weights of a pretrained image model and then conducting end-to-end training on videos.
-
4 Aug 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Extensive experiments demonstrate that our approach is effective and can be generalized to different video recognition scenarios.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections