Browse State-of-the-Art › Video Object Detection
Video Object Detection
72 papers with code · 7 benchmarks · 13 datasets archive 2025-07-28
Video object detection is the task of detecting objects from a video as opposed to images.
( Image credit: Learning Motion Priors for Efficient Video Object Detection )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task (1 more in the archive withheld as spam; see /not-shown), 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ImageNet VID (33 rows) | YOLOV++ | Practical Video Object Detection via Feature Selection and Aggregation | code | — | Compare |
| EPIC KITCHENS-unseen splits (1 row) | Temporal ROI Align | Temporal RoI Align for Video Object Recognition | code | — | Compare |
| EPIC-KITCHENS-55 (1 row) | Ours (Faster RCNN) | Objects do not disappear: Video object detection by single-frame... | code | — | Compare |
| EPIC KITCHENS-seen splits (1 row) | Temporal ROI Align | Temporal RoI Align for Video Object Recognition | code | — | Compare |
| USC-GRAD-STDdb (1 row) | SLTnet FPN-X101 | Short-term anchor linking and long-term self-guided attention for... | code | — | Compare |
| Waymo Open Dataset (1 row) | (unnamed in the archive) | Objects do not disappear: Video object detection by single-frame... | code | — | Compare |
| YT-BB (1 row) | (unnamed in the archive) | Objects do not disappear: Video object detection by single-frame... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
13 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 72 papers with code (147 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Apr 2021 32 repositories listed Syntology ran 5 of 20 samples · 15 unverified · 2 pointer-only (licence)In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets).
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
13 Jan 2022 3 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 3 pointer-only (licence)Detection Transformer (DETR) and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted…
-
14 Apr 2021 3 repositories listedThis paper presents HoughNet, a one-stage, anchor-free, voting-based, bottom-up object detection method.
-
7 Dec 2019 3 repositories listedIn this paper we propose a method that leverages temporal context from the unlabeled frames of a novel camera to improve performance at that camera.
-
30 Nov 2018 3 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Adversarial examples have been demonstrated to threaten many computer vision tasks including object detection.
-
16 Apr 2018 3 repositories listedIn this paper, we present a light weight network architecture for video object detection on mobiles.
-
17 Nov 2017 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)This paper introduces an online model for object detection in videos designed to run in real-time on low-powered mobile and embedded devices.
-
4 Feb 2024 2 repositories listedThen, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction.
-
9 Aug 2023 2 repositories listed2) Improved efficiency by only doing the expensive feature computations on a small subset of all frames.
-
9 Apr 2020 2 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedWeakly supervised learning has emerged as a compelling tool for object detection by reducing the need for strong supervision during training.
-
26 Mar 2020 2 repositories listedWe argue that there are two important cues for humans to recognize objects in videos: the global semantic information and the local localization information.
-
2 Mar 2020 2 repositories listedAs the tracker reuses the features from the detector, it is a very light-weighted increment to the detection network.
-
16 Jan 2020 2 repositories listedWe provide a large-scale drone captured dataset, VisDrone, which includes four tracks, i.
-
26 Aug 2019 2 repositories listedIn this paper, we introduce a new design to capture the interactions across the objects in spatio-temporal context.
-
15 Jul 2019 2 repositories listedIn this work, we argue that aggregating features in the full-sequence level will lead to more discriminative and robust features for video object detection.
-
25 Mar 2019 2 repositories listedModels and examples built with TensorFlow
-
29 Mar 2017 2 repositories listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)The accuracy of detection suffers from degenerated object appearances in videos, e.
-
11 Aug 2024 1 repository listedFalling objects from buildings can cause severe injuries to pedestrians due to the great impact force they exert.
-
29 Jul 2024 1 repository listedIn principle, the detection in a certain frame of a video can benefit from information in other frames.
-
17 Apr 2024 1 repository listedThis paper introduces Multi-Resolution Rescored Byte-Track (MR2-ByteTrack), a novel video object detection framework for ultra-low-power embedded processors.
-
16 Apr 2024 1 repository listedWe present a scalable framework designed to craft efficient lightweight models for video object detection utilizing self-training and knowledge distillation techniques.
-
28 Feb 2024 1 repository listedUrban traffic environments present unique challenges for object detection, particularly with the increasing presence of micromobility vehicles like e-scooters and bikes.
-
16 Feb 2024 1 repository listedThe objective of our work is to leverage this complementary information to improve detection.
-
14 Feb 2024 1 repository listedDeep video models, for example, 3D CNNs or video transformers, have achieved promising performance on sparse video tasks, i.
-
14 Feb 2024 1 repository listedBased on the analysis, we present a simple yet efficient framework to address the computational bottlenecks and achieve efficient one-stage VOD by exploiting the temporal consistency in video frames.
-
18 Jan 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedHowever, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2)…
-
30 Oct 2023 1 repository listedTo effectively refine the box from the degraded images in the videos, we used three novel approaches: cascade refinement, dynamic core-set conditioning, and local batch refinement.
-
25 Aug 2023 1 repository listed Syntology ran 7 of 25 samples · 18 unverifiedIn this work, we exploit temporal redundancy between subsequent inputs to reduce the cost of Transformers for video processing.
-
22 Aug 2023 1 repository listedThe ODD score enhances the VOD system in two ways: 1) it enables the VOD system to select superior global reference frames, thereby improving overall accuracy; and 2) it serves as an indicator in the newly designed ODD…
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections