Browse State-of-the-Art › Video Object Tracking
Video Object Tracking
72 papers with code · 4 benchmarks · 13 datasets archive 2025-07-28
Video Object Detection aims to detect targets in videos using both spatial and temporal information. It's usually deeply integrated with tasks such as Object Detection and Object Tracking.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| NT-VOT211 (43 rows) | ProContEXT | ProContEXT: Exploring Progressive Context Transformer for Tracking | code | — | Compare |
| CATER (7 rows) | Loci | Learning What and Where: Disentangling Location and Identity... | code | — | Compare |
| GOT-10k (1 row) | TATrack-L-GOT | Target-Aware Tracking with Long-term Context Attention | code | Syntology ran 0 of 2 samples · 2 unverified | Compare |
| SoccerNet-v2 (1 row) | CO-MOT | Bridging the Gap Between End-to-end and Non-End-to-end... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
13 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 72 papers with code (98 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 May 2017 34 repositories listed Syntology ran 16 of 28 samples · 12 unverified · 7 pointer-only (licence)The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale…
-
6 Jul 2022 21 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedYOLOv7 surpasses all known object detectors in both speed and accuracy in the range from 5 FPS to 160 FPS and has the highest accuracy 56.
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
30 Apr 2014 9 repositories listedInterestingly, for linear regression our formulation is equivalent to a correlation filter, used by some of the fastest competitive trackers.
-
7 Jan 2019 5 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedSiamese networks have drawn great attention in visual tracking because of their balanced accuracy and speed.
-
19 Oct 2018 5 repositories listedThis paper presents three fully convolutional neural network architectures which perform change detection using a pair of coregistered images.
-
1 Jun 2018 5 repositories listedVisual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks.
-
27 Oct 2022 4 repositories listedExisting Visual Object Tracking (VOT) only takes the target area in the first frame as a template.
-
27 Mar 2022 4 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe present the first comprehensive video polyp segmentation (VPS) study in the deep learning era.
-
18 Jun 2020 4 repositories listedIn this paper, we propose a novel object-aware anchor-free network to address this issue.
-
12 Dec 2018 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach.
-
4 Dec 2015 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedCorrelation Filter-based trackers have recently achieved excellent performance, showing great robustness to challenging situations exhibiting motion blur and illumination changes.
-
1 Aug 2024 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedMedical image segmentation plays a pivotal role in clinical diagnostics and treatment planning, yet existing models often face challenges in generalization and in handling both 2D and 3D data uniformly.
-
28 May 2024 2 repositories listedTechnically, we achieve this by routing samples from one modality to the expert of the others, within a mixture-of-experts framework designed for multimodal video object tracking.
-
22 May 2023 2 repositories listedExisting end-to-end Multi-Object Tracking (e2e-MOT) methods have not surpassed non-end-to-end tracking-by-detection methods.
-
11 Aug 2022 2 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)Despite the extensive adoption of machine learning on the task of visual object tracking, recent learning-based approaches have largely overlooked the fact that visual tracking is a sequence-level task in its nature;…
-
21 Mar 2022 2 repositories listedWe infer a bounding box from the segmentation mask, validate our tracker on challenging tracking datasets and achieve the new state of the art on LaSOT with a success AUC score of 69.
-
17 Dec 2021 2 repositories listedE.
-
10 Oct 2019 2 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedIn this work, we build a video dataset with fully observable and controllable object and scene bias, and which truly requires spatiotemporal understanding in order to be solved.
-
15 Apr 2019 2 repositories listedThe current strive towards end-to-end trainable computer vision systems imposes major challenges for the task of visual tracking.
-
14 Dec 2017 2 repositories listedIn order to efficiently search in such a large 4-DoF space in real-time, we formulate the problem into two 2-DoF sub-problems and apply an efficient Block Coordinates Descent solver to optimize the estimation result.
-
10 Jul 2025 1 repository listedThis paper presents enhancements to the SAM2 framework for video object tracking task, addressing challenges such as occlusions, background clutter, and target reappearance.
-
20 Dec 2024 1 repository listed Syntology ran 3 of 8 samples · 5 unverifiedThese temporal information tokens are used to guide the localization of the target in the next time state, establish long-range contextual relationships between video frames, and capture the temporal trajectory of the…
-
15 Dec 2024 1 repository listedContextual information at the video level has become increasingly crucial for visual object tracking.
-
2 Dec 2024 1 repository listedReferring video object segmentation (RVOS) requires tracking and segmenting an object throughout a video according to a given natural language expression, demanding both complex motion understanding and the alignment of…
-
20 Nov 2024 1 repository listed Syntology ran 0 of 4 samples · 4 unverified · 4 pointer-only (licence)In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with…
-
27 Oct 2024 1 repository listedTo this end, this paper presents NT-VOT211, a new benchmark tailored for evaluating visual object tracking algorithms in the challenging night-time conditions.
-
27 Oct 2024 1 repository listedRGB video object tracking is a fundamental task in computer vision.
-
14 Sep 2024 1 repository listedDifferent from existing tracking-by-detection MOT methods, AED gets rid of prior knowledge (e.
-
28 Feb 2024 1 repository listed Syntology ran 12 of 14 samples · 2 unverified · 14 pointer-only (licence)Despite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections