Browse State-of-the-Art › Visual Object Tracking
Visual Object Tracking
187 papers with code · 25 benchmarks · 29 datasets archive 2025-07-28
Visual Object Tracking is an important research topic in computer vision, image understanding and pattern recognition. Given the initial state (centre location and scale) of a target in the first frame of a video sequence, the aim of Visual Object Tracking is to automatically obtain the states of the object in the subsequent video frames.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
25 leaderboard tables shown for this task, 25 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 25 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
29 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 187 papers with code (341 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Dec 2015 221 repositories listed Syntology ran 19 of 131 samples · 112 unverified · 5 pointer-only (licence)Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified…
-
31 Dec 2018 13 repositories listed Syntology ran 3 of 17 samples · 14 unverifiedMoreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size.
-
1 Aug 2024 11 repositories listed Syntology ran 28 of 49 samples · 21 unverifiedWe present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos.
-
16 Nov 2016 8 repositories listedThis paper tackles the task of semi-supervised video object segmentation, i.
-
14 Nov 2019 6 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Following these guidelines, we design our Fully Convolutional Siamese tracker++ (SiamFC++) by introducing both classification and target state estimation branch(G1), classification score without ambiguity(G2), tracking…
-
23 May 2024 5 repositories listedTo leverage more modalities, some recent efforts have been made to learn a unified visual object tracking model for any modality.
-
31 Mar 2021 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedTo learn generalizable representation for correspondence in large-scale, a variety of self-supervised pretext tasks are proposed to explicitly perform object-level or patch-level similarity learning.
-
7 Jan 2019 5 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedSiamese networks have drawn great attention in visual tracking because of their balanced accuracy and speed.
-
1 Jun 2018 5 repositories listedVisual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks.
-
28 Nov 2016 5 repositories listedMoreover, our fast variant, using hand-crafted features, operates at 60 Hz on a single CPU, while obtaining 65.
-
26 Sep 2023 4 repositories listed Syntology ran 9 of 9 samples · 0 unverified · 9 pointer-only (licence)Tracking using bio-inspired event cameras has drawn more and more attention in recent years.
-
27 Oct 2022 4 repositories listedExisting Visual Object Tracking (VOT) only takes the target area in the first frame as a template.
-
18 Jun 2020 4 repositories listedIn this paper, we propose a novel object-aware anchor-free network to address this issue.
-
7 Feb 2019 4 repositories listedIt combines a Convolutional Neural Network (CNN) backbone and a cross-correlation operator, and takes advantage of the features from exemplary images for more accurate object tracking.
-
19 Nov 2018 4 repositories listed Syntology ran 3 of 9 samples · 6 unverifiedWe argue that this approach is fundamentally limited since target estimation is a complex task, requiring high-level knowledge about the object.
-
3 Sep 2018 4 repositories listedEnd-to-end sequential learning to explore spatial-temporal features for video segmentation is largely limited by the scale of available video segmentation datasets, i.
-
25 Nov 2016 4 repositories listedShort-term tracking is an open and challenging problem for which discriminative correlation filters (DCF) have shown excellent performance.
-
12 Dec 2018 3 repositories listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach.
-
4 Dec 2015 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedCorrelation Filter-based trackers have recently achieved excellent performance, showing great robustness to challenging situations exhibiting motion blur and illumination changes.
-
25 Sep 2024 2 repositories listedTo bridge this gap, we take a step forward by proposing the first large-scale multi-modal underwater camouflaged object tracking dataset, namely UW-COT220.
-
6 Mar 2024 2 repositories listedThe rich annotations of VastTrack enables development of both the vision-only and the vision-language tracking.
-
4 Feb 2024 2 repositories listedThen, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction.
-
11 Aug 2022 2 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)Despite the extensive adoption of machine learning on the task of visual object tracking, recent learning-based approaches have largely overlooked the fact that visual tracking is a sequence-level task in its nature;…
-
21 Mar 2022 2 repositories listedWe infer a bounding box from the segmentation mask, validate our tracker on challenging tracking datasets and achieve the new state of the art on LaSOT with a success AUC score of 69.
-
30 Jan 2022 2 repositories listed Syntology ran 4 of 30 samples · 26 unverifiedThe canonical object representation is learned solely in simulation and then used to parse a category-level, task trajectory from a single demonstration video.
-
17 Dec 2021 2 repositories listedE.
-
4 Jun 2021 2 repositories listedThe state-of-the-art methods learn to decode features with a single positive object and thus have to match and segment each target separately under multi-object scenarios, consuming multiple times computing resources.
-
31 Mar 2021 2 repositories listedWe believe this benchmark will greatly boost related researches on natural language guided tracking.
-
24 Dec 2020 2 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)We further show that this change in orientation can be used to impose an additional motion constraint in Siamese tracking through imposing restriction on the change in orientation between two consecutive frames.
-
8 Dec 2020 2 repositories listedVisual object tracking, as a fundamental task in computer vision, has drawn much attention in recent years.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections