Browse State-of-the-Art › Visual Object Tracking

Visual Object Tracking

187 papers with code · 25 benchmarks · 29 datasets archive 2025-07-28

Computer Vision

Visual Object Tracking is an important research topic in computer vision, image understanding and pattern recognition. Given the initial state (centre location and scale) of a target in the first frame of a video sequence, the aim of Visual Object Tracking is to automatically obtain the states of the object in the subsequent video frames.

Source: Learning Adaptive Discriminative Correlation Filters via Temporal Consistency Preserving Spatial Feature Selection for Robust Visual Object Tracking

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

25 leaderboard tables shown for this task, 25 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 25 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
LaSOT (46 rows) SPMTrack-G SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with... code Syntology ran 0 of 1 samples · 1 unverified Compare
GOT-10k (42 rows) SAMURAI-L SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual... code Syntology ran 1 of 1 samples · 0 unverified Compare
TrackingNet (40 rows) MCITrack-L384 Exploring Enhanced Contextual Information for Video-Level Object Tracking code — Compare
LaSOT-ext (18 rows) SAMURAI-L SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual... code Syntology ran 1 of 1 samples · 0 unverified Compare
OTB-2015 (18 rows) SPMTrack-B SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with... code Syntology ran 0 of 1 samples · 1 unverified Compare
TNL2K (16 rows) MCITrack-L384 Exploring Enhanced Contextual Information for Video-Level Object Tracking code — Compare
UAV123 (16 rows) LoRAT-g-378 Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance code Syntology ran 3 of 5 samples · 2 unverified Compare
VOT2017/18 (15 rows) SiamMask_E Fast Visual Object Tracking with Rotated Bounding Boxes code — Compare
DiDi (11 rows) DAM4SAM A Distractor-Aware Memory for Visual Object Tracking with SAM2 code Syntology ran 3 of 11 samples · 8 unverified Compare
NeedForSpeed (10 rows) SAMURAI-L SAMURAI: Adapting Segment Anything Model for Zero-Shot Visual... code Syntology ran 1 of 1 samples · 0 unverified Compare
YouTube-VOS 2018 (9 rows) OSVOS One-Shot Video Object Segmentation code — Compare
AVisT (7 rows) PiVOT-L Improving Visual Object Tracking through Visual Prompting code — Compare
OTB-2013 (7 rows) SE-SiamFC Scale Equivariance Improves Siamese Tracking code Syntology ran 1 of 12 samples · 11 unverified Compare
VOT2016 (6 rows) SiamMask_E Fast Visual Object Tracking with Rotated Bounding Boxes code — Compare
VOT2017 (6 rows) GFS-DCF Joint Group Feature Selection and Discriminative Filter Learning... code — Compare
VOT2022 (5 rows) DAM4SAM A Distractor-Aware Memory for Visual Object Tracking with SAM2 code Syntology ran 3 of 11 samples · 8 unverified Compare
OTB-50 (4 rows) SiamVGG SiamVGG: Visual Tracking using Deeper Siamese Networks code — Compare
VOT2019 (3 rows) TREG Target Transformed Regression for Accurate Tracking code — Compare
OTB-100 (2 rows) DiMP-NCE+ How to Train Your Energy-Based Model for Regression code Syntology ran 0 of 3 samples · 3 unverified Compare
VOT2018 (2 rows) TREG Target Transformed Regression for Accurate Tracking code — Compare
YouTube-VOS (2 rows) AOC-MF Towards Robust Video Object Segmentation with Adaptive Object Calibration code Syntology ran 2 of 4 samples · 2 unverified Compare
ITB (1 row) DropTrack DropMAE: Masked Autoencoders with Spatial-Attention Dropout for... code — Compare
TempleColor128 (1 row) AAA AAA: Adaptive Aggregation of Arbitrary Online Trackers with... code — Compare
VideoCube (1 row) RTracker-L RTracker: Recoverable Tracking via PN Tree Structured Memory code — Compare
VOT2014 (1 row) (unnamed in the archive) Efficient Counterfactual Learning from Bandit Feedback — — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

29 datasets whose archive record lists this task, ordered by the archive's paper count.

Subtasks archive 2025-07-28

1 subtask in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 187 papers with code (341 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections