Browse State-of-the-Art › Point Tracking
Point Tracking
61 papers with code · 8 benchmarks · 4 datasets archive 2025-07-28
Point Tracking, often referred to as Tracking any Point (TAP) involves acquiring, focusing on, and continuously tracking specific target point/points across video frames. The system identifies the target point, maintains focus, and predicts its movement, enabling smooth tracking even if the target moves unpredictably, or through occlusions. TAP has wide applications like object tracking, surveillance, and autonomous navigation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 61 papers with code (151 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 May 2023 7 repositories listed Syntology ran 7 of 16 samples · 9 unverified · 3 pointer-only (licence)Synthesizing visual content that meets users' needs often requires flexible and precise controllability of the pose, shape, expression, and layout of the generated objects.
-
27 Jul 2023 3 repositories listed Syntology ran 10 of 13 samples · 3 unverified · 8 pointer-only (licence)Our goal is to advance the state-of-the-art by placing emphasis on long videos with naturalistic motion.
-
14 Jun 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present a novel model for Tracking Any Point (TAP) that effectively tracks any queried point on any physical surface throughout a video sequence.
-
7 Nov 2022 3 repositories listed Syntology ran 4 of 12 samples · 8 unverified · 2 pointer-only (licence)Generic motion understanding from video involves not only tracking objects, but also perceiving how their surfaces deform and move.
-
23 Jul 2024 2 repositories listed Syntology ran 20 of 23 samples · 3 unverifiedDespite the impressive progress, research on metrics for evaluating the quality of generated videos, especially concerning temporal and motion consistency, remains underexplored.
-
22 Jul 2024 2 repositories listed Syntology ran 6 of 12 samples · 6 unverifiedWe introduce LocoTrack, a highly accurate and efficient model designed for the task of tracking any point (TAP) across video sequences.
-
8 Jul 2024 2 repositories listedWe introduce a new benchmark, TAPVid-3D, for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D).
-
1 Feb 2024 2 repositories listedTo endow models with greater understanding of physics and motion, it is useful to enable them to perceive how solid surfaces move and deform in real scenes.
-
14 Jul 2023 2 repositories listedWe introduce CoTracker, a transformer-based model that tracks a large number of 2D points in long video sequences.
-
16 Jul 2025 1 repository listedWe present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos.
-
15 Jul 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications.
-
18 May 2025 1 repository listedDrag-based editing within pretrained diffusion model provides a precise and flexible way to manipulate foreground objects.
-
20 Apr 2025 1 repository listed Syntology ran 2 of 8 samples · 6 unverifiedWe introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos.
-
20 Apr 2025 1 repository listedAccurate depth estimation from monocular videos remains challenging due to ambiguities inherent in single-view geometry, as crucial depth cues like stereopsis are absent.
-
8 Apr 2025 1 repository listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Furthermore, we introduce a temporal motion module for dynamic motions that ensures scale consistency across different frames and enhances performance in tasks requiring both precise geometry and reliable matching, most…
-
8 Apr 2025 1 repository listedTracking Any Point (TAP) in a video is a challenging computer vision problem with many demonstrated applications in robotics, video editing, and 3D reconstruction.
-
31 Mar 2025 1 repository listedUnderstanding tissue motion in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe-based scanning, and subtask autonomy.
-
14 Mar 2025 1 repository listedWe present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views.
-
13 Mar 2025 1 repository listedThe limits of agreement were wider for both CoTracker2 and EchoTracker, worse than the interobserver variability.
-
9 Mar 2025 1 repository listed Syntology ran 5 of 12 samples · 7 unverifiedDense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence of occlusions.
-
30 Jan 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedIn this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and…
-
21 Jan 2025 1 repository listedPoint tracking in videos is a fundamental task with applications in robotics, video editing, and more.
-
1 Jan 2025 1 repository listedInspired by this flexibility, we propose MATCHA, a unified feature model designed to "rule them all", establishing robust correspondences across diverse matching tasks.
-
15 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Most state-of-the-art point trackers are trained on synthetic data due to the difficulty of annotating real videos for this task.
-
24 Sep 2024 1 repository listed Syntology ran 4 of 8 samples · 4 unverifiedWe present a simple, self-supervised approach to the Tracking Any Point (TAP) problem.
-
21 Sep 2024 1 repository listedThis way, while basing our method on an amodal instance segmentation, we nevertheless obtain video-level amodal instance segmentation results.
-
9 Sep 2024 1 repository listedPoint tracking is a fundamental problem in computer vision with numerous applications in AR and robotics.
-
11 Aug 2024 1 repository listedWe propose OneShotLP, a training-free framework for video-based license plate detection and recognition, leveraging these advanced models.
-
30 Jul 2024 1 repository listedPoint tracking is a challenging task in computer vision, aiming to establish point-wise correspondence across long video sequences.
-
15 Jul 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedIn optical flow estimation, our method elevates a simple UNet to achieve state-of-the-art performance among self-supervised methods on the DSEC optical flow benchmark.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections