Browse State-of-the-Art › Pose Tracking
Pose Tracking
76 papers with code · 3 benchmarks · 14 datasets archive 2025-07-28
Pose Tracking is the task of estimating multi-person human poses in videos and assigning unique instance IDs for each keypoint across frames. Accurate estimation of human keypoint-trajectories is useful for human action recognition, human interaction understanding, motion capture and animation.
Source: LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| PoseTrack2017 (10 rows) | DetTrack | Combining detection and tracking for human pose estimation in videos | — | — | Compare |
| PoseTrack2018 (5 rows) | DetTrack | Combining detection and tracking for human pose estimation in videos | — | — | Compare |
| Multi-Person PoseTrack (1 row) | PoseTrack | PoseTrack: Joint Multi-Person Pose Estimation and Tracking | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
14 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 76 papers with code (191 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Feb 2019 39 repositories listed Syntology ran 8 of 25 samples · 17 unverifiedWe start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel.
-
17 Apr 2018 27 repositories listedThere has been significant progress on pose estimation and increasing interests on pose tracking in recent years.
-
17 Jun 2020 7 repositories listedWe present BlazePose, a lightweight convolutional neural network architecture for human pose estimation that is tailored for real-time inference on mobile devices.
-
25 Jul 2024 2 repositories listedInspired by recent work on prompting in vision, we introduce Keypoint Promptable ReID (KPR), a novel formulation of the ReID problem that explicitly complements the input bounding box with a set of semantic keypoints…
-
7 May 2024 2 repositories listedIn this paper, we improve our previous direct pipeline \textit{Event-based Stereo Visual Odometry} in terms of accuracy and efficiency.
-
30 Jan 2022 2 repositories listed Syntology ran 4 of 30 samples · 26 unverifiedThe canonical object representation is learned solely in simulation and then used to parse a category-level, task trajectory from a single demonstration video.
-
6 Nov 2021 2 repositories listedIn this work, we introduce ROFT, a Kalman filtering approach for 6D object pose and velocity tracking from a stream of RGB-D images.
-
25 Oct 2021 2 repositories listedFinally, we use a pre-rendered sparse viewpoint model to create a joint posterior probability for the object pose.
-
11 Apr 2021 2 repositories listedIn this work, we focus our attention on the similarity among works of art based on human poses and the actions they represent, moving from the concept of Pathosformel in Aby Warburg.
-
23 Oct 2019 2 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedWe present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data.
-
7 May 2019 2 repositories listed Syntology ran 2 of 16 samples · 14 unverifiedTo the best of our knowledge, this is the first paper to propose an online human pose tracking framework in a top-down fashion.
-
2 Apr 2019 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)We introduce multigrid Predictive Filter Flow (mgPFF), a framework for unsupervised learning on videos.
-
27 Oct 2017 2 repositories listedIn this work, we aim to further advance the state of the art by establishing "PoseTrack", a new large-scale benchmark for video-based human pose estimation and articulated tracking, and bringing together the community…
-
3 Apr 2017 2 repositories listedIn this work, we propose a framework for hand tracking that can capture the motion of two interacting hands using only a single, inexpensive RGB-D camera.
-
23 Nov 2016 2 repositories listedIn this work, we introduce the challenging problem of joint multi-person pose estimation and tracking of an unknown number of persons in unconstrained videos.
-
7 Oct 2015 2 repositories listedEvent-based vision sensors mimic the operation of biological retina and they represent a major paradigm shift from traditional cameras.
-
20 Jun 2025 1 repository listedWe introduce a robust framework, RGBTrack, for real-time 6D pose estimation and tracking that operates solely on RGB data, thereby eliminating the need for depth input for such dynamic and precise object pose tracking…
-
15 May 2025 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedBuilding on SDHA, we further analyze various spike-driven space-time attention designs and identify an optimal scheme that delivers appealing performance for video tasks, while maintaining only linear temporal…
-
7 Apr 2025 1 repository listedSimultaneous localization and mapping (SLAM) technology now has photorealistic mapping capabilities thanks to the real-time high-fidelity rendering capability of 3D Gaussian splatting (3DGS).
-
25 Mar 2025 1 repository listedIn the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent.
-
17 Mar 2025 1 repository listedTo address the standard on-orbit observation missions, we propose a line-based pose tracking method for uncooperative spacecraft utilizing a stereo event camera.
-
27 Nov 2024 1 repository listedOur results demonstrate the effectiveness of G3Flow in enhancing real-time dynamic semantic feature understanding for robotic manipulation policies.
-
21 Nov 2024 1 repository listedIn this work, we present a novel light field segmentation method that adapts SAM 2 to the light field domain without retraining or modifying the model.
-
12 Oct 2024 1 repository listedTo this end, a compact back-end is proposed for continuously updating the IMU bias and predicting the linear velocity, enabling an accurate motion prediction for camera pose tracking.
-
28 Sep 2024 1 repository listedReliable self-localization is a foundational skill for many intelligent mobile platforms.
-
11 Jul 2024 1 repository listedTwo-view pose estimation is essential for map-free visual relocalization and object pose tracking tasks.
-
10 Jun 2024 1 repository listedWe introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos.
-
19 Mar 2024 1 repository listedWe propose a dense RGBD SLAM system based on 3D Gaussian Splatting that provides metrically accurate pose tracking and visually realistic reconstruction.
-
29 Feb 2024 1 repository listedIn this paper, we propose a new approach termed as \textbf{VideoMAC}, which combines video masked autoencoders with resource-friendly ConvNets.
-
9 Jan 2024 1 repository listedIn this paper, we investigate the real-world robot task of aerial vision guidance for aerial robotics manipulation, utilizing category-level 6-DoF pose tracking.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections