Papers › IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation

IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation

6 Aug 2022arXiv:2208.03431archive 2025-07-28

Zhongwei Qiu, Qiansheng Yang, Jian Wang, Dongmei Fu

Video 3D human pose estimation aims to localize the 3D coordinates of human joints from videos. Recent transformer-based approaches focus on capturing the spatiotemporal information from sequential 2D poses, which cannot model the contextual depth feature effectively since the visual depth features are lost in the step of 2D pose estimation. In this paper, we simplify the paradigm into an end-to-end framework, Instance-guided Video Transformer (IVT), which enables learning spatiotemporal contextual depth information from visual features effectively and predicts 3D poses directly from video frames. In particular, we firstly formulate video frames as a series of instance-guided tokens and each token is in charge of predicting the 3D pose of a human instance. These tokens contain body structure information since they are extracted by the guidance of joint offsets from the human center to the corresponding body joints. Then, these tokens are sent into IVT for learning spatiotemporal contextual depth. In addition, we propose a cross-scale instance-guided attention mechanism to handle the variational scales among multiple persons. Finally, the 3D poses of each person are decoded from instance-guided tokens by coordinate regression. Experiments on three widely-used 3D pose estimation benchmarks show that the proposed IVT achieves state-of-the-art performances.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

2D Pose Estimation3D Human Pose Estimation3D Multi-Person Pose Estimation3D Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation 3DPW IVT (f=5) PA-MPJPE 46 #69 of 119 Archive leaderboard report
3D Human Pose Estimation Human3.6M IVT (f=5) Average MPJPE (mm) 40.2 #19 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M IVT (f=5) Multi-View or Monocular Monocular #19 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M IVT (f=5) Using 2D ground-truth joints No #19 of 88 Archive leaderboard report
3D Multi-Person Pose Estimation Panoptic IVT (f=5) Average MPJPE (mm) 48.4 #11 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections