Papers › VTP: Volumetric Transformer for Multi-view Multi-person 3D Pose Estimation

VTP: Volumetric Transformer for Multi-view Multi-person 3D Pose Estimation

25 May 2022arXiv:2205.12602archive 2025-07-28

Yuxing Chen, Renshu Gu, Ouhan Huang, Gangyong Jia

This paper presents Volumetric Transformer Pose estimator (VTP), the first 3D volumetric transformer framework for multi-view multi-person 3D human pose estimation. VTP aggregates features from 2D keypoints in all camera views and directly learns the spatial relationships in the 3D voxel space in an end-to-end fashion. The aggregated 3D features are passed through 3D convolutions before being flattened into sequential embeddings and fed into a transformer. A residual structure is designed to further improve the performance. In addition, the sparse Sinkhorn attention is empowered to reduce the memory cost, which is a major bottleneck for volumetric representations, while also achieving excellent performance. The output of the transformer is again concatenated with 3D convolutional features by a residual design. The proposed VTP framework integrates the high performance of the transformer with volumetric representations, which can be used as a good alternative to the convolutional backbones. Experiments on the Shelf, Campus and CMU Panoptic benchmarks show promising results in terms of both Mean Per Joint Position Error (MPJPE) and Percentage of Correctly estimated Parts (PCP). Our code will be available.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose Estimation3D Multi-Person Pose Estimation3D Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation Panoptic VTP Average MPJPE (mm) 17.62 #4 of 9 Archive leaderboard report
3D Multi-Person Pose Estimation Campus VTP Mean mAP 80.1 #12 of 16 Archive leaderboard report
3D Multi-Person Pose Estimation Campus VTP PCP3D 96.3 #12 of 16 Archive leaderboard report
3D Multi-Person Pose Estimation Shelf VTP MPJPE 56.3 #15 of 27 Archive leaderboard report
3D Multi-Person Pose Estimation Shelf VTP PCP3D 97.3 #15 of 27 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutFeedforward NetworkLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerReLUResidual ConnectionSoftmaxSparse Sinkhorn AttentionTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections