Papers › XFormer: Fast and Accurate Monocular 3D Body Capture
XFormer: Fast and Accurate Monocular 3D Body Capture
Lihui Qian, Xintong Han, Faqiang Wang, Hongyu Liu, Haoye Dong, Zhiwen Li, Huawei Wei, Zhe Lin, Cheng-Bin Jin
We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an image branch that makes predictions directly from the RGB image features. At the core of our method is a cross-modal transformer block that allows information to flow across these two branches by modeling the attention between 2D keypoint coordinates and image spatial features. Our architecture is smartly designed, which enables us to train on various types of datasets including images with 2D/3D annotations, images with 3D pseudo labels, and motion capture datasets that do not have associated images. This effectively improves the accuracy and generalization ability of our system. Built on a lightweight backbone (MobileNetV3), our method runs blazing fast (over 30fps on a single CPU core) and still yields competitive accuracy. Furthermore, with an HRNet backbone, XFormer delivers state-of-the-art performance on Huamn3.6 and 3DPW datasets.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | 3DPW | XFormer (HRNet) | MPJPE | 75 | #34 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | XFormer (HRNet) | MPVPE | 87.1 | #34 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | XFormer (HRNet) | PA-MPJPE | 45.7 | #34 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | XFormer (HRNet) | MPJPE | 109.8 | #75 of 108 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | XFormer (HRNet) | PA-MPJPE | 64.5 | #75 of 108 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections