Papers › Lightweight Multi-View 3D Pose Estimation through Camera-Disentangled Representation
Lightweight Multi-View 3D Pose Estimation through Camera-Disentangled Representation
Edoardo Remelli, Shangchen Han, Sina Honari, Pascal Fua, Robert Wang
We present a lightweight solution to recover 3D pose from multi-view images captured with spatially calibrated cameras. Building upon recent advances in interpretable representation learning, we exploit 3D geometry to fuse input images into a unified latent representation of pose, which is disentangled from camera view-points. This allows us to reason effectively about 3D pose across different views without using compute-intensive volumetric grids. Our architecture then conditions the learned representation on camera projection operators to produce accurate per-view 2d detections, that can be simply lifted to 3D via a differentiable Direct Linear Transform (DLT) layer. In order to do it efficiently, we propose a novel implementation of DLT that is orders of magnitude faster on GPU architectures than standard SVD-based triangulation methods. We evaluate our approach on two large-scale human pose datasets (H36M and Total Capture): our method outperforms or performs comparably to the state-of-the-art volumetric methods, while, unlike them, yielding real-time performance.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | Human3.6M | LWCDR | Average MPJPE (mm) | 30.2 | #8 of 88 | Archive leaderboard | report |
| 3D Human Pose Estimation | Human3.6M | LWCDR | Multi-View or Monocular | Multi-View | #8 of 88 | Archive leaderboard | report |
| 3D Human Pose Estimation | Human3.6M | LWCDR | Using 2D ground-truth joints | No | #8 of 88 | Archive leaderboard | report |
| 3D Human Pose Estimation | Total Capture | LWCDR | Average MPJPE (mm) | 27.5 | #4 of 14 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections