Papers › XNect: Real-time Multi-Person 3D Motion Capture with a Single RGB Camera

XNect: Real-time Multi-Person 3D Motion Capture with a Single RGB Camera

1 Jul 2019arXiv:1907.00837archive 2025-07-28

Dushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu, Mohamed Elgharib, Pascal Fua, Hans-Peter Seidel, Helge Rhodin, Gerard Pons-Moll, Christian Theobalt

We present a real-time approach for multi-person 3D motion capture at over 30 fps using a single RGB camera. It operates successfully in generic scenes which may contain occlusions by objects and by other people. Our method operates in subsequent stages. The first stage is a convolutional neural network (CNN) that estimates 2D and 3D pose features along with identity assignments for all visible joints of all individuals.We contribute a new architecture for this CNN, called SelecSLS Net, that uses novel selective long and short range skip connections to improve the information flow allowing for a drastically faster network without compromising accuracy. In the second stage, a fully connected neural network turns the possibly partial (on account of occlusion) 2Dpose and 3Dpose features for each subject into a complete 3Dpose estimate per individual. The third stage applies space-time skeletal model fitting to the predicted 2D and 3D pose per subject to further reconcile the 2D and 3D pose, and enforce temporal coherence. Our method returns the full skeletal pose in joint angles for each subject. This is a further key distinction from previous work that do not produce joint angle results of a coherent skeleton in real time for multi-person scenes. The proposed system runs on consumer hardware at a previously unseen speed of more than 30 fps given 512x320 images as input while achieving state-of-the-art accuracy, which we will demonstrate on a range of challenging real-world scenes.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mehtadushy/SelecSLS-Pytorch mentioned on GitHubpytorch report
osmr/imgclsmob mentioned on GitHubmxnetMIT report
rwightman/pytorch-image-models mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose Estimation3D Multi-Person Human Pose Estimation3D Multi-Person Pose EstimationMonocular 3D Human Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation MPI-INF-3DHP XNect (SelecSLS) AUC 45.3 #67 of 108 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP XNect (SelecSLS) MPJPE 98.4 #67 of 108 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP XNect (SelecSLS) PCK 82.8 #67 of 108 Archive leaderboard report
3D Multi-Person Pose Estimation MuPoTS-3D SelecSLS 3DPCK 75.8 #7 of 10 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M SelecSLS Average MPJPE (mm) 63.6 #33 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M SelecSLS Frames Needed 1 #33 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M SelecSLS Need Ground Truth 2D Pose No #33 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M SelecSLS Use Video Sequence No #33 of 52 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

SPEED

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections