{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vnect-real-time-3d-human-pose-estimation-with","title":"VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera","arxiv_id":"1705.01583","date":"2017-05-03","proceeding":null,"authors":["Dushyant Mehta","Srinath Sridhar","Oleksandr Sotnychenko","Helge Rhodin","Mohammad Shafiei","Hans-Peter Seidel","Weipeng Xu","Dan Casas","Christian Theobalt"],"abstract":"We present the first real-time method to capture the full global 3D skeletal\npose of a human in a stable, temporally consistent manner using a single RGB\ncamera. Our method combines a new convolutional neural network (CNN) based pose\nregressor with kinematic skeleton fitting. Our novel fully-convolutional pose\nformulation regresses 2D and 3D joint positions jointly in real time and does\nnot require tightly cropped input frames. A real-time kinematic skeleton\nfitting method uses the CNN output to yield temporally stable 3D global pose\nreconstructions on the basis of a coherent kinematic skeleton. This makes our\napproach the first monocular RGB method usable in real-time applications such\nas 3D character control---thus far, the only monocular methods for such\napplications employed specialized RGB-D cameras. Our method's accuracy is\nquantitatively on par with the best offline 3D monocular RGB pose estimation\nmethods. Our results are qualitatively comparable to, and sometimes better\nthan, results from monocular RGB-D approaches, such as the Kinect. However, we\nshow that our approach is more broadly applicable than RGB-D solutions, i.e. it\nworks for outdoor scenes, community videos, and low quality commodity RGB\ncameras.","url_abs":"http://arxiv.org/abs/1705.01583v1","url_pdf":"http://arxiv.org/pdf/1705.01583v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vnect-real-time-3d-human-pose-estimation-with","repo_url":"https://github.com/XinArkh/VNect","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"VNect (Augm.)","rank_in_archive_order":82,"of":108,"metrics":{"AUC":"40.4","MPJPE":"124.7","PCK":"76.6"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-mpi-inf-3dhp","task":"3D Human Pose Estimation","dataset":"MPI-INF-3DHP","model":"VNect (ResNet 50 GT)","rank_in_archive_order":96,"of":108,"metrics":{"AUC":"41.6","PCK":"79.4"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-leeds-sports-poses","task":"Pose Estimation","dataset":"Leeds Sports Poses","model":"VNect (ResNet 50)","rank_in_archive_order":16,"of":18,"metrics":{"PCK":"79.4"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}