{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tessetrack-end-to-end-learnable-multi-person","title":"TesseTrack: End-to-End Learnable Multi-Person Articulated 3D Pose Tracking","arxiv_id":null,"date":"2021-06-16","proceeding":"CVPR 2021 1","authors":["N. Dinesh Reddy","Laurent Guigues","Leonid Pischulini","Jayan Eledath","Srinivasa Narasimhan"],"abstract":"We consider the task of 3D pose estimation and tracking of multiple people seen in an arbitrary number of camera feeds. We propose TesseTrack, a novel top-down approach that simultaneously reasons about multiple individuals’ 3D body joint reconstructions and associations in space and time in a single end-to-end learnable framework. At the core of our approach is a novel spatio-temporal formulation that operates in a common voxelized feature space aggregated from single- or multiple camera views. After a person detection step, a 4D CNN produces short-term person-specific representations which are then linked across time by a differentiable matcher. The linked descriptions are then merged and deconvolved into 3D poses. This joint spatio-temporal formulation contrasts with previous piece-wise strategies that treat 2D pose estimation, 2D-to-3D lifting, and 3D pose tracking as independent sub-problems that are error-prone when solved in isolation. Furthermore, unlike previous methods, TesseTrack is robust to changes in the number of camera views and achieves very good results even if a single view is available at inference time. Quantitative evaluation of 3D pose reconstruction accuracy on standard benchmarks shows significant improvements over the state of the art. Evaluation of multi-person articulated 3D pose tracking in our novel evaluation framework demonstrates the superiority of TesseTrack over strong baselines.","url_abs":"http://www.cs.cmu.edu/~ILIM/projects/IM/TesseTrack/","url_pdf":"https://dineshreddy91.github.io/papers/TesseTrack.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"2d-pose-estimation","task_name":"2D Pose Estimation"},{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"3d-human-pose-tracking","task_name":"3D Human Pose Tracking"},{"task_slug":"3d-multi-object-tracking","task_name":"3D Multi-Object Tracking"},{"task_slug":"3d-multi-person-pose-estimation","task_name":"3D Multi-Person Pose Estimation"},{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"human-detection","task_name":"Human Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-tracking","task_name":"Pose Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-human36m","task":"3D Human Pose Estimation","dataset":"Human3.6M","model":"TesseTrack (Monocular)","rank_in_archive_order":39,"of":88,"metrics":{"Average MPJPE (mm)":"44.6","Multi-View or Monocular":"Monocular","Using 2D ground-truth joints":"No"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-cmu-panoptic","task":"3D Human Pose Estimation","dataset":"Panoptic","model":"TesseTrack Multi-View (5 views)","rank_in_archive_order":1,"of":9,"metrics":{"Average MPJPE (mm)":"7.3"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-pose-estimation-on-cmu-panoptic","task":"3D Human Pose Estimation","dataset":"Panoptic","model":"TesseTrack Monocular","rank_in_archive_order":5,"of":9,"metrics":{"Average MPJPE (mm)":"18.9"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-pose-tracking-on-cmu-panoptic","task":"3D Human Pose Tracking","dataset":"Panoptic","model":"TesseTrack","rank_in_archive_order":1,"of":1,"metrics":{"3DMOTA":"94.1"},"uses_additional_data":false},{"leaderboard":"/sota/3d-multi-person-pose-estimation-on-campus","task":"3D Multi-Person Pose Estimation","dataset":"Campus","model":"TesseTrack","rank_in_archive_order":1,"of":16,"metrics":{"PCP3D":"97.4"},"uses_additional_data":false},{"leaderboard":"/sota/3d-multi-person-pose-estimation-on-cmu","task":"3D Multi-Person Pose Estimation","dataset":"Panoptic","model":"TesseTrack","rank_in_archive_order":1,"of":20,"metrics":{"Average MPJPE (mm)":"7.3"},"uses_additional_data":true},{"leaderboard":"/sota/3d-multi-person-pose-estimation-on-shelf","task":"3D Multi-Person Pose Estimation","dataset":"Shelf","model":"TesseTrack (correct)","rank_in_archive_order":6,"of":27,"metrics":{"PCP3D":"97.9"},"uses_additional_data":false},{"leaderboard":"/sota/3d-pose-estimation-on-human3-6m","task":"3D Pose Estimation","dataset":"Human3.6M","model":"TesseTrack","rank_in_archive_order":1,"of":4,"metrics":{"Average MPJPE (mm)":"18.7"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}