{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/neural-scene-decomposition-for-multi-person","title":"Neural Scene Decomposition for Multi-Person Motion Capture","arxiv_id":"1903.05684","date":"2019-03-13","proceeding":"CVPR 2019 6","authors":["Helge Rhodin","Victor Constantin","Isinsu Katircioglu","Mathieu Salzmann","Pascal Fua"],"abstract":"Learning general image representations has proven key to the success of many\ncomputer vision tasks. For example, many approaches to image understanding\nproblems rely on deep networks that were initially trained on ImageNet, mostly\nbecause the learned features are a valuable starting point to learn from\nlimited labeled data. However, when it comes to 3D motion capture of multiple\npeople, these features are only of limited use.\n  In this paper, we therefore propose an approach to learning features that are\nuseful for this purpose. To this end, we introduce a self-supervised approach\nto learning what we call a neural scene decomposition (NSD) that can be\nexploited for 3D pose estimation. NSD comprises three layers of abstraction to\nrepresent human subjects: spatial layout in terms of bounding-boxes and\nrelative depth; a 2D shape representation in terms of an instance segmentation\nmask; and subject-specific appearance and 3D pose information. By exploiting\nself-supervision coming from multiview data, our NSD model can be trained\nend-to-end without any 2D or 3D supervision. In contrast to previous\napproaches, it works for multiple persons and full-frame images. Because it\nencodes 3D geometry, NSD can then be effectively leveraged to train a 3D pose\nestimation network from small amounts of annotated data.","url_abs":"http://arxiv.org/abs/1903.05684v1","url_pdf":"http://arxiv.org/pdf/1903.05684v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"neural-scene-decomposition-for-multi-person","repo_url":"https://github.com/hrhodin/NeuralSceneDecomposition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"3d-geometry","task_name":"3D geometry"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1903.05684","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}