{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maskfusion-real-time-recognition-tracking-and","title":"MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects","arxiv_id":"1804.09194","date":"2018-04-24","proceeding":null,"authors":["Martin Rünz","Maud Buffier","Lourdes Agapito"],"abstract":"We present MaskFusion, a real-time, object-aware, semantic and dynamic RGB-D\nSLAM system that goes beyond traditional systems which output a purely\ngeometric map of a static scene. MaskFusion recognizes, segments and assigns\nsemantic class labels to different objects in the scene, while tracking and\nreconstructing them even when they move independently from the camera.\n  As an RGB-D camera scans a cluttered scene, image-based instance-level\nsemantic segmentation creates semantic object masks that enable real-time\nobject recognition and the creation of an object-level representation for the\nworld map. Unlike previous recognition-based SLAM systems, MaskFusion does not\nrequire known models of the objects it can recognize, and can deal with\nmultiple independent motions. MaskFusion takes full advantage of using\ninstance-level semantic segmentation to enable semantic labels to be fused into\nan object-aware map, unlike recent semantics enabled SLAM systems that perform\nvoxel-level semantic segmentation. We show augmented-reality applications that\ndemonstrate the unique features of the map output by MaskFusion:\ninstance-aware, semantic and dynamic.","url_abs":"http://arxiv.org/abs/1804.09194v2","url_pdf":"http://arxiv.org/pdf/1804.09194v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"maskfusion-real-time-recognition-tracking-and","repo_url":"https://github.com/martinruenz/maskfusion","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"object-slam","task_name":"Object SLAM"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-slam","task_name":"Semantic SLAM"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"simultaneous-localization-and-mapping","task_name":"Simultaneous Localization and Mapping"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.09194","atlas_url":"https://app.syntology.ai/?focus=1804.09194","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}