{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/open-set-3d-semantic-instance-maps-for-vision","title":"Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM","arxiv_id":"2404.17922","date":"2024-04-27","proceeding":null,"authors":["Laksh Nanwani","Kumaraditya Gupta","Aditya Mathur","Swayam Agrawal","A. H. Abdul Hafez","K. Madhava Krishna"],"abstract":"Humans excel at forming mental maps of their surroundings, equipping them to understand object relationships and navigate based on language queries. Our previous work SI Maps [1] showed that having instance-level information and the semantic understanding of an environment helps significantly improve performance for language-guided tasks. We extend this instance-level approach to 3D while increasing the pipeline's robustness and improving quantitative and qualitative results. Our method leverages foundational models for object recognition, image segmentation, and feature extraction. We propose a representation that results in a 3D point cloud map with instance-level embeddings, which bring in the semantic understanding that natural language commands can query. Quantitatively, the work improves upon the success rate of language-guided tasks. At the same time, we qualitatively observe the ability to identify instances more clearly and leverage the foundational models and language and image-aligned embeddings to identify objects that, otherwise, a closed-set approach wouldn't be able to identify.","url_abs":"https://arxiv.org/abs/2404.17922v1","url_pdf":"https://arxiv.org/pdf/2404.17922v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"open-set-3d-semantic-instance-maps-for-vision","repo_url":"https://github.com/Smart-Wheelchair-RRC/o3d-sim","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"vision-language-navigation","task_name":"Vision-Language Navigation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}