{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2512-07599","title":"Online Segment Any 3D Thing as Instance Tracking","arxiv_id":"2512.07599","date":"2025-12-08","proceeding":"NeurIPS","authors":["Hanshi Wang","Zijian Cai","Jin Gao","Yiwei Zhang","Weiming Hu","Ke Wang","Zhipeng Zhang"],"abstract":"Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VFMs) outputs that are lifted into 3D point clouds, facilitating spatial information propagation through inter-query interactions. Nevertheless, perception is an inherently dynamic process, rendering temporal understanding a critical yet overlooked dimension within these prevailing query-based pipelines. Therefore, to further unlock the temporal environmental perception capabilities of embodied agents, our work reconceptualizes online 3D segmentation as an instance tracking problem (AutoSeg3D). Our core strategy involves utilizing object queries for temporal information propagation, where long-term instance association promotes the coherence of features and object identities, while short-term instance update enriches instant observations. Given that viewpoint variations in embodied robotics often lead to partial object visibility across frames, this mechanism aids the model in developing a holistic object understanding beyond incomplete instantaneous views. Furthermore, we introduce spatial consistency learning to mitigate the fragmentation problem inherent in VFMs, yielding more comprehensive instance information for enhancing the efficacy of both long-term and short-term temporal learning. The temporal information exchange and consistency learning facilitated by these sparse object queries not only enhance spatial comprehension but also circumvent the computational burden associated with dense temporal point cloud interactions. Our method establishes a new state-of-the-art, surpassing ESAM by 2.8 AP on ScanNet200 and delivering consistent gains on ScanNet, SceneNN, and 3RScan datasets.","url_abs":"https://arxiv.org/abs/2512.07599","url_pdf":"https://arxiv.org/pdf/2512.07599","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2512.07599","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2512.07599"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/AutoLab-SAI-SJTU/AutoSeg3D","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":4},"by_repo_kind":{"found_in_text":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"81a52f12291c03c6","entry":"adjust_intrinsic","repo":"AutoLab-SAI-SJTU/AutoSeg3D","repo_kind":"found_in_text","path":"oneformer3d/formatting.py","file_url":"https://github.com/AutoLab-SAI-SJTU/AutoSeg3D/blob/HEAD/oneformer3d/formatting.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"81a52f12291c03c6"}},{"code_sha256_prefix":"dfcc5c5f08448dec","entry":"build_pairwise_mask","repo":"AutoLab-SAI-SJTU/AutoSeg3D","repo_kind":"found_in_text","path":"oneformer3d/dq_utils.py","file_url":"https://github.com/AutoLab-SAI-SJTU/AutoSeg3D/blob/HEAD/oneformer3d/dq_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dfcc5c5f08448dec"}},{"code_sha256_prefix":"0d03cb242852337e","entry":"cluster_with_threshold","repo":"AutoLab-SAI-SJTU/AutoSeg3D","repo_kind":"found_in_text","path":"oneformer3d/dq_utils.py","file_url":"https://github.com/AutoLab-SAI-SJTU/AutoSeg3D/blob/HEAD/oneformer3d/dq_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0d03cb242852337e"}},{"code_sha256_prefix":"f47c8e54943ce833","entry":"make_intrinsic","repo":"AutoLab-SAI-SJTU/AutoSeg3D","repo_kind":"found_in_text","path":"oneformer3d/formatting.py","file_url":"https://github.com/AutoLab-SAI-SJTU/AutoSeg3D/blob/HEAD/oneformer3d/formatting.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f47c8e54943ce833"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CV","source":"arxiv_api"},"syntology_extracted_results":null}