{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/object-pose-estimation-via-the-aggregation-of","title":"Object Pose Estimation via the Aggregation of Diffusion Features","arxiv_id":"2403.18791","date":"2024-03-27","proceeding":"CVPR 2024 1","authors":["Tianfu Wang","Guosheng Hu","Hongguang Wang"],"abstract":"Estimating the pose of objects from images is a crucial task of 3D scene understanding, and recent approaches have shown promising results on very large benchmarks. However, these methods experience a significant performance drop when dealing with unseen objects. We believe that it results from the limited generalizability of image features. To address this problem, we have an in-depth analysis on the features of diffusion models, e.g. Stable Diffusion, which hold substantial potential for modeling unseen objects. Based on this analysis, we then innovatively introduce these diffusion features for object pose estimation. To achieve this, we propose three distinct architectures that can effectively capture and aggregate diffusion features of different granularity, greatly improving the generalizability of object pose estimation. Our approach outperforms the state-of-the-art methods by a considerable margin on three popular benchmark datasets, LM, O-LM, and T-LESS. In particular, our method achieves higher accuracy than the previous best arts on unseen objects: 97.9% vs. 93.5% on Unseen LM, 85.9% vs. 76.3% on Unseen O-LM, showing the strong generalizability of our method. Our code is released at https://github.com/Tianfu18/diff-feats-pose.","url_abs":"https://arxiv.org/abs/2403.18791v3","url_pdf":"https://arxiv.org/pdf/2403.18791v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"object-pose-estimation-via-the-aggregation-of","repo_url":"https://github.com/tianfu18/diff-feats-pose","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2403.18791","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2403.18791"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/Tianfu18/diff-feats-pose","reach":null}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"6157c2d41ff3f0ee","entry":"AggregationNetwork","repo":"Tianfu18/diff-feats-pose","repo_kind":"official","path":"lib/models/diffusion/aggregation_network.py","file_url":"https://github.com/Tianfu18/diff-feats-pose/blob/HEAD/lib/models/diffusion/aggregation_network.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6157c2d41ff3f0ee"}},{"code_sha256_prefix":"6e59b6ddbd63fc3e","entry":"FusionModule","repo":"Tianfu18/diff-feats-pose","repo_kind":"official","path":"lib/models/diffusion/aggregation_network.py","file_url":"https://github.com/Tianfu18/diff-feats-pose/blob/HEAD/lib/models/diffusion/aggregation_network.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6e59b6ddbd63fc3e"}},{"code_sha256_prefix":"ddb1b4b26e423b71","entry":"ResBlock","repo":"Tianfu18/diff-feats-pose","repo_kind":"official","path":"lib/models/diffusion/aggregation_network.py","file_url":"https://github.com/Tianfu18/diff-feats-pose/blob/HEAD/lib/models/diffusion/aggregation_network.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ddb1b4b26e423b71"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}