{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/affordance-grounding-from-demonstration-video-1","title":"Affordance Grounding from Demonstration Video to Target Image","arxiv_id":"2303.14644","date":"2023-03-26","proceeding":"CVPR 2023 1","authors":["Joya Chen","Difei Gao","Kevin Qinghong Lin","Mike Zheng Shou"],"abstract":"Humans excel at learning from expert demonstrations and solving their own problems. To equip intelligent robots and assistants, such as AR glasses, with this ability, it is essential to ground human hand interactions (i.e., affordances) from demonstration videos and apply them to a target image like a user's AR glass view. The video-to-image affordance grounding task is challenging due to (1) the need to predict fine-grained affordances, and (2) the limited training data, which inadequately covers video-image discrepancies and negatively impacts grounding. To tackle them, we propose Affordance Transformer (Afformer), which has a fine-grained transformer-based decoder that gradually refines affordance grounding. Moreover, we introduce Mask Affordance Hand (MaskAHand), a self-supervised pre-training technique for synthesizing video-image data and simulating context changes, enhancing affordance grounding across video-image discrepancies. Afformer with MaskAHand pre-training achieves state-of-the-art performance on multiple benchmarks, including a substantial 37% improvement on the OPRA dataset. Code is made available at https://github.com/showlab/afformer.","url_abs":"https://arxiv.org/abs/2303.14644v1","url_pdf":"https://arxiv.org/pdf/2303.14644v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"affordance-grounding-from-demonstration-video-1","repo_url":"https://github.com/showlab/afformer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"video-to-image-affordance-grounding","task_name":"Video-to-image Affordance Grounding"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-to-image-affordance-grounding-on-epic","task":"Video-to-image Affordance Grounding","dataset":"EPIC-Hotspot","model":"Afformer","rank_in_archive_order":1,"of":3,"metrics":{"AUC-J":"0.88","KLD":"0.97","SIM":"0.56"},"uses_additional_data":false},{"leaderboard":"/sota/video-to-image-affordance-grounding-on-opra","task":"Video-to-image Affordance Grounding","dataset":"OPRA","model":"Afformer (ViTDet-B encoder)","rank_in_archive_order":1,"of":3,"metrics":{"KLD":"1.51","Top-1 Action Accuracy":"52.27"},"uses_additional_data":false},{"leaderboard":"/sota/video-to-image-affordance-grounding-on-opra","task":"Video-to-image Affordance Grounding","dataset":"OPRA","model":"Afformer (ResNet-50-FPN encoder)","rank_in_archive_order":2,"of":3,"metrics":{"KLD":"1.55","Top-1 Action Accuracy":"52.14"},"uses_additional_data":false},{"leaderboard":"/sota/video-to-image-affordance-grounding-on-opra-1","task":"Video-to-image Affordance Grounding","dataset":"OPRA (28x28)","model":"Afformer","rank_in_archive_order":1,"of":4,"metrics":{"AUC-J":"0.89","KLD":"1.05","SIM":"0.53"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2303.14644","atlas_url":"https://app.syntology.ai/?focus=2303.14644","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2303.14644"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/showlab/afformer","reach":null}],"summary":{"ran":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"1fa04425ef20ddae","entry":"Afformer","repo":"showlab/afformer","repo_kind":"official","path":"afformer/afformer.py","file_url":"https://github.com/showlab/afformer/blob/HEAD/afformer/afformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1fa04425ef20ddae"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}