{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/not-all-tokens-are-equal-human-centric-visual","title":"Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer","arxiv_id":"2204.08680","date":"2022-04-19","proceeding":"CVPR 2022 1","authors":["Wang Zeng","Sheng Jin","Wentao Liu","Chen Qian","Ping Luo","Wanli Ouyang","Xiaogang Wang"],"abstract":"Vision transformers have achieved great successes in many computer vision tasks. Most methods generate vision tokens by splitting an image into a regular and fixed grid and treating each cell as a token. However, not all regions are equally important in human-centric vision tasks, e.g., the human body needs a fine representation with many tokens, while the image background can be modeled by a few tokens. To address this problem, we propose a novel Vision Transformer, called Token Clustering Transformer (TCFormer), which merges tokens by progressive clustering, where the tokens can be merged from different locations with flexible shapes and sizes. The tokens in TCFormer can not only focus on important areas but also adjust the token shapes to fit the semantic concept and adopt a fine resolution for regions containing critical details, which is beneficial to capturing detailed information. Extensive experiments show that TCFormer consistently outperforms its counterparts on different challenging human-centric tasks and datasets, including whole-body pose estimation on COCO-WholeBody and 3D human mesh reconstruction on 3DPW. Code is available at https://github.com/zengwang430521/TCFormer.git","url_abs":"https://arxiv.org/abs/2204.08680v3","url_pdf":"https://arxiv.org/pdf/2204.08680v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"not-all-tokens-are-equal-human-centric-visual","repo_url":"https://github.com/zengwang430521/tcformer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"2d-human-pose-estimation","task_name":"2D Human Pose Estimation"},{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"all","task_name":"All"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/2d-human-pose-estimation-on-coco-wholebody-1","task":"2D Human Pose Estimation","dataset":"COCO-WholeBody","model":"TCFormer","rank_in_archive_order":5,"of":15,"metrics":{"WB":"64.2","body":"71.8","face":"79.0","foot":"74.4","hand":"61.4"},"uses_additional_data":true},{"leaderboard":"/sota/3d-human-pose-estimation-on-3dpw","task":"3D Human Pose Estimation","dataset":"3DPW","model":"TCFormer","rank_in_archive_order":71,"of":119,"metrics":{"MPJPE":"80.6","PA-MPJPE":"49.3"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2204.08680","atlas_url":"https://app.syntology.ai/?focus=2204.08680","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2204.08680"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zengwang430521/tcformer","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":6,"unverified":1},"by_repo_kind":{"official":{"samples":7,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e3c853a7511a411a","entry":"drop_block_2d","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/transformer_utils.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/transformer_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e3c853a7511a411a"}},{"code_sha256_prefix":"4162565de6ffa9da","entry":"drop_block_fast_2d","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/transformer_utils.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/transformer_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4162565de6ffa9da"}},{"code_sha256_prefix":"9d3c18b6f3ac3dc1","entry":"drop_path","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/transformer_utils.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/transformer_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9d3c18b6f3ac3dc1"}},{"code_sha256_prefix":"428c8cc732ef3657","entry":"flops_to_string","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/custom_flops_counter.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/custom_flops_counter.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"428c8cc732ef3657"}},{"code_sha256_prefix":"306896b703835e2c","entry":"get_grid_index","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/tcformer_utils.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/tcformer_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"306896b703835e2c"}},{"code_sha256_prefix":"e0247439ad10578e","entry":"params_to_string","repo":"zengwang430521/tcformer","repo_kind":"official","path":"tcformer_module/custom_flops_counter.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/tcformer_module/custom_flops_counter.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e0247439ad10578e"}},{"code_sha256_prefix":"3bc74137ab36aa79","entry":"build_transform","repo":"zengwang430521/tcformer","repo_kind":"official","path":"classification/datasets.py","file_url":"https://github.com/zengwang430521/tcformer/blob/HEAD/classification/datasets.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3bc74137ab36aa79"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}