{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-human-action-recognition","title":"Learning Human Action Recognition Representations Without Real Humans","arxiv_id":"2311.06231","date":"2023-11-10","proceeding":"NeurIPS 2023 11","authors":["Howard Zhong","Samarth Mishra","Donghyun Kim","SouYoung Jin","Rameswar Panda","Hilde Kuehne","Leonid Karlinsky","Venkatesh Saligrama","Aude Oliva","Rogerio Feris"],"abstract":"Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets contain images of people and hence are accompanied with issues related to privacy, ethics, and data protection, often preventing them from being publicly shared for reproducible research. Existing work has attempted to alleviate these problems by blurring faces, downsampling videos, or training on synthetic data. On the other hand, analysis on the transferability of privacy-preserving pre-trained models to downstream tasks has been limited. In this work, we study this problem by first asking the question: can we pre-train models for human action recognition with data that does not include real humans? To this end, we present, for the first time, a benchmark that leverages real-world videos with humans removed and synthetic data containing virtual humans to pre-train a model. We then evaluate the transferability of the representation learned on this data to a diverse set of downstream action recognition benchmarks. Furthermore, we propose a novel pre-training strategy, called Privacy-Preserving MAE-Align, to effectively combine synthetic data and human-removed real data. Our approach outperforms previous baselines by up to 5% and closes the performance gap between human and no-human action recognition representations on downstream tasks, for both linear probing and fine-tuning. Our benchmark, code, and models are available at https://github.com/howardzh01/PPMA .","url_abs":"https://arxiv.org/abs/2311.06231v1","url_pdf":"https://arxiv.org/pdf/2311.06231v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-human-action-recognition","repo_url":"https://github.com/howardzh01/ppma","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"ethics","task_name":"Ethics"},{"task_slug":"privacy-preserving","task_name":"Privacy Preserving"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2311.06231","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.06231"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/howardzh01/PPMA","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/howardzh01/ppma","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran":6,"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":8,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"eecc5dc23f84ce7c","entry":"Block","repo":"howardzh01/ppma","repo_kind":"official","path":"code/omnivision/models/vision_transformer.py","file_url":"https://github.com/howardzh01/ppma/blob/HEAD/code/omnivision/models/vision_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"eecc5dc23f84ce7c"}},{"code_sha256_prefix":"ee6120b797d547d3","entry":"PatchEmbed","repo":"howardzh01/ppma","repo_kind":"official","path":"code/omnivision/models/vision_transformer.py","file_url":"https://github.com/howardzh01/ppma/blob/HEAD/code/omnivision/models/vision_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ee6120b797d547d3"}},{"code_sha256_prefix":"c85bbc1f8ae5a2df","entry":"PatchEmbedGeneric","repo":"howardzh01/ppma","repo_kind":"official","path":"code/omnivision/models/vision_transformer.py","file_url":"https://github.com/howardzh01/ppma/blob/HEAD/code/omnivision/models/vision_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"c85bbc1f8ae5a2df"}},{"code_sha256_prefix":"35a743e50e17e66d","entry":"get_sinusoid_encoding_table","repo":"howardzh01/ppma","repo_kind":"official","path":"code/omnivision/models/vision_transformer.py","file_url":"https://github.com/howardzh01/ppma/blob/HEAD/code/omnivision/models/vision_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"35a743e50e17e66d"}},{"code_sha256_prefix":"de6af75fa2d1e4ad","entry":"init_parameters","repo":"howardzh01/PPMA","repo_kind":"official","path":"code/omnivision/model/model_init_utils.py","file_url":"https://github.com/howardzh01/PPMA/blob/HEAD/code/omnivision/model/model_init_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"de6af75fa2d1e4ad"}},{"code_sha256_prefix":"e627ac7e4b6f7972","entry":"layer_decay_param_modifier","repo":"howardzh01/PPMA","repo_kind":"official","path":"code/omnivision/optim/layer_decay_param_modifier.py","file_url":"https://github.com/howardzh01/PPMA/blob/HEAD/code/omnivision/optim/layer_decay_param_modifier.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e627ac7e4b6f7972"}},{"code_sha256_prefix":"b43c65d053fdd761","entry":"make_functional","repo":"howardzh01/PPMA","repo_kind":"official","path":"code/omnivision/model/model_wrappers.py","file_url":"https://github.com/howardzh01/PPMA/blob/HEAD/code/omnivision/model/model_wrappers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b43c65d053fdd761"}},{"code_sha256_prefix":"3e2d901955f04262","entry":"VisionTransformer","repo":"howardzh01/ppma","repo_kind":"official","path":"code/omnivision/models/vision_transformer.py","file_url":"https://github.com/howardzh01/ppma/blob/HEAD/code/omnivision/models/vision_transformer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"3e2d901955f04262"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}