{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/videomae-masked-autoencoders-are-data-1","title":"VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training","arxiv_id":"2203.12602","date":"2022-03-23","proceeding":null,"authors":["Zhan Tong","Yibing Song","Jue Wang","LiMin Wang"],"abstract":"Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets. In this paper, we show that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP). We are inspired by the recent ImageMAE and propose customized video tube masking with an extremely high ratio. This simple design makes video reconstruction a more challenging self-supervision task, thus encouraging extracting more effective video representations during this pre-training process. We obtain three important findings on SSVP: (1) An extremely high proportion of masking ratio (i.e., 90% to 95%) still yields favorable performance of VideoMAE. The temporally redundant video content enables a higher masking ratio than that of images. (2) VideoMAE achieves impressive results on very small datasets (i.e., around 3k-4k videos) without using any extra data. (3) VideoMAE shows that data quality is more important than data quantity for SSVP. Domain shift between pre-training and target datasets is an important issue. Notably, our VideoMAE with the vanilla ViT can achieve 87.4% on Kinetics-400, 75.4% on Something-Something V2, 91.3% on UCF101, and 62.6% on HMDB51, without using any extra data. Code is available at https://github.com/MCG-NJU/VideoMAE.","url_abs":"https://arxiv.org/abs/2203.12602v3","url_pdf":"https://arxiv.org/pdf/2203.12602v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MCG-NJU/VideoMAE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/huggingface/transformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/innat/VideoMAE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MS-P3/code7/tree/main/videomae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MindCode-4/code-1/tree/main/videomae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MindSpore-scientific/code-13/tree/main/token_learner","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/MindSpore-scientific/code-7/tree/main/VideoMAE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"videomae-masked-autoencoders-are-data-1","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/5/videomae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"4k","task_name":"4k"},{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"self-supervised-action-recognition","task_name":"Self-Supervised Action Recognition"},{"task_slug":"self-supervised-action-recognition-linear","task_name":"Self-Supervised Action Recognition Linear"},{"task_slug":"video-reconstruction","task_name":"Video Reconstruction"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"VideoMAE (no extra data, ViT-H, 32x320x320)","rank_in_archive_order":36,"of":207,"metrics":{"Acc@1":"87.4","Acc@5":"97.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"VideoMAE (no extra data, ViT-H)","rank_in_archive_order":46,"of":207,"metrics":{"Acc@1":"86.6","Acc@5":"97.1"},"uses_additional_data":false},{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"VideoMAE (no extra data, ViT-L, 32x320x320)","rank_in_archive_order":50,"of":207,"metrics":{"Acc@1":"86.1","Acc@5":"97.3"},"uses_additional_data":false},{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"VideoMAE (no extra data, ViT-L, 16x4)","rank_in_archive_order":57,"of":207,"metrics":{"Acc@1":"85.2","Acc@5":"96.8"},"uses_additional_data":false},{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"VideoMAE (no extra data, ViT-B, 16x4)","rank_in_archive_order":82,"of":207,"metrics":{"Acc@1":"81.5","Acc@5":"95.1"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain+finetune, ViT-H, 16x4)","rank_in_archive_order":10,"of":38,"metrics":{"mAP":"39.5"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K700 pretrain+finetune, ViT-L, 16x4)","rank_in_archive_order":11,"of":38,"metrics":{"mAP":"39.3"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain+finetune, ViT-L, 16x4)","rank_in_archive_order":13,"of":38,"metrics":{"mAP":"37.8"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain, ViT-H, 16x4)","rank_in_archive_order":15,"of":38,"metrics":{"mAP":"36.5"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K700 pretrain, ViT-L, 16x4)","rank_in_archive_order":16,"of":38,"metrics":{"mAP":"36.1"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain, ViT-L, 16x4)","rank_in_archive_order":19,"of":38,"metrics":{"mAP":"34.3"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain+finetune, ViT-B, 16x4)","rank_in_archive_order":23,"of":38,"metrics":{"mAP":"31.8"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset":"AVA v2.2","model":"VideoMAE (K400 pretrain, ViT-B, 16x4)","rank_in_archive_order":33,"of":38,"metrics":{"mAP":"26.7"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"VideoMAE (no extra data, ViT-L, 32x2)","rank_in_archive_order":9,"of":123,"metrics":{"GFLOPs":"1436x3","Parameters":"305","Top-1 Accuracy":"75.4","Top-5 Accuracy":"95.2"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"VideoMAE (no extra data, ViT-L, 16frame)","rank_in_archive_order":15,"of":123,"metrics":{"GFLOPs":"597x6","Parameters":"305","Top-1 Accuracy":"74.3","Top-5 Accuracy":"94.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"VideoMAE (no extra data, ViT-B, 16frame)","rank_in_archive_order":36,"of":123,"metrics":{"GFLOPs":"180x6","Parameters":"87","Top-1 Accuracy":"70.8","Top-5 Accuracy":"92.4"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-hmdb51","task":"Self-Supervised Action Recognition","dataset":"HMDB51","model":"VideoMAE","rank_in_archive_order":5,"of":48,"metrics":{"Frozen":"false","Pre-Training Dataset":"Kinetics400","Top-1 Accuracy":"73.3"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-hmdb51","task":"Self-Supervised Action Recognition","dataset":"HMDB51","model":"VideoMAE(no extra data)","rank_in_archive_order":22,"of":48,"metrics":{"Frozen":"false","Pre-Training Dataset":"no extra data","Top-1 Accuracy":"62.6"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101","task":"Self-Supervised Action Recognition","dataset":"UCF101","model":"VideoMAE","rank_in_archive_order":6,"of":53,"metrics":{"3-fold Accuracy":"96.1","Frozen":"false","Pre-Training Dataset":"Kinetics400"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101","task":"Self-Supervised Action Recognition","dataset":"UCF101","model":"VideoMAE(no extra data)","rank_in_archive_order":20,"of":53,"metrics":{"3-fold Accuracy":"91.3","Frozen":"false","Pre-Training Dataset":"no extra data"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2203.12602","atlas_url":"https://app.syntology.ai/?focus=2203.12602","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2203.12602"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/huggingface/transformers","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/MCG-NJU/VideoMAE","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-7/tree/main/VideoMAE","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-9/tree/main/5/videomae","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/innat/VideoMAE","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MS-P3/code7/tree/main/videomae","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-1/tree/main/videomae","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-13/tree/main/token_learner","reach":null}],"summary":{"ran":7,"ran_honours":1,"ran_draft_wrong":2,"unverified":3},"by_repo_kind":{"official":{"samples":12,"ran":9,"repositories":2},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":12,"samples":[{"code_sha256_prefix":"3510cd2667e7b0ec","entry":"PatchEmbed","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"3510cd2667e7b0ec"}},{"code_sha256_prefix":"b4a6e4ea06502d03","entry":"Pooler3d","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b4a6e4ea06502d03"}},{"code_sha256_prefix":"81bb1c42f767f787","entry":"ROIAlign3d","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"81bb1c42f767f787"}},{"code_sha256_prefix":"2648ece3ee513f86","entry":"ROIPool3d","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"2648ece3ee513f86"}},{"code_sha256_prefix":"6589a88b092ff9b4","entry":"ROIPoolingCfg","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"6589a88b092ff9b4"}},{"code_sha256_prefix":"095c2b760f9d3919","entry":"_ROIPool3d","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"095c2b760f9d3919"}},{"code_sha256_prefix":"da651e3979a18f84","entry":"get_sinusoid_encoding_table","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"da651e3979a18f84"}},{"code_sha256_prefix":"2a5af7ff657cc9fd","entry":"get_sinusoid_encoding_table_tf","repo":"innat/VideoMAE","repo_kind":"listed","path":"videomae/utils/sinusoid_encoding_table.py","file_url":"https://github.com/innat/VideoMAE/blob/HEAD/videomae/utils/sinusoid_encoding_table.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2a5af7ff657cc9fd"}},{"code_sha256_prefix":"b8c0a309c689d127","entry":"interpolate_pos_embed_online","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"b8c0a309c689d127"}},{"code_sha256_prefix":"0a01080f8a3a31e8","entry":"make_3d_pooler","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"0a01080f8a3a31e8"}},{"code_sha256_prefix":"de43f6cdccf9a26b","entry":"PretrainVisionTransformer","repo":"MCG-NJU/VideoMAE","repo_kind":"official","path":"modeling_pretrain.py","file_url":"https://github.com/MCG-NJU/VideoMAE/blob/HEAD/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"de43f6cdccf9a26b"}},{"code_sha256_prefix":"0b10363b9191132c","entry":"PretrainVisionTransformerEncoder","repo":"MCG-NJU/VideoMAE","repo_kind":"official","path":"modeling_pretrain.py","file_url":"https://github.com/MCG-NJU/VideoMAE/blob/HEAD/modeling_pretrain.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"0b10363b9191132c"}},{"code_sha256_prefix":"15190c9e8134c543","entry":"VisionTransformer","repo":"MCG-NJU/VideoMAE-Action-Detection","repo_kind":"official","path":"modeling_finetune.py","file_url":"https://github.com/MCG-NJU/VideoMAE-Action-Detection/blob/HEAD/modeling_finetune.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"15190c9e8134c543"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}