{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modist-motion-distillation-for-self","title":"MaCLR: Motion-aware Contrastive Learning of Representations for Videos","arxiv_id":"2106.09703","date":"2021-06-17","proceeding":null,"authors":["Fanyi Xiao","Joseph Tighe","Davide Modolo"],"abstract":"We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learning methods that mostly focus on learning motion cues implicitly from RGB inputs, MaCLR enriches standard contrastive learning objectives for RGB video clips with a cross-modal learning objective between a Motion pathway and a Visual pathway. We show that the representation learned with our MaCLR method focuses more on foreground motion regions and thus generalizes better to downstream tasks. To demonstrate this, we evaluate MaCLR on five datasets for both action recognition and action detection, and demonstrate state-of-the-art self-supervised performance on all datasets. Furthermore, we show that MaCLR representation can be as effective as representations learned with full supervision on UCF101 and HMDB51 action recognition, and even outperform the supervised representation for action recognition on VidSitu and SSv2, and action detection on AVA.","url_abs":"https://arxiv.org/abs/2106.09703v2","url_pdf":"https://arxiv.org/pdf/2106.09703v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modist-motion-distillation-for-self","repo_url":"https://github.com/amazon-science/self-supervised-maclr","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2106.09703","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2106.09703"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/amazon-science/self-supervised-maclr","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":1,"unverified":5},"by_repo_kind":{"official":{"samples":6,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c90b5f606f54d4f7","entry":"get_loss_func","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/losses.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/losses.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c90b5f606f54d4f7"}},{"code_sha256_prefix":"a4723eab8df61dbe","entry":"attention_pool","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/attention.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/attention.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a4723eab8df61dbe"}},{"code_sha256_prefix":"06282ee2b8b703dc","entry":"construct_optimizer","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/optimizer.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/optimizer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"06282ee2b8b703dc"}},{"code_sha256_prefix":"bcc1cdae3bb3212c","entry":"drop_path","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/common.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/common.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"bcc1cdae3bb3212c"}},{"code_sha256_prefix":"1c1fa7933939a08f","entry":"get_trans_func","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/resnet_helper.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/resnet_helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1c1fa7933939a08f"}},{"code_sha256_prefix":"690eba49225ad86f","entry":"wrap_distributed_model","repo":"amazon-science/self-supervised-maclr","repo_kind":"official","path":"slowfast/models/build.py","file_url":"https://github.com/amazon-science/self-supervised-maclr/blob/HEAD/slowfast/models/build.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"690eba49225ad86f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}