{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-contrastive-learning-with-global","title":"Video Contrastive Learning with Global Context","arxiv_id":"2108.02722","date":"2021-08-05","proceeding":null,"authors":["Haofei Kuang","Yi Zhu","Zhi Zhang","Xinyu Li","Joseph Tighe","Sören Schwertfeger","Cyrill Stachniss","Mu Li"],"abstract":"Contrastive learning has revolutionized self-supervised image representation learning field, and recently been adapted to video domain. One of the greatest advantages of contrastive learning is that it allows us to flexibly define powerful loss objectives as long as we can find a reasonable way to formulate positive and negative samples to contrast. However, existing approaches rely heavily on the short-range spatiotemporal salience to form clip-level contrastive signals, thus limit themselves from using global context. In this paper, we propose a new video-level contrastive learning method based on segments to formulate positive pairs. Our formulation is able to capture global context in a video, thus robust to temporal content change. We also incorporate a temporal order regularization term to enforce the inherent sequential structure of videos. Extensive experiments show that our video-level contrastive learning framework (VCLR) is able to outperform previous state-of-the-arts on five video datasets for downstream action classification, action localization and video retrieval. Code is available at https://github.com/amazon-research/video-contrastive-learning.","url_abs":"https://arxiv.org/abs/2108.02722v1","url_pdf":"https://arxiv.org/pdf/2108.02722v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-contrastive-learning-with-global","repo_url":"https://github.com/amazon-research/video-contrastive-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2108.02722","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2108.02722"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/amazon-research/video-contrastive-learning","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_draft_wrong":2,"unverified":6},"by_repo_kind":{"official":{"samples":8,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d9def42110729a85","entry":"conv1x1","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"models/resnet_mlp.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/models/resnet_mlp.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d9def42110729a85"}},{"code_sha256_prefix":"160bb14bd76201b4","entry":"conv3x3","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"models/resnet_mlp.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/models/resnet_mlp.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"160bb14bd76201b4"}},{"code_sha256_prefix":"b135855d925db2bf","entry":"evaluations","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"liblinear/commonutil.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/liblinear/commonutil.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b135855d925db2bf"}},{"code_sha256_prefix":"45c60fc38e15ae86","entry":"evaluations_scipy","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"liblinear/commonutil.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/liblinear/commonutil.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"45c60fc38e15ae86"}},{"code_sha256_prefix":"9377bcc54e512f5b","entry":"resnet18","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"models/resnet_mlp.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/models/resnet_mlp.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9377bcc54e512f5b"}},{"code_sha256_prefix":"b0b94b20ebbcfeed","entry":"retrieval_imgs","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"eval_retrieval_store_imgs.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/eval_retrieval_store_imgs.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b0b94b20ebbcfeed"}},{"code_sha256_prefix":"5349a6905964c533","entry":"shuffle_list","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"dataset/dataset_kinetics.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/dataset/dataset_kinetics.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5349a6905964c533"}},{"code_sha256_prefix":"27f574a9cb80a79f","entry":"svm_read_problem","repo":"amazon-research/video-contrastive-learning","repo_kind":"official","path":"liblinear/commonutil.py","file_url":"https://github.com/amazon-research/video-contrastive-learning/blob/HEAD/liblinear/commonutil.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"27f574a9cb80a79f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}