{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-clustering-networks-for-self","title":"Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos","arxiv_id":"2104.12671","date":"2021-04-26","proceeding":"ICCV 2021 10","authors":["Brian Chen","Andrew Rouditchenko","Kevin Duarte","Hilde Kuehne","Samuel Thomas","Angie Boggust","Rameswar Panda","Brian Kingsbury","Rogerio Feris","David Harwath","James Glass","Michael Picheny","Shih-Fu Chang"],"abstract":"Multimodal self-supervised learning is getting more and more attention as it allows not only to train large networks without human supervision but also to search and retrieve data across various modalities. In this context, this paper proposes a self-supervised training framework that learns a common multimodal embedding space that, in addition to sharing representations across different modalities, enforces a grouping of semantically similar instances. To this end, we extend the concept of instance-level contrastive learning with a multimodal clustering step in the training pipeline to capture semantic similarities across modalities. The resulting embedding space enables retrieval of samples across all modalities, even from unseen datasets and different domains. To evaluate our approach, we train our model on the HowTo100M dataset and evaluate its zero-shot retrieval capabilities in two challenging domains, namely text-to-video retrieval, and temporal action localization, showing state-of-the-art results on four different datasets.","url_abs":"https://arxiv.org/abs/2104.12671v3","url_pdf":"https://arxiv.org/pdf/2104.12671v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-clustering-networks-for-self","repo_url":"https://github.com/brian7685/Multimodal-Clustering-Network","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"long-video-retrieval-background-removed","task_name":"Long Video Retrieval (Background Removed)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"text-to-video-retrieval","task_name":"Text to Video Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/long-video-retrieval-background-removed-on","task":"Long Video Retrieval (Background Removed)","dataset":"YouCook2","model":"MCN","rank_in_archive_order":4,"of":6,"metrics":{"Cap. Avg. R@1":"53.4","Cap. Avg. R@10":"81.4","Cap. Avg. R@5":"75.0"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.12671","atlas_url":"https://app.syntology.ai/?focus=2104.12671","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2104.12671"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/brian7685/Multimodal-Clustering-Network","reach":null}],"summary":{"ran_fixture":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"d7b23e73c9749fe8","entry":"update_queue","repo":"brian7685/Multimodal-Clustering-Network","repo_kind":"official","path":"train_tri_kmeans.py","file_url":"https://github.com/brian7685/Multimodal-Clustering-Network/blob/HEAD/train_tri_kmeans.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d7b23e73c9749fe8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}