{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-multimodal-versatile-networks","title":"Self-Supervised MultiModal Versatile Networks","arxiv_id":"2006.16228","date":"2020-06-29","proceeding":"NeurIPS 2020 12","authors":["Jean-Baptiste Alayrac","Adrià Recasens","Rosalia Schneider","Relja Arandjelović","Jason Ramapuram","Jeffrey De Fauw","Lucas Smaira","Sander Dieleman","Andrew Zisserman"],"abstract":"Videos are a rich source of multi-modal supervision. In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visual, audio and language streams. To this end, we introduce the notion of a multimodal versatile network -- a network that can ingest multiple modalities and whose representations enable downstream tasks in multiple modalities. In particular, we explore how best to combine the modalities, such that fine-grained representations of the visual and audio modalities can be maintained, whilst also integrating text into a common embedding. Driven by versatility, we also introduce a novel process of deflation, so that the networks can be effortlessly applied to the visual data in the form of video or a static image. We demonstrate how such networks trained on large collections of unlabelled video data can be applied on video, video-text, image and audio tasks. Equipped with these representations, we obtain state-of-the-art performance on multiple challenging benchmarks including UCF101, HMDB51, Kinetics600, AudioSet and ESC-50 when compared to previous self-supervised work. Our models are publicly available.","url_abs":"https://arxiv.org/abs/2006.16228v2","url_pdf":"https://arxiv.org/pdf/2006.16228v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-supervised-multimodal-versatile-networks","repo_url":"https://github.com/deepmind/deepmind-research/tree/master/mmv","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"jax","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"audio-classification","task_name":"Audio Classification"},{"task_slug":"self-supervised-action-recognition","task_name":"Self-Supervised Action Recognition"},{"task_slug":"self-supervised-audio-classification","task_name":"Self-Supervised Audio Classification"}],"methods":[{"method_slug":"deflation","method_name":"Deflation"}],"datasets_introduced":[],"methods_introduced":[{"slug":"deflation","name":"Deflation","full_name":"Deflation"}],"results":[{"leaderboard":"/sota/audio-classification-on-audioset","task":"Audio Classification","dataset":"AudioSet","model":"MMV","rank_in_archive_order":48,"of":51,"metrics":{"Test mAP":"0.309"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-hmdb51-1","task":"Self-Supervised Action Recognition","dataset":"HMDB51 (finetuned)","model":"MMV","rank_in_archive_order":2,"of":14,"metrics":{"Top-1 Accuracy":"70.1"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on","task":"Self-Supervised Action Recognition","dataset":"Kinetics-600","model":"MMV","rank_in_archive_order":5,"of":5,"metrics":{"Top-1 Accuracy":"55.5"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101","task":"Self-Supervised Action Recognition","dataset":"UCF101","model":"MMV TSM-50x2","rank_in_archive_order":8,"of":53,"metrics":{"3-fold Accuracy":"95.2","Frozen":"false","Pre-Training Dataset":"Audioset + Howto100M"},"uses_additional_data":false},{"leaderboard":"/sota/self-supervised-action-recognition-on-ucf101-1","task":"Self-Supervised Action Recognition","dataset":"UCF101 (finetuned)","model":"MMV","rank_in_archive_order":8,"of":14,"metrics":{"3-fold Accuracy":"91.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2006.16228","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}