{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/merlot-reserve-neural-script-knowledge","title":"MERLOT Reserve: Neural Script Knowledge through Vision and Language and Sound","arxiv_id":"2201.02639","date":"2022-01-07","proceeding":"CVPR 2022 1","authors":["Rowan Zellers","Jiasen Lu","Ximing Lu","Youngjae Yu","Yanpeng Zhao","Mohammadreza Salehi","Aditya Kusupati","Jack Hessel","Ali Farhadi","Yejin Choi"],"abstract":"As humans, we navigate a multimodal world, building a holistic understanding from all our senses. We introduce MERLOT Reserve, a model that represents videos jointly over time -- through a new training objective that learns from audio, subtitles, and video frames. Given a video, we replace snippets of text and audio with a MASK token; the model learns by choosing the correct masked-out snippet. Our objective learns faster than alternatives, and performs well at scale: we pretrain on 20 million YouTube videos. Empirical results show that MERLOT Reserve learns strong multimodal representations. When finetuned, it sets state-of-the-art on Visual Commonsense Reasoning (VCR), TVQA, and Kinetics-600; outperforming prior work by 5%, 7%, and 1.5% respectively. Ablations show that these tasks benefit from audio pretraining -- even VCR, a QA task centered around images (without sound). Moreover, our objective enables out-of-the-box prediction, revealing strong multimodal commonsense understanding. In a fully zero-shot setting, our model obtains competitive results on four video tasks, even outperforming supervised approaches on the recently proposed Situated Reasoning (STAR) benchmark. We analyze why audio enables better vision-language representations, suggesting significant opportunities for future research. We conclude by discussing ethical and societal implications of multimodal pretraining.","url_abs":"https://arxiv.org/abs/2201.02639v4","url_pdf":"https://arxiv.org/pdf/2201.02639v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"visual-commonsense-reasoning","task_name":"Visual Commonsense Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-600","task":"Action Classification","dataset":"Kinetics-600","model":"🍷MerlotReserve-Large (+Audio)","rank_in_archive_order":6,"of":65,"metrics":{"Top-1 Accuracy":"91.1","Top-5 Accuracy":"97.1"},"uses_additional_data":true},{"leaderboard":"/sota/action-classification-on-kinetics-600","task":"Action Classification","dataset":"Kinetics-600","model":"🍷MerlotReserve-Base (+Audio)","rank_in_archive_order":14,"of":65,"metrics":{"Top-1 Accuracy":"89.7","Top-5 Accuracy":"96.6"},"uses_additional_data":true},{"leaderboard":"/sota/action-classification-on-kinetics-600","task":"Action Classification","dataset":"Kinetics-600","model":"🍷MerlotReserve-Large (no Audio)","rank_in_archive_order":15,"of":65,"metrics":{"Top-1 Accuracy":"89.4","Top-5 Accuracy":"96.3"},"uses_additional_data":true},{"leaderboard":"/sota/action-classification-on-kinetics-600","task":"Action Classification","dataset":"Kinetics-600","model":"🍷MerlotReserve-Base (no Audio)","rank_in_archive_order":22,"of":65,"metrics":{"Top-1 Accuracy":"88.1","Top-5 Accuracy":"95.8"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2201.02639","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}