{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-pretraining-for-dense-video","title":"Multimodal Pretraining for Dense Video Captioning","arxiv_id":"2011.11760","date":"2020-11-10","proceeding":"Asian Chapter of the Association for Computational Linguistics 2020","authors":["Gabriel Huang","Bo Pang","Zhenhai Zhu","Clara Rivera","Radu Soricut"],"abstract":"Learning specific hands-on skills such as cooking, car maintenance, and home repairs increasingly happens via instructional videos. The user experience with such videos is known to be improved by meta-information such as time-stamped annotations for the main steps involved. Generating such annotations automatically is challenging, and we describe here two relevant contributions. First, we construct and release a new dense video captioning dataset, Video Timeline Tags (ViTT), featuring a variety of instructional videos together with time-stamped annotations. Second, we explore several multimodal sequence-to-sequence pretraining strategies that leverage large unsupervised datasets of videos and caption-like texts. We pretrain and subsequently finetune dense video captioning models using both YouCook2 and ViTT. We show that such models generalize well and are robust over a wide variety of instructional videos.","url_abs":"https://arxiv.org/abs/2011.11760v1","url_pdf":"https://arxiv.org/pdf/2011.11760v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-pretraining-for-dense-video","repo_url":"https://github.com/google-research-datasets/Video-Timeline-Tags-ViTT","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"dense-video-captioning","task_name":"Dense Video Captioning"},{"task_slug":"video-captioning","task_name":"Video Captioning"}],"methods":[],"datasets_introduced":[{"slug":"vitt","name":"ViTT","full_name":"Video Timeline Tags"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/dense-video-captioning-on-youcook2","task":"Dense Video Captioning","dataset":"YouCook2","model":"E2vidD6-MASSalign-BiD","rank_in_archive_order":7,"of":7,"metrics":{"ROUGE-L":"39.03"},"uses_additional_data":true},{"leaderboard":"/sota/video-captioning-on-youcook2","task":"Video Captioning","dataset":"YouCook2","model":"E2vidD6-MASSvid-BiD","rank_in_archive_order":6,"of":14,"metrics":{"BLEU-4":"12.04","CIDEr":"1.22","METEOR":"18.32","ROUGE-L":"39.03"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2011.11760","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}