{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-from-narrated","title":"Unsupervised Learning from Narrated Instruction Videos","arxiv_id":"1506.09215","date":"2015-06-30","proceeding":"CVPR 2016 6","authors":["Jean-Baptiste Alayrac","Piotr Bojanowski","Nishant Agrawal","Josef Sivic","Ivan Laptev","Simon Lacoste-Julien"],"abstract":"We address the problem of automatically learning the main steps to complete a\ncertain task, such as changing a car tire, from a set of narrated instruction\nvideos. The contributions of this paper are three-fold. First, we develop a new\nunsupervised learning approach that takes advantage of the complementary nature\nof the input video and the associated narration. The method solves two\nclustering problems, one in text and one in video, applied one after each other\nand linked by joint constraints to obtain a single coherent sequence of steps\nin both modalities. Second, we collect and annotate a new challenging dataset\nof real-world instruction videos from the Internet. The dataset contains about\n800,000 frames for five different tasks that include complex interactions\nbetween people and objects, and are captured in a variety of indoor and outdoor\nsettings. Third, we experimentally demonstrate that the proposed method can\nautomatically discover, in an unsupervised manner, the main steps to achieve\nthe task and locate the steps in the input videos.","url_abs":"http://arxiv.org/abs/1506.09215v4","url_pdf":"http://arxiv.org/pdf/1506.09215v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/temporal-action-localization-on-crosstask","task":"Temporal Action Localization","dataset":"CrossTask","model":"Alayrac","rank_in_archive_order":7,"of":7,"metrics":{"Recall":"13.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.09215","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}