{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/whats-cookin-interpreting-cooking-videos-1","title":"What's Cookin'? Interpreting Cooking Videos using Text, Speech and Vision","arxiv_id":"1503.01558","date":"2015-03-05","proceeding":null,"authors":["Jonathan Malmaud","Jonathan Huang","Vivek Rathod","Nick Johnston","Andrew Rabinovich","Kevin Murphy"],"abstract":"We present a novel method for aligning a sequence of instructions to a video\nof someone carrying out a task. In particular, we focus on the cooking domain,\nwhere the instructions correspond to the recipe. Our technique relies on an HMM\nto align the recipe steps to the (automatically generated) speech transcript.\nWe then refine this alignment using a state-of-the-art visual food detector,\nbased on a deep convolutional neural network. We show that our technique\noutperforms simpler techniques based on keyword spotting. It also enables\ninteresting applications, such as automatically illustrating recipes with\nkeyframes, and searching within a video for events of interest.","url_abs":"http://arxiv.org/abs/1503.01558v3","url_pdf":"http://arxiv.org/pdf/1503.01558v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"whats-cookin-interpreting-cooking-videos-1","repo_url":"https://github.com/malmaud/whats_cookin","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1503.01558","atlas_url":"https://app.syntology.ai/?focus=1503.01558","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}