{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-hierarchical-script-knowledge","title":"Benchmarking Hierarchical Script Knowledge","arxiv_id":null,"date":"2019-06-01","proceeding":"NAACL 2019 6","authors":["Yonatan Bisk","Jan Buys","Karl Pichotta","Yejin Choi"],"abstract":"Understanding procedural language requires reasoning about both hierarchical and temporal relations between events. For example, {``}boiling pasta{''} is a sub-event of {``}making a pasta dish{''}, typically happens before {``}draining pasta,{''} and requires the use of omitted tools (e.g. a strainer, sink...). While people are able to choose when and how to use abstract versus concrete instructions, the NLP community lacks corpora and tasks for evaluating if our models can do the same. In this paper, we introduce KidsCook, a parallel script corpus, as well as a cloze task which matches video captions with missing procedural details. Experimental results show that state-of-the-art models struggle at this task, which requires inducing functional commonsense knowledge not explicitly stated in text.","url_abs":"https://aclanthology.org/N19-1412","url_pdf":"https://aclanthology.org/N19-1412.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-hierarchical-script-knowledge","repo_url":"https://github.com/janmbuys/ScriptTransduction","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}