{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-goal-step-inference-using-wikihow","title":"Visual Goal-Step Inference using wikiHow","arxiv_id":"2104.05845","date":"2021-04-12","proceeding":"EMNLP 2021 11","authors":["Yue Yang","Artemis Panagopoulou","Qing Lyu","Li Zhang","Mark Yatskar","Chris Callison-Burch"],"abstract":"Understanding what sequence of steps are needed to complete a goal can help artificial intelligence systems reason about human activities. Past work in NLP has examined the task of goal-step inference for text. We introduce the visual analogue. We propose the Visual Goal-Step Inference (VGSI) task, where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal. With a new dataset harvested from wikiHow consisting of 772,277 images representing human actions, we show that our task is challenging for state-of-the-art multimodal models. Moreover, the multimodal representation learned from our data can be effectively transferred to other datasets like HowTo100m, increasing the VGSI accuracy by 15 - 20%. Our task will facilitate multimodal reasoning about procedural events.","url_abs":"https://arxiv.org/abs/2104.05845v2","url_pdf":"https://arxiv.org/pdf/2104.05845v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-goal-step-inference-using-wikihow","repo_url":"https://github.com/yueyang1996/wikihow-vgsi","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"multimodal-reasoning","task_name":"Multimodal Reasoning"},{"task_slug":"vgsi","task_name":"VGSI"}],"methods":[{"method_slug":"siamese-network","method_name":"Siamese Network"}],"datasets_introduced":[{"slug":"wikihow-image","name":"wikiHow-image","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/vgsi-on-wikihow-image","task":"VGSI","dataset":"wikiHow-image","model":"Triplet Network","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"0.7494"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2104.05845","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}