{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/shapeworld-a-new-test-methodology-for","title":"ShapeWorld - A new test methodology for multimodal language understanding","arxiv_id":"1704.04517","date":"2017-04-14","proceeding":null,"authors":["Alexander Kuhnle","Ann Copestake"],"abstract":"We introduce a novel framework for evaluating multimodal deep learning models\nwith respect to their language understanding and generalization abilities. In\nthis approach, artificial data is automatically generated according to the\nexperimenter's specifications. The content of the data, both during training\nand evaluation, can be controlled in detail, which enables tasks to be created\nthat require true generalization abilities, in particular the combination of\npreviously introduced concepts in novel ways. We demonstrate the potential of\nour methodology by evaluating various visual question answering models on four\ndifferent tasks, and show how our framework gives us detailed insights into\ntheir capabilities and limitations. By open-sourcing our framework, we hope to\nstimulate progress in the field of multimodal language understanding.","url_abs":"http://arxiv.org/abs/1704.04517v1","url_pdf":"http://arxiv.org/pdf/1704.04517v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"shapeworld-a-new-test-methodology-for","repo_url":"https://github.com/AlexKuhnle/ShapeWorld","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"shapeworld-a-new-test-methodology-for","repo_url":"https://github.com/dhruvyad/MultimodalGame","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"shapeworld-a-new-test-methodology-for","repo_url":"https://github.com/lgraesser/MultimodalGame","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"multimodal-deep-learning","task_name":"Multimodal Deep Learning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"shapeworld","name":"ShapeWorld","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.04517","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}