{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/clear-a-dataset-for-compositional-language","title":"CLEAR: A Dataset for Compositional Language and Elementary Acoustic Reasoning","arxiv_id":"1811.10561","date":"2018-11-26","proceeding":null,"authors":["Jerome Abdelnour","Giampiero Salvi","Jean Rouat"],"abstract":"We introduce the task of acoustic question answering (AQA) in the area of\nacoustic reasoning. In this task an agent learns to answer questions on the\nbasis of acoustic context. In order to promote research in this area, we\npropose a data generation paradigm adapted from CLEVR (Johnson et al. 2017). We\ngenerate acoustic scenes by leveraging a bank elementary sounds. We also\nprovide a number of functional programs that can be used to compose questions\nand answers that exploit the relationships between the attributes of the\nelementary sounds in each scene. We provide AQA datasets of various sizes as\nwell as the data generation code. As a preliminary experiment to validate our\ndata, we report the accuracy of current state of the art visual question\nanswering models when they are applied to the AQA task without modifications.\nAlthough there is a plethora of question answering tasks based on text, image\nor video data, to our knowledge, we are the first to propose answering\nquestions directly on audio streams. We hope this contribution will facilitate\nthe development of research in the area.","url_abs":"http://arxiv.org/abs/1811.10561v1","url_pdf":"http://arxiv.org/pdf/1811.10561v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"clear-a-dataset-for-compositional-language","repo_url":"https://github.com/IGLU-CHISTERA/CLEAR-dataset-generation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"acoustic-question-answering","task_name":"Acoustic Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.10561","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}