{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tallyqa-answering-complex-counting-questions","title":"TallyQA: Answering Complex Counting Questions","arxiv_id":"1810.12440","date":"2018-10-29","proceeding":null,"authors":["Manoj Acharya","Kushal Kafle","Christopher Kanan"],"abstract":"Most counting questions in visual question answering (VQA) datasets are\nsimple and require no more than object detection. Here, we study algorithms for\ncomplex counting questions that involve relationships between objects,\nattribute identification, reasoning, and more. To do this, we created TallyQA,\nthe world's largest dataset for open-ended counting. We propose a new algorithm\nfor counting that uses relation networks with region proposals. Our method lets\nrelation networks be efficiently used with high-resolution imagery. It yields\nstate-of-the-art results compared to baseline and recent systems on both\nTallyQA and the HowMany-QA benchmark.","url_abs":"http://arxiv.org/abs/1810.12440v2","url_pdf":"http://arxiv.org/pdf/1810.12440v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tallyqa-answering-complex-counting-questions","repo_url":"https://github.com/manoja328/tallyqacode","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"object-counting","task_name":"Object Counting"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[{"slug":"tallyqa","name":"TallyQA","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-counting-on-howmany-qa","task":"Object Counting","dataset":"HowMany-QA","model":"RCN","rank_in_archive_order":3,"of":3,"metrics":{"Accuracy":"60.3","RMSE":"2.35"},"uses_additional_data":false},{"leaderboard":"/sota/object-counting-on-tallyqa-complex","task":"Object Counting","dataset":"TallyQA-Complex","model":"RCN","rank_in_archive_order":5,"of":6,"metrics":{"Accuracy":"56.2","RMSE":"1.43"},"uses_additional_data":false},{"leaderboard":"/sota/object-counting-on-tallyqa-simple","task":"Object Counting","dataset":"TallyQA-Simple","model":"RCN","rank_in_archive_order":5,"of":6,"metrics":{"Accuracy":"71.8","RMSE":"1.13"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.12440","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}