{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/screenqa-large-scale-question-answer-pairs","title":"ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots","arxiv_id":"2209.08199","date":"2022-09-16","proceeding":null,"authors":["Yu-Chung Hsiao","Fedir Zubach","Gilles Baechler","Srinivas Sunkara","Victor Carbune","Jason Lin","Maria Wang","Yun Zhu","Jindong Chen"],"abstract":"We introduce ScreenQA, a novel benchmarking dataset designed to advance screen content understanding through question answering. The existing screen datasets are focused either on low-level structural and component understanding, or on a much higher-level composite task such as navigation and task completion for autonomous agents. ScreenQA attempts to bridge this gap. By annotating 86k question-answer pairs over the RICO dataset, we aim to benchmark the screen reading comprehension capacity, thereby laying the foundation for vision-based automation over screenshots. Our annotations encompass full answers, short answer phrases, and corresponding UI contents with bounding boxes, enabling four subtasks to address various application scenarios. We evaluate the dataset's efficacy using both open-weight and proprietary models in zero-shot, fine-tuned, and transfer learning settings. We further demonstrate positive transfer to web applications, highlighting its potential beyond mobile applications.","url_abs":"https://arxiv.org/abs/2209.08199v4","url_pdf":"https://arxiv.org/pdf/2209.08199v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"screenqa-large-scale-question-answer-pairs","repo_url":"https://github.com/google-research-datasets/screen_qa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2209.08199","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2209.08199"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google-research-datasets/screen_qa","reach":null}],"summary":{"ran_honours":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"a76475def1944b88","entry":"sqa_s_metrics","repo":"google-research-datasets/screen_qa","repo_kind":"official","path":"code/metrics.py","file_url":"https://github.com/google-research-datasets/screen_qa/blob/HEAD/code/metrics.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"a76475def1944b88"}},{"code_sha256_prefix":"64d2dd90e3e9ceee","entry":"sqa_uic_bb_metrics","repo":"google-research-datasets/screen_qa","repo_kind":"official","path":"code/metrics.py","file_url":"https://github.com/google-research-datasets/screen_qa/blob/HEAD/code/metrics.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"64d2dd90e3e9ceee"}},{"code_sha256_prefix":"f0ff8458c564145c","entry":"sqa_uic_metrics","repo":"google-research-datasets/screen_qa","repo_kind":"official","path":"code/metrics.py","file_url":"https://github.com/google-research-datasets/screen_qa/blob/HEAD/code/metrics.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"f0ff8458c564145c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}