{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/winogavil-gamified-association-benchmark-to","title":"WinoGAViL: Gamified Association Benchmark to Challenge Vision-and-Language Models","arxiv_id":"2207.12576","date":"2022-07-25","proceeding":null,"authors":["Yonatan Bitton","Nitzan Bitton Guetta","Ron Yosef","Yuval Elovici","Mohit Bansal","Gabriel Stanovsky","Roy Schwartz"],"abstract":"While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we introduce WinoGAViL: an online game of vision-and-language associations (e.g., between werewolves and a full moon), used as a dynamic evaluation benchmark. Inspired by the popular card game Codenames, a spymaster gives a textual cue related to several visual candidates, and another player tries to identify them. Human players are rewarded for creating associations that are challenging for a rival AI model but still solvable by other human players. We use the game to collect 3.5K instances, finding that they are intuitive for humans (>90% Jaccard index) but challenging for state-of-the-art AI models, where the best model (ViLT) achieves a score of 52%, succeeding mostly where the cue is visually salient. Our analysis as well as the feedback we collect from players indicate that the collected associations require diverse reasoning skills, including general knowledge, common sense, abstraction, and more. We release the dataset, the code and the interactive game, allowing future data collection that can be used to develop models with better association abilities.","url_abs":"https://arxiv.org/abs/2207.12576v2","url_pdf":"https://arxiv.org/pdf/2207.12576v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"winogavil-gamified-association-benchmark-to","repo_url":"https://github.com/winogavil/winogavil-experiments","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"general-knowledge","task_name":"General Knowledge"},{"task_slug":"multimodal-association","task_name":"Multimodal Association"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[{"slug":"winogavil","name":"WinoGAViL","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/common-sense-reasoning-on-winogavil","task":"Common Sense Reasoning","dataset":"WinoGAViL","model":"ViLT","rank_in_archive_order":1,"of":1,"metrics":{"Jaccard Index":"52"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"Humans","rank_in_archive_order":1,"of":8,"metrics":{"Jaccard Index":"90"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"ViLT (Zero-Shot)","rank_in_archive_order":2,"of":8,"metrics":{"Jaccard Index":"52"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"X-VLM (Zero-Shot)","rank_in_archive_order":3,"of":8,"metrics":{"Jaccard Index":"46"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"CLIP-ViT-B/32 (Zero-Shot)","rank_in_archive_order":4,"of":8,"metrics":{"Jaccard Index":"41"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"CLIP-ViT-L/14 (Zero-Shot)","rank_in_archive_order":5,"of":8,"metrics":{"Jaccard Index":"40"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"CLIP-RN50x64/14 (Zero-Shot)","rank_in_archive_order":6,"of":8,"metrics":{"Jaccard Index":"38"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"CLIP-RN50 (Zero-Shot)","rank_in_archive_order":7,"of":8,"metrics":{"Jaccard Index":"35"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winogavil","task":"Visual Reasoning","dataset":"WinoGAViL","model":"CLIP-ViL (Zero-Shot)","rank_in_archive_order":8,"of":8,"metrics":{"Jaccard Index":"15"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2207.12576","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2207.12576"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/WinoGAViL/WinoGAViL-experiments","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/winogavil/winogavil-experiments","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"fb34eee819c2938e","entry":"get_scores","repo":"WinoGAViL/WinoGAViL-experiments","repo_kind":"official","path":"run_zero_shot.py","file_url":"https://github.com/WinoGAViL/WinoGAViL-experiments/blob/HEAD/run_zero_shot.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fb34eee819c2938e"}},{"code_sha256_prefix":"fcb30c44e7bcdb29","entry":"test_epoch","repo":"WinoGAViL/WinoGAViL-experiments","repo_kind":"official","path":"run_trainable.py","file_url":"https://github.com/WinoGAViL/WinoGAViL-experiments/blob/HEAD/run_trainable.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fcb30c44e7bcdb29"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}