{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/broaden-the-vision-geo-diverse-visual","title":"Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning","arxiv_id":"2109.06860","date":"2021-09-14","proceeding":"EMNLP 2021 11","authors":["Da Yin","Liunian Harold Li","Ziniu Hu","Nanyun Peng","Kai-Wei Chang"],"abstract":"Commonsense is defined as the knowledge that is shared by everyone. However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally. For example, the scenarios of wedding ceremonies vary across regions due to different customs influenced by historical and religious factors. Such regional characteristics, however, are generally omitted in prior work. In this paper, we construct a Geo-Diverse Visual Commonsense Reasoning dataset (GD-VCR) to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense. In particular, we study two state-of-the-art Vision-and-Language models, VisualBERT and ViLBERT trained on VCR, a standard multimodal commonsense benchmark with images primarily from Western regions. We then evaluate how well the trained models can generalize to answering the questions in GD-VCR. We find that the performance of both models for non-Western regions including East Asia, South Asia, and Africa is significantly lower than that for Western region. We analyze the reasons behind the performance disparity and find that the performance gap is larger on QA pairs that: 1) are concerned with culture-related scenarios, e.g., weddings, religious activities, and festivals; 2) require high-level geo-diverse commonsense reasoning rather than low-order perception and recognition. Dataset and code are released at https://github.com/WadeYin9712/GD-VCR.","url_abs":"https://arxiv.org/abs/2109.06860v1","url_pdf":"https://arxiv.org/pdf/2109.06860v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"broaden-the-vision-geo-diverse-visual","repo_url":"https://github.com/wadeyin9712/gd-vcr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"culture","task_name":"Cultural Vocal Bursts Intensity Prediction"},{"task_slug":"visual-commonsense-reasoning","task_name":"Visual Commonsense Reasoning"}],"methods":[{"method_slug":"vilbert","method_name":"ViLBERT"},{"method_slug":"visualbert","method_name":"VisualBERT"}],"datasets_introduced":[{"slug":"gd-vcr","name":"GD-VCR","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-commonsense-reasoning-on-gd-vcr","task":"Visual Commonsense Reasoning","dataset":"GD-VCR","model":"Human","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy":"88.84"},"uses_additional_data":false},{"leaderboard":"/sota/visual-commonsense-reasoning-on-gd-vcr","task":"Visual Commonsense Reasoning","dataset":"GD-VCR","model":"ViLBERT","rank_in_archive_order":2,"of":4,"metrics":{"Accuracy":"59.99","Gap (West)":"-7.28"},"uses_additional_data":false},{"leaderboard":"/sota/visual-commonsense-reasoning-on-gd-vcr","task":"Visual Commonsense Reasoning","dataset":"GD-VCR","model":"VisualBERT","rank_in_archive_order":3,"of":4,"metrics":{"Accuracy":"53.95","Gap (West)":"-10.42"},"uses_additional_data":false},{"leaderboard":"/sota/visual-commonsense-reasoning-on-gd-vcr","task":"Visual Commonsense Reasoning","dataset":"GD-VCR","model":"Text-only BERT","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy":"35.33"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2109.06860","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2109.06860"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wadeyin9712/gd-vcr","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":9},"by_repo_kind":{"official":{"samples":9,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9905a41b46f307b0","entry":"converId","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"vilbert_beta/script/convert_lmdb_VCR.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/vilbert_beta/script/convert_lmdb_VCR.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9905a41b46f307b0"}},{"code_sha256_prefix":"46ddc1aeff61e561","entry":"fix_item","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"visualbert/dataloaders/vcr_data_utils.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/visualbert/dataloaders/vcr_data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"46ddc1aeff61e561"}},{"code_sha256_prefix":"c96b2567710f4ba3","entry":"generate_answer_choices","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"build_dataset/similarity.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/build_dataset/similarity.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c96b2567710f4ba3"}},{"code_sha256_prefix":"78e86f22ea46c2c2","entry":"getEuclidean","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"build_dataset/question_cluster.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/build_dataset/question_cluster.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"78e86f22ea46c2c2"}},{"code_sha256_prefix":"34b7e4798f39686c","entry":"k_means","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"build_dataset/question_cluster.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/build_dataset/question_cluster.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"34b7e4798f39686c"}},{"code_sha256_prefix":"56ced5a70a36ad9e","entry":"limit_range","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"build_dataset/similarity.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/build_dataset/similarity.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"56ced5a70a36ad9e"}},{"code_sha256_prefix":"ad276e7a73c8649e","entry":"process_ctx_ans_for_bert","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"visualbert/dataloaders/vcr_data_utils.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/visualbert/dataloaders/vcr_data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ad276e7a73c8649e"}},{"code_sha256_prefix":"be08e5e375ccf6be","entry":"retokenize_with_alignment","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"visualbert/dataloaders/vcr_data_utils.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/visualbert/dataloaders/vcr_data_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"be08e5e375ccf6be"}},{"code_sha256_prefix":"ea2f5a5e3b004c3a","entry":"text_preprocessing","repo":"wadeyin9712/gd-vcr","repo_kind":"official","path":"build_dataset/relevance_model.py","file_url":"https://github.com/wadeyin9712/gd-vcr/blob/HEAD/build_dataset/relevance_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ea2f5a5e3b004c3a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}