{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gem-a-general-evaluation-benchmark-for","title":"GEM: A General Evaluation Benchmark for Multimodal Tasks","arxiv_id":"2106.09889","date":"2021-06-18","proceeding":"Findings (ACL) 2021 8","authors":["Lin Su","Nan Duan","Edward Cui","Lei Ji","Chenfei Wu","Huaishao Luo","Yongfei Liu","Ming Zhong","Taroon Bharti","Arun Sacheti"],"abstract":"In this paper, we present GEM as a General Evaluation benchmark for Multimodal tasks. Different from existing datasets such as GLUE, SuperGLUE, XGLUE and XTREME that mainly focus on natural language tasks, GEM is a large-scale vision-language benchmark, which consists of GEM-I for image-language tasks and GEM-V for video-language tasks. Comparing with existing multimodal datasets such as MSCOCO and Flicker30K for image-language tasks, YouCook2 and MSR-VTT for video-language tasks, GEM is not only the largest vision-language dataset covering image-language tasks and video-language tasks at the same time, but also labeled in multiple languages. We also provide two baseline models for this benchmark. We will release the dataset, code and baseline models, aiming to advance the development of multilingual multimodal research.","url_abs":"https://arxiv.org/abs/2106.09889v1","url_pdf":"https://arxiv.org/pdf/2106.09889v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gem-a-general-evaluation-benchmark-for","repo_url":"https://github.com/microsoft/GEM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"gem-a-general-evaluation-benchmark-on-multi","name":"GEM (A General Evaluation Benchmark on Multi-modal Tasks)","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2106.09889","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2106.09889"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/Maluuba/nlg-eval","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/microsoft/UniVL","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/microsoft/GEM","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"593322fdfd62e04e","entry":"compute_metrics","repo":"microsoft/GEM","repo_kind":"official","path":"evaluation/metric-retrieval.py","file_url":"https://github.com/microsoft/GEM/blob/HEAD/evaluation/metric-retrieval.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"593322fdfd62e04e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}