{"url":"/dataset/gem-a-general-evaluation-benchmark-on-multi","name":"GEM (A General Evaluation Benchmark on Multi-modal Tasks)","full_name":null,"description_markdown":"**GEM** (A General Evaluation Benchmark on Multi-modal Tasks) is a significant benchmark dataset designed to evaluate the performance of cross-modal pre-trained models, including both understanding and generation tasks. Unlike existing datasets such as GLUE, SuperGLUE, XGLUE, and XTREME, which primarily focus on natural language tasks, **GEM** stands out as a large-scale vision-language benchmark.\r\n\r\nHere are the key features of **GEM**:\r\n\r\n1. **Multimodal Focus**: **GEM** covers both vision and language domains. It consists of two main components:\r\n   - **GEM-I**: This part focuses on image-language tasks.\r\n   - **GEM-V**: This part focuses on video-language tasks.\r\n\r\n2. **Large-Scale Dataset**: **GEM** is one of the largest vision-language datasets available. It encompasses both image-language and video-language tasks simultaneously.\r\n\r\n3. **Multilingual Labeling**: The dataset is labeled in multiple languages, making it versatile for multilingual multimodal research.\r\n\r\n4. **Baseline Models**: The creators of **GEM** provide two baseline models to facilitate research and development in this area.\r\n\r\nThe goal of **GEM** is to advance the field of multimodal research by providing a comprehensive evaluation benchmark that spans vision and language modalities. Researchers can use this dataset to assess the capabilities of their models across different tasks and languages¹².\r\n\r\n(1) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://arxiv.org/abs/2106.09889.\r\n(2) GEM Submission Instructions - GitHub Pages. https://microsoft.github.io/GEM/.\r\n(3) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://www.microsoft.com/en-us/research/publication/gem-a-general-evaluation-benchmark-for-multimodal-tasks/.\r\n(4) undefined. https://doi.org/10.48550/arXiv.2106.09889.","description_withheld":null,"homepage":"https://microsoft.github.io/GEM/","introduced_date":"2021-06-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/gem-a-general-evaluation-benchmark-for","title":"GEM: A General Evaluation Benchmark for Multimodal Tasks","first_author":"Lin Su","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["GEM (A General Evaluation Benchmark on Multi-modal Tasks)"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}