Datasets › GEM (A General Evaluation Benchmark on Multi-modal Tasks)
GEM (A General Evaluation Benchmark on Multi-modal Tasks)
GEM (A General Evaluation Benchmark on Multi-modal Tasks) is a significant benchmark dataset designed to evaluate the performance of cross-modal pre-trained models, including both understanding and generation tasks. Unlike existing datasets such as GLUE, SuperGLUE, XGLUE, and XTREME, which primarily focus on natural language tasks, GEM stands out as a large-scale vision-language benchmark.
Here are the key features of GEM:
- Multimodal Focus: GEM covers both vision and language domains. It consists of two main components:
- GEM-I: This part focuses on image-language tasks.
-
GEM-V: This part focuses on video-language tasks.
-
Large-Scale Dataset: GEM is one of the largest vision-language datasets available. It encompasses both image-language and video-language tasks simultaneously.
-
Multilingual Labeling: The dataset is labeled in multiple languages, making it versatile for multilingual multimodal research.
-
Baseline Models: The creators of GEM provide two baseline models to facilitate research and development in this area.
The goal of GEM is to advance the field of multimodal research by providing a comprehensive evaluation benchmark that spans vision and language modalities. Researchers can use this dataset to assess the capabilities of their models across different tasks and languages¹².
(1) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://arxiv.org/abs/2106.09889. (2) GEM Submission Instructions - GitHub Pages. https://microsoft.github.io/GEM/. (3) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://www.microsoft.com/en-us/research/publication/gem-a-general-evaluation-benchmark-for-multimodal-tasks/. (4) undefined. https://doi.org/10.48550/arXiv.2106.09889.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- GEM (A General Evaluation Benchmark on Multi-modal Tasks)
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections