Datasets › GEM (A General Evaluation Benchmark on Multi-modal Tasks)

GEM (A General Evaluation Benchmark on Multi-modal Tasks)

Introduced by Lin Su et al. in GEM: A General Evaluation Benchmark for Multimodal Tasks18 Jun 2021 archive 2025-07-28

GEM (A General Evaluation Benchmark on Multi-modal Tasks) is a significant benchmark dataset designed to evaluate the performance of cross-modal pre-trained models, including both understanding and generation tasks. Unlike existing datasets such as GLUE, SuperGLUE, XGLUE, and XTREME, which primarily focus on natural language tasks, GEM stands out as a large-scale vision-language benchmark.

Here are the key features of GEM:

  1. Multimodal Focus: GEM covers both vision and language domains. It consists of two main components:
  2. GEM-I: This part focuses on image-language tasks.
  3. GEM-V: This part focuses on video-language tasks.

  4. Large-Scale Dataset: GEM is one of the largest vision-language datasets available. It encompasses both image-language and video-language tasks simultaneously.

  5. Multilingual Labeling: The dataset is labeled in multiple languages, making it versatile for multilingual multimodal research.

  6. Baseline Models: The creators of GEM provide two baseline models to facilitate research and development in this area.

The goal of GEM is to advance the field of multimodal research by providing a comprehensive evaluation benchmark that spans vision and language modalities. Researchers can use this dataset to assess the capabilities of their models across different tasks and languages¹².

(1) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://arxiv.org/abs/2106.09889. (2) GEM Submission Instructions - GitHub Pages. https://microsoft.github.io/GEM/. (3) GEM: A General Evaluation Benchmark for Multimodal Tasks. https://www.microsoft.com/en-us/research/publication/gem-a-general-evaluation-benchmark-for-multimodal-tasks/. (4) undefined. https://doi.org/10.48550/arXiv.2106.09889.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • GEM (A General Evaluation Benchmark on Multi-modal Tasks)

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections