Datasets › GRIT
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources. GRIT hopes to encourage our research community to pursue the following research directions:
- General purpose vision models - GRIT facilitates the evaluation of unified and general-purpose vision models that demonstrate a wide range of skills across a diverse set of concepts.
- Robust specialized models - GRIT simplifies and unifies quantification of misinformation, calibration, and generalization under distribution shifts due to novel concepts, novel data sources or image distortions for 7 standard vision and vision-language tasks.
- Efficient learning - GRIT includes a
restrictedand anunrestrictedtrack. Therestrictedtrack constrains the allowed training data to a selected but rich set of data sources that allows more scientific and meaningful comparison between models. This is meant to encourage resource constrained researchers to participate in the GRIT challenge and to spur interest in efficient learning methods as opposed to the dominant paradigm of training larger models on ever increasing amounts of training data. Theunrestrictedtrack allows much more flexibility in training data selection to test the capability of vision models trained with massive data and compute.
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Object Categorization | GRIT | Unified-IOXL Categorization (ablation) 61.7 | Unified-IO: A Unified Model for Vision, Language, and... | — | 4 | Compare |
| Object Localization | GRIT | Unified-IOXL Localization (ablation) 67.0 | Unified-IO: A Unified Model for Vision, Language, and... | — | 3 | Compare |
| Object Segmentation | GRIT | Unified-IOXL Segmentation (ablation) 56.3 | Unified-IO: A Unified Model for Vision, Language, and... | — | 2 | Compare |
| Visual Question Answering (VQA) | GRIT | Unified-IOXL VQA (ablation) 74.5 | Unified-IO: A Unified Model for Vision, Language, and... | — | 2 | Compare |
| Visual Question Answering | GRIT | OFA VQA (ablation) 72.4 | OFA: Unifying Architectures, Tasks, and Modalities... | modelscope/modelscope +3 | 1 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 16. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks | 0 | 4 | 17 Jun 2022 | not harvested |
| OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework | 4 | 2 | 7 Feb 2022 | ran 1 of 1 samples (0 unverified) |
| Webly Supervised Concept Expansion for General Purpose Vision Models | 0 | 3 | 4 Feb 2022 | not harvested |
| Learning Transferable Visual Models From Natural Language Supervision | 82 | 1 | 26 Feb 2021 | ran 16 of 20 samples (4 unverified; 16 pointer-only for licence) |
| Mask R-CNN | 179 | 2 | 20 Mar 2017 | ran 42 of 140 samples (98 unverified; 23 pointer-only for licence) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- GRIT
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections