{"url":"/dataset/grit","name":"GRIT","full_name":"General Robust Image Task Benchmark","description_markdown":"The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources. GRIT hopes to encourage our research community to pursue the following research directions:\r\n\r\n1. **General purpose vision models** - GRIT facilitates the evaluation of unified and general-purpose vision models that demonstrate a wide range of skills across a diverse set of concepts.\r\n2. **Robust specialized models** - GRIT simplifies and unifies quantification of misinformation, calibration, and generalization under distribution shifts due to novel concepts, novel data sources or image distortions for 7 standard vision and vision-language tasks.\r\n3. **Efficient learning** - GRIT includes a `restricted` and an `unrestricted` track. The `restricted `track constrains the allowed training data to a selected but rich set of data sources that allows more scientific and meaningful comparison between models. This is meant to encourage resource constrained researchers to participate in the GRIT challenge and to spur interest in efficient learning methods as opposed to the dominant paradigm of training larger models on ever increasing amounts of training data. The `unrestricted` track allows much more flexibility in training data selection to test the capability of vision models trained with massive data and compute.","description_withheld":null,"homepage":"https://grit-benchmark.org/","introduced_date":"2022-04-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/grit-general-robust-image-task-benchmark","title":"GRIT: General Robust Image Task Benchmark","first_author":"Tanmay Gupta","url":null},"license":{"name":"Apache License 2.0","url":"https://github.com/allenai/grit_official/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Instance Segmentation","url":"/task/instance-segmentation","datasets_with_task":"/datasets/task/instance-segmentation"},{"name":"Visual Question Answering","url":"/task/visual-question-answering-1","datasets_with_task":"/datasets/task/visual-question-answering-1"},{"name":"Object Localization","url":"/task/object-localization","datasets_with_task":"/datasets/task/object-localization"},{"name":"Object Segmentation","url":"/task/object-segmentation","datasets_with_task":"/datasets/task/object-segmentation"},{"name":"Referring Expression Comprehension","url":"/task/referring-expression-comprehension","datasets_with_task":"/datasets/task/referring-expression-comprehension"},{"name":"Keypoint Detection","url":"/task/keypoint-detection","datasets_with_task":"/datasets/task/keypoint-detection"},{"name":"Surface Normals Estimation","url":"/task/surface-normals-estimation","datasets_with_task":"/datasets/task/surface-normals-estimation"},{"name":"Surface Normal Estimation","url":"/task/surface-normal-estimation","datasets_with_task":"/datasets/task/surface-normal-estimation"},{"name":"Object Categorization","url":"/task/object-categorization","datasets_with_task":"/datasets/task/object-categorization"},{"name":"Referring Expression","url":"/task/referring-expression","datasets_with_task":"/datasets/task/referring-expression"},{"name":"Keypoint Estimation","url":"/task/keypoint-estimation","datasets_with_task":"/datasets/task/keypoint-estimation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["GRIT"],"data_loaders":[],"num_papers_in_archive":16,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/object-categorization-on-grit","task":"Object Categorization","dataset_variant":"GRIT","rows":4,"metrics":["Categorization (ablation)","Categorization (test)"],"first_row_in_archive_order":{"model":"Unified-IOXL","paper":"/paper/unified-io-a-unified-model-for-vision","metrics":{"Categorization (ablation)":"61.7","Categorization (test)":"60.8"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/object-localization-on-grit","task":"Object Localization","dataset_variant":"GRIT","rows":3,"metrics":["Localization (ablation)","Localization (test)"],"first_row_in_archive_order":{"model":"Unified-IOXL","paper":"/paper/unified-io-a-unified-model-for-vision","metrics":{"Localization (ablation)":"67.0","Localization (test)":"67.1"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/object-segmentation-on-grit","task":"Object Segmentation","dataset_variant":"GRIT","rows":2,"metrics":["Segmentation (ablation)","Segmentation (test)"],"first_row_in_archive_order":{"model":"Unified-IOXL","paper":"/paper/unified-io-a-unified-model-for-vision","metrics":{"Segmentation (ablation)":"56.3","Segmentation (test)":"56.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-grit","task":"Visual Question Answering (VQA)","dataset_variant":"GRIT","rows":2,"metrics":["VQA (ablation)","VQA (test)"],"first_row_in_archive_order":{"model":"Unified-IOXL","paper":"/paper/unified-io-a-unified-model-for-vision","metrics":{"VQA (ablation)":"74.5","VQA (test)":"74.5"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-on-grit-1","task":"Visual Question Answering","dataset_variant":"GRIT","rows":1,"metrics":["VQA (ablation)"],"first_row_in_archive_order":{"model":"OFA","paper":"/paper/unifying-architectures-tasks-and-modalities","metrics":{"VQA (ablation)":"72.4"},"code_links":[{"title":"modelscope/modelscope","url":"https://github.com/modelscope/modelscope"},{"title":"ofa-sys/ofa","url":"https://github.com/ofa-sys/ofa"},{"title":"JHKim-snu/GVCCI","url":"https://github.com/JHKim-snu/GVCCI"},{"title":"JHKim-snu/PGA","url":"https://github.com/JHKim-snu/PGA"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/unified-io-a-unified-model-for-vision","title":"Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks","date":"2022-06-17","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/unifying-architectures-tasks-and-modalities","title":"OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework","date":"2022-02-07","rows_on_this_dataset":2,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/webly-supervised-concept-expansion-for","title":"Webly Supervised Concept Expansion for General Purpose Vision Models","date":"2022-02-04","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/learning-transferable-visual-models-from","title":"Learning Transferable Visual Models From Natural Language Supervision","date":"2021-02-26","rows_on_this_dataset":1,"code_links":82,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":20,"samples_ran":16,"samples_unverified":4,"pointer_only_for_licence":16,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mask-r-cnn","title":"Mask R-CNN","date":"2017-03-20","rows_on_this_dataset":2,"code_links":179,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":140,"samples_ran":42,"samples_unverified":98,"pointer_only_for_licence":23,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":161,"samples_ran":59,"samples_unverified":102,"pointer_only_for_licence":39,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}