Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 183 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8737–8784 of 12,172
A collection of prior text datasets assembled for hypothesis generation.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Introduced by Singh, Sumeet S..
1 paper · 1 benchmark
The image collection of the IAPR TC-12 Benchmark consists of 20,000 still natural images taken from locations around the world and comprising an assorted cross-section of still natural images.
1 paper · 0 benchmarks
Archivos con audios de toses de personas grabadas por celular, segmentados por COVID positivo y negativo según resultado de test RT-PCR.
1 paper · 0 benchmarks
The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g.
1 paper · 0 benchmarks
This dataset contains general and named entities annotations on both clean written text and on noisy speech data.
1 paper · 0 benchmarks
The IC13 dataset contains 561 images: 420 for training and 141 for testing.
1 paper · 1 benchmark
- Revision: v1.0.0-full-20210527a - DOI: 10.5281/zenodo.4817662 - Authors: J.
1 paper · 0 benchmarks
HIghly Heterogeneous KGs for entity alignment research.
1 paper · 0 benchmarks
The dataset is composed of Hematoxylin and eosin (H&E) stained breast histology microscopy and whole-slide images.
1 paper · 1 benchmark
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
ICM (Image-Caption Matching Dataset)
ICM is curated for the image-text matching task.
1 paper · 0 benchmarks
ICR (Image-Caption Retrieval Dataset)
In this dataset, we collect 200,000 image-text pairs.
1 paper · 0 benchmarks
Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for explainable driving decision-making and safe and efficient navigation.
1 paper · 0 benchmarks
IDE (Identifying spatio-temporal drivers of extreme events)
This data set allows to systematically evaluate approaches for the task of identifying anomalies and extreme events in water cycle components by developing deep neural networks that detect anomalies and drivers of extremes in simulated…
1 paper · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
The IDMT-SMT-Bass database is a large database for automatic bass transcription and signal processing.
1 paper · 0 benchmarks
We introduce the IDRCell100K image dataset, a collection of biological images, purposefully curated from the extensive and varied Image Data Resource platform.
1 paper · 0 benchmarks
The dataset is taken from the First shared task on Information Extractor for Conversational Systems in Indian Languages (IECSIL) .
1 paper · 1 benchmark
IEE is a financial-domain dataset of the Insurance-entity extraction task.
1 paper · 0 benchmarks
The Vesta dataset was released for use in the IEEE CIS Fraud Detection competition.
1 paper · 0 benchmarks
The IEEE Computational Intelligence Society ran a competition from July to November 2021 for predicting and optimizing based on renewable energy data.
1 paper · 0 benchmarks
IEIs (Ion and Electron Insulators)
We would like to introduce three types of ion and electron insulators, i.e.
1 paper · 0 benchmarks
IGC (Image-Grounded Conversation dataset)
111
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
IHEval (Evaluation on Instruction Hierarchy)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
IISc VEED (Indian Institute of Science Virtual Environment Exploration Database of Static Scenes)
IISc VEED consists of 200 diverse indoor and outdoor scenes (see samples below).
1 paper · 0 benchmarks
IISc VINE (Indian Institute of Science VIdeo Naturalness Evaluation)
Indian Institute of Science VIdeo Naturalness Evaluation (IISc VINE) is a database consisting of 300 videos, obtained by applying different prediction models on different datasets, and accompanying human opinion scores.
1 paper · 0 benchmarks
IIT-JEE (2009 IIT-JEE Scores Dataset)
Dataset consists of scores of 384,977 students in the Mathematics, Physics, and Chemistry sections of 2009 IIT-JEE (The Joint Entrance Exam of Indian Institutes of Technology), along with their gender, birth category, disability status,…
1 paper · 0 benchmarks
IITM-Bandersnatch is a dataset to evaluate traffic analysis techniques.
1 paper · 0 benchmarks
ILIAS (ILIAS: Instance-Level Image retrieval At Scale)
ILIAS is a large-scale test dataset for evaluation on Instance-Level Image retrieval At Scale.
1 paper · 0 benchmarks
A large dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
A small dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
A collection of test sets for evaluating base and chat LLMs (incl.
1 paper · 0 benchmarks
IM-SportingBehaviors Dataset Dataset Overview The IM-SportingBehaviors dataset, developed by researchers at Air University Pakistan, provides detailed motion data from participants engaged in various sports activities.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
IMDB-WIKI-SbS is a new large-scale dataset for evaluation pairwise comparisons, building on the success of a well-known benchmark for computer vision systems IMDB-WIKI.
1 paper · 0 benchmarks
IMFW (Indian Masked Faces In The Wild)
Indian Masked faces in the wild Database is collected into three sets:(i) Indian Celebrity, (ii) Instagram and (iii) Indian Crowd.
1 paper · 0 benchmarks
IMPACT Patent (A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents)
It is a large-scale multimodal patent dataset with detailed captions for design patent figures.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
IMaSC (ICFOSS Malayalam Speech Corpus)
IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech.
1 paper · 0 benchmarks
IN2LAAMA is a set of lidar-inertial datasets collected with a Velodyne VLP-16 lidar and a Xsens MTi-3 IMU.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
INDRA (INdian Dataset for RoAd crossing)
INDRA is a dataset capturing videos of Indian roads from the pedestrian point-of-view.
1 paper · 0 benchmarks
A representative event-based eye-tracking dataset, collected with two event cameras mounted on a glass frame.
1 paper · 2 benchmarks
The INRIA Dense Light Field Dataset (DLFD) is a dataset for testing depth estimation methods in a light field.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.