Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 184 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8785–8832 of 12,172
The Holidays dataset is a set of images which mainly contains some of the authors' personal holidays photos.
1 paper · 1 benchmark
The INRIA Sprse Light Field Dataset (SLFD) is a dataset for testing depth estimation methods in a light field.
1 paper · 0 benchmarks
This data set contains over 600GB of multimodal data from a Mars analog mission, including accurate 6DoF outdoor ground truth, indoor-outdoor transitions with continuous cross-domain ground truth, and indoor data with Optitrack…
1 paper · 0 benchmarks
This dataset contains 65 DFIs acquired from patients with POAG at the University of Iowa Hospitals and Clinics.
1 paper · 2 benchmarks
INSTANCE (the Italian seismic dataset for machine learning)
INSTANCE is a data collection of more than 1.3 million seismic waveforms originating from a selection of about 54,000 earthquakes occurred since 2005 in Italy and surrounding regions and seismic noise recordings randomly extracted from…
1 paper · 0 benchmarks
IPOD (Industrial and Professional Occupation Dataset)
Comprises 192k job titles belonging to 56k LinkedIn users.
1 paper · 0 benchmarks
These are the summary crosstabular data of the 2024 ISPSOS survey on which the paper, "Automation from the Worker's Perspective" is based.
1 paper · 0 benchmarks
IQM (Image-Query Matching Dataset)
IQM is curated for the image-text matching task.
1 paper · 0 benchmarks
IQR (Image-Query Retrieval Dataset)
IQR is proposed for the image-text retrieval task.
1 paper · 0 benchmarks
We establish the first large benchmark called IRBFD to facilitate the research in the area of nonuniformity correction and infrared UAV target detection, which consists of 50,000 manually labeled infrared images with various nonuniformity…
1 paper · 0 benchmarks
IRLCov19 is a multilingual Twitter dataset related to Covid-19 collected in the period between February 2020 to July 2020 specifically for regional languages in India.
1 paper · 0 benchmarks
IRMA (15,363 IRMA images of 193 categories for ImageCLEFmed 2009)
This collection compiles anonymous radiographs, which have been arbitrarly selected from routine at the Department of Diagnostic Radiology, Aachen University of Technology (RWTH), Aachen, Germany.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
IRV2V (IRregular V2V Dataset)
To facilitate research on asynchrony for collaborative perception, we simulate the first collaborative perception dataset with different temporal asynchronies based on CARLA, named IRregular V2V(IRV2V).
1 paper · 1 benchmark
We introduce a new synthetic test set named IS3 for interactive sound source localization.
1 paper · 0 benchmarks
This repository holds two datasets: one with both the original binaries and the code sections extracted from them (“full dataset”), and one with only the code sections (“only code sections”).
1 paper · 0 benchmarks
ISBNet is a dataset of images of recyclables.
1 paper · 1 benchmark
ISEKAI dataset’s images are generated by Midjourney’s text-to-image model using well-crafted instructions.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ISO17 (ISO17 - MD Trajectories of C7O2H10 with total energies and atomic forces)
Description The molecules were randomly drawn from the largest set of isomers in the QM9 dataset [1] which consists of molecules with a fixed composition of atoms (C7O2H10) arranged in different chemically valid structures.
1 paper · 0 benchmarks
ISOD (Indoor Small Object Dataset)
ISOD contains 2,000 manually labelled RGB-D images from 20 diverse sites, each featuring over 30 types of small objects randomly placed amidst the items already present in the scenes.
1 paper · 0 benchmarks
ISP-AD (The Industrial Screen Printing Anomaly Detection Dataset)
The ISP-AD Dataset is a large-scale anomaly detection dataset, representing a real-world industrial use case.
1 paper · 0 benchmarks
Contains 208,104 images with the same size of 10241024.
1 paper · 0 benchmarks
The ITCPR dataset is a comprehensive collection specifically designed for the Zero-Shot Composed Person Retrieval (ZS-CPR) task.
1 paper · 1 benchmark
ITD (Industrial Textile Dataset)
This dataset aims to provide a color dataset with real industrial fabric defect gathered in a visiting machine with several industrial cameras.
1 paper · 0 benchmarks
ITDD (Industrial Textile Defect Detection)
The Industrial Textile Defect Detection (ITDD) dataset includes 1885 industrial textile images categorized into 4 categories: cotton fabric, dyed fabric, hemp fabric, and plaid fabric.
1 paper · 2 benchmarks
IU ShareView dataset consists of 9 sets of two 5-10 minute first-person videos.
1 paper · 0 benchmarks
IUPAC Standards Online is a database built from IUPAC’s standards and recommendations, extracted from the journal Pure and Applied Chemistry (PAC).
1 paper · 0 benchmarks
The IUSTPersonReID dataset was developed to address limitations in existing person re-identification datasets by including cultural and environmental contexts unique to Islamic countries, especially Iran and Iraq.
1 paper · 1 benchmark
IVM-Mix-1M provide over 1M image-instruction pairs with corresponding instruction-relevant mask labels.
1 paper · 0 benchmarks
Ice Hockey News Dataset is a corpus of Finnish ice hockey news, edited to be suitable for training of end-to-end news generation methods, as well as demonstrate generation of text, which was judged by journalists to be relatively close to…
1 paper · 0 benchmarks
Icon645 is a large-scale dataset of icon images that cover a wide range of objects: 645,687 colored icons 377 different icon classes These collected icon classes are frequently mentioned in the IconQA questions.
1 paper · 0 benchmarks
After defining a taxonomy of the main stone deterioration patterns and anomalies, we selected 354 highly representative images of stone-built heritage, offering them a careful selection of labels to choose from.
1 paper · 1 benchmark
We release 280 synthetic IAM graphs generated using IAM graphs of commercial companies.
1 paper · 0 benchmarks
IgboNLP is a standard machine translation benchmark dataset for Igbo.
1 paper · 0 benchmarks
IllusionAnimalstest Dataset Characteristics IllusionAnimalstest is a generated dataset based on a synthetic collection of animal images, including 10 animal classes: cat, dog, pigeon, butterfly, elephant, horse, deer, snake, fish, and…
1 paper · 0 benchmarks
IllusionChartest Dataset Characteristics IllusionChartest is a generated dataset containing 3,300 samples of images that feature sequences of 3 to 5 random characters.
1 paper · 0 benchmarks
IllusionFashionMNISTtest Dataset Characteristics IllusionFashionMNISTtest is a generated dataset derived from the FashionMNIST dataset.
1 paper · 0 benchmarks
IllusionMNISTtest Dataset Characteristics IllusionMNISTtest is a generated dataset derived from the MNIST dataset.
1 paper · 0 benchmarks
Introduced by Singh, Sumeet S..
1 paper · 0 benchmarks
Image Caption Quality Dataset is a dataset of crowdsourced ratings for machine-generated image captions.
1 paper · 0 benchmarks
This publicly available dataset contains 1613 RGB-D images of field-grown broccoli plants.
1 paper · 0 benchmarks
This ImageNet version contains only 50 training images per class while the original testing set remains unchanged.
1 paper · 1 benchmark
This dataset was presented as part of the ICLR 2023 paper 𝘈 𝘧𝘳𝘢𝘮𝘦𝘸𝘰𝘳𝘬 𝘧𝘰𝘳 𝘣𝘦𝘯𝘤𝘩𝘮𝘢𝘳𝘬𝘪𝘯𝘨 𝘊𝘭𝘢𝘴𝘴-𝘰𝘶𝘵-𝘰𝘧-𝘥𝘪𝘴𝘵𝘳𝘪𝘣𝘶𝘵𝘪𝘰𝘯 𝘥𝘦𝘵𝘦𝘤𝘵𝘪𝘰𝘯 𝘢𝘯𝘥 𝘪𝘵𝘴 𝘢𝘱𝘱𝘭𝘪𝘤𝘢𝘵𝘪𝘰𝘯 𝘵𝘰 𝘐𝘮𝘢𝘨𝘦𝘕𝘦𝘵.
1 paper · 1 benchmark
The training and validation data are subsets of the training split of the Imagenet 2012.
1 paper · 0 benchmarks
This split was introduced in TEMP (BMVC 2023) Adaloglou, Nikolas, Felix Michels, Hamza Kalisch, and Markus Kollmann.
1 paper · 0 benchmarks
We build a new evaluation set by adding spotting words to the images of ImageNet 2012 evaluation sets.
1 paper · 0 benchmarks
A dataset of A 3D Computed Tomography (CT) image dataset, ImageTBAD, for segmentation of Type-B Aortic Dissection is published.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.