Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 44 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2065–2112 of 12,172
One of the largest commonsense knowledge bases available, describing over 2 million disambiguated concepts and activities, connected by over 18 million assertions.
20 papers · 0 benchmarks
WebFace260M is a million-scale face benchmark, which is constructed for the research community towards closing the data gap behind the industry.
20 papers · 0 benchmarks
XFUND (A Multilingual Form Understanding Benchmark)
XFUND is a multilingual form understanding benchmark dataset that includes human-labeled forms with key-value pairs in 7 languages (Chinese, Japanese, Spanish, French, Italian, German, Portuguese).
20 papers · 0 benchmarks
YAGO3-10 (Yet Another Great Ontology 3-10)
YAGO3-10 is benchmark dataset for knowledge base completion.
20 papers · 1 benchmark
The iKala dataset is a singing voice separation dataset that comprises of 252 30-second excerpts sampled from 206 iKala songs (plus 100 hidden excerpts reserved for MIREX data mining contest).
20 papers · 1 benchmark
The iLIDS-VID dataset is a person re-identification dataset which involves 300 different pedestrians observed across two disjoint camera views in public open space.
20 papers · 2 benchmarks
The Iris flower data set or Fisher's Iris data set is a multivariate data set introduced by the British statistician, eugenicist, and biologist Ronald Fisher in his 1936 paper The use of multiple measurements in taxonomic problems as an…
20 papers · 8 benchmarks
AGENDA (Abstract GENeration DAtaset)
Abstract GENeration DAtaset (AGENDA) is a dataset of knowledge graphs paired with scientific abstracts.
19 papers · 1 benchmark
The Abt-Buy dataset for entity resolution derives from the online retailers Abt.com and Buy.com.
19 papers · 2 benchmarks
Amazon Clothing (Amazon Clothing 5-core)
19 papers · 1 benchmark
Argoverse-HD is a dataset built for streaming object detection, which encompasses real-time object detection, video object detection, tracking, and short-term forecasting.
19 papers · 4 benchmarks
BCI (Breast Cancer Immunohistochemical Image Generation)
The evaluation of human epidermal growth factor receptor 2 (HER2) expression is essential to formulate a precise treatment for breast cancer.
19 papers · 1 benchmark
BRATS 2021 (RSNA-ASNR-MICCAI Brain Tumor Segmentation (BraTS) Challenge 2021)
The RSNA-ASNR-MICCAI BraTS 2021 challenge utilizes multi-institutional pre-operative baseline multi-parametric magnetic resonance imaging (mpMRI) scans, and focuses on the evaluation of state-of-the-art methods for (Task 1) the…
19 papers · 0 benchmarks
The purpose of this dataset was to study gender bias in occupations.
19 papers · 1 benchmark
CMD (Condensed Movies Dataset)
Consists of the key scenes from over 3K movies: each key scene is accompanied by a high level semantic description of the scene, character face-tracks, and metadata about the movie.
19 papers · 0 benchmarks
CValues is a Chinese human values evaluation benchmark designed to assess the alignment of Chinese Large Language Models (LLMs) with human values.
19 papers · 0 benchmarks
Ten years (2008-2018) ChFinAnn documents and human-summarized event knowledge bases to conduct the DS-based event labeling.
19 papers · 1 benchmark
Chaoyang dataset contains 1111 normal, 842 serrated, 1404 adenocarcinoma, 664 adenoma, and 705 normal, 321 serrated, 840 adenocarcinoma, 273 adenoma samples for training and testing, respectively.
19 papers · 2 benchmarks
CoAuthor is a dataset designed for revealing GPT-3's capabilities in assisting creative and argumentative writing.
19 papers · 1 benchmark
The CoNLL04 dataset is a benchmark dataset used for relation extraction tasks.
19 papers · 3 benchmarks
CRONQUESTIONS, the Temporal KGQA dataset consists of two parts: a KG with temporal annotations, and a set of natural language questions requiring temporal reasoning.
19 papers · 1 benchmark
CurveLanes is a new benchmark lane detection dataset with 150K lanes images for difficult scenarios such as curves and multi-lanes in traffic lane detection.
19 papers · 1 benchmark
DDXPlus (DDXPlus: A New Dataset For Automatic Medical Diagnosis)
There has been a rapidly growing interest in Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the machine learning research literature, aiming to assist doctors in telemedicine services.
19 papers · 0 benchmarks
DELIVER is an arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB.
19 papers · 3 benchmarks
DIRHA (Distant-speech Interaction for Robust Home Applications)
DIRHA-English is a multi-microphone database composed of real and simulated sequences of 1-minute.
19 papers · 1 benchmark
DeepFish as a benchmark suite with a large-scale dataset to train and test methods for several computer vision tasks.
19 papers · 1 benchmark
The National Institutes of Health’s Clinical Center has made a large-scale dataset of CT images publicly available to help the scientific community improve detection accuracy of lesions.
19 papers · 1 benchmark
Node classification on Deezer Europe with 50%/25%/25% random splits for training/validation/test.
19 papers · 1 benchmark
Dinstinctions-646 are composed of 646 foreground images with manually annotated alpha mattes
19 papers · 1 benchmark
FBMS-59 (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is a dataset for motion segmentation, which extends the BMS-26 dataset with 33 additional video sequences.
19 papers · 3 benchmarks
FNC-1 (Fake News Challenge Stage 1)
FNC-1 was designed as a stance detection dataset and it contains 75,385 labeled headline and article pairs.
19 papers · 2 benchmarks
The FSDnoisy18k dataset is an open dataset containing 42.5 hours of audio across 20 sound event classes, including a small amount of manually-labeled data and a larger quantity of real-world noisy data.
19 papers · 0 benchmarks
Node classification on Film with 60%/20%/20% random splits for training/validation/test.
19 papers · 1 benchmark
FoCus (Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge)
We introduce a new dataset, called FoCus, that supports knowledge-grounded answers that reflect user’s persona.
19 papers · 0 benchmarks
FoodSeg103 is a new food image dataset containing 7,118 images.
19 papers · 1 benchmark
Funcom is a collection of ~2.1 million Java methods and their associated Javadoc comments.
19 papers · 0 benchmarks
H3DS a high-resolution 3D full head textured scans and 360º images dataset collected with a structured light scanner, consisting of 23 3D full-head scans containing images, masks and camera poses.
19 papers · 0 benchmarks
Hi4D contains 4D textured scans of 20 subject pairs, 100 sequences, and a total of more than 11K frames.
19 papers · 0 benchmarks
A new large-scale dataset for understanding human motions, poses, and actions in a variety of realistic events, especially crowd & complex events.
19 papers · 1 benchmark
The data set contains 33 patches (of different sizes), each consisting of a true orthophoto (TOP) extracted from a larger TOP mosaic.
19 papers · 1 benchmark
IU X-ray (Demner-Fushman et al., 2016) is a set of chest X-ray images paired with their corresponding diagnostic reports.
19 papers · 2 benchmarks
A benchmark dataset for out-of-distribution detection.
19 papers · 1 benchmark
InfiMM-Eval (Complex Open-ended Reasoning Evaluation for Multi-Modal Language Models)
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence.
19 papers · 1 benchmark
This data set provides Light Detection and Ranging (LiDAR) data and stereo image with various position sensors targeting a highly complex urban environment.
19 papers · 0 benchmarks
KinFaceW-I dataset contains 533 pairs of facial images of persons with a kin relation.
19 papers · 1 benchmark
Consists of annotated frames containing GI procedure tools such as snares, balloons and biopsy forceps, etc.
19 papers · 3 benchmarks
A suite of open-source federated datasets, a rigorous evaluation framework, and a set of reference implementations, all geared towards capturing the obstacles and intricacies of practical federated environments.
19 papers · 0 benchmarks
LVOS is a dataset for long-term video object segmentation (VOS).
19 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.