Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 35 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1633–1680 of 12,172
A realistic dataset composed of 27 episodes from 6 popular TV series.
29 papers · 1 benchmark
Includes 5000 spatially aligned RGBT image pairs with ground truth annotations.
29 papers · 0 benchmarks
The VidSTG dataset is a spatio-temporal video grounding dataset constructed based on the video relation dataset VidOR.
29 papers · 1 benchmark
WI-LOCNESS (Cambridge English Write & Improve & LOCNESS)
WI-LOCNESS is part of the Building Educational Applications 2019 Shared Task for Grammatical Error Correction.
29 papers · 2 benchmarks
ADNI (Alzheimer's Disease NeuroImaging Initiative)
Alzheimer's Disease Neuroimaging Initiative (ADNI) is a multisite study that aims to improve clinical trials for the prevention and treatment of Alzheimer’s disease (AD).[1] This cooperative study combines expertise and funding from the…
28 papers · 5 benchmarks
Inspired by cognitive development studies on intuitive psychology, we present a benchmark consisting of a large dataset of procedurally generated 3D animations, AGENT (Action, Goal, Efficiency, coNstraint, uTility), structured around four…
28 papers · 0 benchmarks
The BBBP dataset comes from a study focused on modeling and predicting the permeability of the blood-brain barrier.
28 papers · 5 benchmarks
Semi-Supervised Object Detection on COCO 10% labeled data
28 papers · 2 benchmarks
Contains 349 COVID-19 CT images from 216 patients and 463 non-COVID-19 CTs.
28 papers · 0 benchmarks
CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors.
28 papers · 7 benchmarks
CelebA-Spoof is a large-scale face anti-spoofing dataset with the following properties: 1.
28 papers · 0 benchmarks
A human-to-human Chinese dialog dataset (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).
28 papers · 0 benchmarks
Dynamic FAUST extends the FAUST dataset to dynamic 4D data.
28 papers · 1 benchmark
EVALution dataset is evenly distributed among the three classes (hypernyms, co-hyponyms and random) and involves three types of parts of speech (noun, verb, adjective).
28 papers · 0 benchmarks
Dataset with 28,792 retinal images from the EyePACS dataset, based on a three-level quality grading system (i.e., Good', Usable' and Reject') for evaluating RIQA methods.
28 papers · 0 benchmarks
The Gaming 3D Dataset (G3D) focuses on real-time action recognition in a gaming scenario.
28 papers · 2 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
IndicCorp is a large monolingual corpora with around 9 billion tokens covering 12 of the major Indian languages.
28 papers · 0 benchmarks
KAIST (High-quality hyperspectral reconstruction using a spectral prior)
High-quality hyperspectral reconstruction using a spectral prior
28 papers · 1 benchmark
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
KaggleDBQA (KaggleDBQA: Realistic Text-to-SQL dataset)
KaggleDBQA is a challenging cross-domain and complex evaluation dataset of real Web databases, with domain-specific data types, original formatting, and unrestricted questions.
28 papers · 1 benchmark
MHIST (Minimalist Histopathology image analysis dataset)
The minimalist histopathology image analysis dataset (MHIST) is a binary classification dataset of 3,152 fixed-size images of colorectal polyps, each with a gold-standard label determined by the majority vote of seven board-certified…
28 papers · 1 benchmark
The MIT-Adobe FiveK dataset consists of 5,000 photographs taken with SLR cameras by a set of different photographers.
28 papers · 4 benchmarks
MP20 (Metastable crystal structures from Materials Project)
MP20 (Xie et al., 2022) contains 45,231 metastable crystal structures from the Materials Project (Jain et al., 2013), each with up to 20 atoms and spanning 89 different element types.
28 papers · 1 benchmark
MR Movie Reviews is a dataset for use in sentiment-analysis experiments.
28 papers · 3 benchmarks
MVSEC (Multi Vehicle Stereo Event Camera)
The Multi Vehicle Stereo Event Camera (MVSEC) dataset is a collection of data designed for the development of novel 3D perception algorithms for event based cameras.
28 papers · 2 benchmarks
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
MedICaT is a dataset of medical images, captions, subfigure-subcaption annotations, and inline textual references.
28 papers · 0 benchmarks
MetaShift is a collection of 12,868 sets of natural images across 410 classes.
28 papers · 0 benchmarks
Dataset is constructed from single intent dataset ATIS.
28 papers · 2 benchmarks
MuseData is an electronic library of orchestral and piano classical music from CCARH.
28 papers · 0 benchmarks
OASIS (Open Annotations of Single Image Surfaces)
A dataset for single-image 3D in the wild consisting of annotations of detailed 3D geometry for 140,000 images.
28 papers · 3 benchmarks
The Oxford Radar RobotCar Dataset is a radar extension to The Oxford RobotCar Dataset.
28 papers · 4 benchmarks
Consists of parallel sentences which pair 13 major languages of India with English.
28 papers · 0 benchmarks
PSG dataset has 48749 images with 133 object classes (80 objects and 53 stuff) and 56 predicate classes.
28 papers · 1 benchmark
Resume contains eight fine-grained entity categories -score from 74.5% to 86.88%.
28 papers · 1 benchmark
An open database for sharing robotic experience, which provides an initial pool of 15 million video frames, from 7 different robot platforms, and study how it can be used to learn generalizable models for vision-based robotic manipulation.
28 papers · 0 benchmarks
SBU-Kinect-Interaction dataset version 2.0 comprises of RGB-D video sequences of humans performing interaction activities that are recording using the Microsoft Kinect sensor.
28 papers · 4 benchmarks
SafetyBench is a comprehensive benchmark designed to evaluate the safety of large language models (LLMs) using multiple-choice questions.
28 papers · 0 benchmarks
SegTHOR (Segmentation of THoracic Organs at Risk)
SegTHOR (Segmentation of THoracic Organs at Risk) is a dataset dedicated to the segmentation of organs at risk (OARs) in the thorax, i.e.
28 papers · 0 benchmarks
The SensatUrbat dataset is an urban-scale photogrammetric point cloud dataset with nearly three billion richly annotated points, which is five times the number of labeled points than the existing largest point cloud dataset.
28 papers · 1 benchmark
The ShareGPT4Video dataset is a large-scale resource designed to improve video understanding and generation¹.
28 papers · 0 benchmarks
THuman2.0 Dataset contains 500 high-quality human scans captured by a dense DLSR rig.
28 papers · 1 benchmark
TextZoom is a super-resolution dataset that consists of paired Low Resolution – High Resolution scene text images.
28 papers · 2 benchmarks
TvSum (TVSum: Summarizing Web Videos Using Titles)
Introduced by Song et al.
28 papers · 4 benchmarks
WNUT-2020 Task 2 (WNUT-2020 Task 2: Identification of Informative COVID-19 English Tweets)
Briefly describe the dataset.
28 papers · 1 benchmark
WPC (Waterloo Point Cloud)
The WPC (Waterloo Point Cloud) database is a dataset for subjective and objective quality assessment of point clouds.
28 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.