12,172 datasets listed, ordered by the archive's paper count. Page 14 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Cambridge Landmarks, a large scale outdoor visual relocalisation dataset taken around Cambridge University.
109 papers · 0 benchmarks
The Chairs dataset contains rendered images of around 1000 different three-dimensional chair models.
109 papers · 1 benchmark
The Evaluation framework of Raganato et al.
109 papers · 3 benchmarks
CommonGen is constructed through a combination of crowdsourced and existing caption corpora, consists of 79k commonsense descriptions over 35k unique concept-sets.
108 papers · 1 benchmark
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
WFLW (Wider Facial Landmarks in the Wild)
The Wider Facial Landmarks in the Wild or WFLW database contains 10000 faces (7500 for training and 2500 for testing) with 98 annotated landmarks.
108 papers · 4 benchmarks
iSUN is a ground truth of gaze traces on images from the SUN dataset.
108 papers · 0 benchmarks
Rationale and objectives: Computer-aided detection and diagnosis (CAD) systems have been developed in the past two decades to assist radiologists in the detection and diagnosis of lesions seen on breast imaging exams, thus providing a…
107 papers · 3 benchmarks
The LIMA dataset is a valuable resource used in natural language processing (NLP) research.
107 papers · 0 benchmarks
MGSM (Multilingual Grade School Math)
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems.
107 papers · 1 benchmark
CORNELL NEWSROOM is a large dataset for training and evaluating summarization systems.
107 papers · 0 benchmarks
The RVL-CDIP dataset consists of scanned document images belonging to 16 classes such as letter, form, email, resume, memo, etc.
107 papers · 2 benchmarks
SegTrack v2 is a video segmentation dataset with full pixel-level annotations on multiple objects at each frame within each video.
107 papers · 5 benchmarks
Sentiment analysis is increasingly viewed as a vital task both from an academic and a commercial standpoint.
107 papers · 4 benchmarks
A dataset for robot navigation task and more.
107 papers · 0 benchmarks
VIST (Visual Storytelling)
The Visual Storytelling Dataset (VIST) consists of 210,819 unique photos and 50,000 stories.
107 papers · 2 benchmarks
The MUSDB18 is a dataset of 150 full lengths music tracks (~10h duration) of different genres along with their isolated drums, bass, vocals and others stems.
106 papers · 2 benchmarks
The Newsela dataset was introduced by Xu et al.
106 papers · 1 benchmark
SLURP (Spoken Language Understanding Resource Package)
A new challenging dataset in English spanning 18 domains, which is substantially bigger and linguistically more diverse than existing datasets.
106 papers · 2 benchmarks
The Stylized-ImageNet dataset is created by removing local texture cues in ImageNet while retaining global shape information on natural images via AdaIN style transfer.
106 papers · 1 benchmark
Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse…
105 papers · 1 benchmark
BlendedMVS is a novel large-scale dataset, to provide sufficient training ground truth for learning-based MVS.
105 papers · 0 benchmarks
The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.
105 papers · 2 benchmarks
Covers multiple aspects of the issue.
105 papers · 2 benchmarks
JIGSAWS (JHU-ISI Gesture and Skill Assessment Working Set)
The JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) is a surgical activity dataset for human motion modeling.
105 papers · 3 benchmarks
MD17 (Molecular Dynamics 17)
Energies and forces for molecular dynamics trajectories of eight organic molecules.
105 papers · 0 benchmarks
ReDial (Recommendation Dialogues) is an annotated dataset of dialogues, where users recommend movies to each other.
105 papers · 2 benchmarks
RewardBench is a benchmark designed to evaluate the capabilities and safety of reward models, including those trained with Direct Preference Optimization (DPO).
105 papers · 0 benchmarks
Consists of a dataset with 1000 whole scanned receipt images and annotations for the competition on scanned receipts OCR and key information extraction (SROIE).
105 papers · 2 benchmarks
ToolBench is an instruction-tuning dataset for tool use, which is created automatically using ChatGPT.
105 papers · 1 benchmark
The ABC Dataset is a collection of one million Computer-Aided Design (CAD) models for research of geometric deep learning methods and applications.
104 papers · 0 benchmarks
The BP4D-Spontaneous dataset is a 3D video database of spontaneous facial expressions in a diverse group of young adults.
104 papers · 3 benchmarks
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
ReferIt3D provides two large-scale and complementary visio-linguistic datasets: i) Sr3D, which contains 83.5K template-based utterances leveraging spatial relations among fine-grained object classes to localize a referred object in a…
104 papers · 1 benchmark
AVE (Audio-Visual Event Localization)
To investigate three temporal localization tasks: supervised and weakly-supervised audio-visual event localization, and cross-modality localization.
103 papers · 0 benchmarks
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
DeepMIMO is a generic dataset for mmWave/massive MIMO channels.
103 papers · 0 benchmarks
GYAFC (Grammarly’s Yahoo Answers Formality Corpus)
Grammarly’s Yahoo Answers Formality Corpus (GYAFC) is the largest dataset for any style containing a total of 110K informal / formal sentence pairs.
103 papers · 2 benchmarks
The PoseTrack dataset is a large-scale benchmark for multi-person pose estimation and tracking in videos.
103 papers · 5 benchmarks
The image dataset TinyImages contains 80 million images of size 32×32 collected from the Internet, crawling the words in WordNet.
103 papers · 0 benchmarks
As autonomous driving systems mature, motion forecasting has received increasing attention as a critical requirement for planning.
103 papers · 0 benchmarks
Uses structured and unstructured data.
102 papers · 0 benchmarks
CosmosQA is a large-scale dataset of 35.6K problems that require commonsense-based reading comprehension, formulated as multiple-choice questions.
102 papers · 0 benchmarks
FRGC (Face Recognition Grand Challenge)
The data for FRGC consists of 50,000 recordings divided into training and validation partitions.
102 papers · 1 benchmark
Fashion IQ support and advance research on interactive fashion image retrieval.
102 papers · 6 benchmarks
GSO (Google Scanned Objects)
Scanned Objects by Google Research is a dataset of common household objects that have been 3D scanned for use in robotic simulation and synthetic perception research.
102 papers · 1 benchmark
QASPER is a dataset for question answering on scientific research papers.
102 papers · 1 benchmark
The Synthetic Rain Datasets consists of 13,712 clean-rain image pairs gathered from multiple datasets (Rain14000, Rain1800, Rain800, Rain12).
102 papers · 6 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.