Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 19 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 865–912 of 12,172
ScanQA (ScanQA: 3D Question Answering for Spatial Scene Understanding)
We collected 41,363 questions and 58,191 answers, in- cluding 32,337 unique questions and 16,999 unique an- swers.
70 papers · 1 benchmark
The UT-Kinect dataset is a dataset for action recognition from depth sequences.
70 papers · 2 benchmarks
This YouTube dataset is a sampling from thousands of User Generated Content (UGC) as uploaded to YouTube distributed under the Creative Commons license.
70 papers · 1 benchmark
The data is related with direct marketing campaigns (phone calls) of a Portuguese banking institution.
69 papers · 0 benchmarks
The BraTS 2015 dataset is a dataset for brain tumor image segmentation.
69 papers · 1 benchmark
CACD (Cross-Age Celebrity Dataset)
The Cross-Age Celebrity Dataset (CACD) contains 163,446 images from 2,000 celebrities collected from the Internet.
69 papers · 1 benchmark
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
Multimodal Opinionlevel Sentiment Intensity (MOSI) contains: (1) multimodal observations including transcribed speech and visual gestures as well as automatic audio and visual features, (2) opinion-level subjectivity segmentation, (3)…
69 papers · 1 benchmark
QMSum is a new human-annotated benchmark for query-based multi-domain meeting summarisation task, which consists of 1,808 query-summary pairs over 232 meetings in multiple domains.
69 papers · 1 benchmark
RegDB (Dongguk Body-based Person Recognition Database (DBPerson-Recog-DB1))
RegDB is used for Visible-Infrared Re-ID which handles the cross-modality matching between the daytime visible and night-time infrared images.
69 papers · 2 benchmarks
UBFC-rPPG (Univ. Bourgogne Franche-Comté Remote PhotoPlethysmoGraphy)
We introduce here a new database called UBFC-rPPG (stands for Univ.
69 papers · 1 benchmark
VSR (Visual Spatial Reasoning)
The Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels.
69 papers · 1 benchmark
WikiLarge comprise 359 test sentences, 2000 development sentences and 300k training sentences.
69 papers · 0 benchmarks
AGORA is a synthetic human dataset with high realism and accurate ground truth.
68 papers · 4 benchmarks
BeaverTails is a dataset aimed at fostering research on safety alignment in large language models (LLMs).
68 papers · 0 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
CN-Celeb is a large-scale speaker recognition dataset collected in the wild'.
68 papers · 1 benchmark
CoNSeP (Colorectal Nuclear Segmentation and Phenotypes)
The colorectal nuclear segmentation and phenotypes (CoNSeP) dataset consists of 41 H&E stained image tiles, each of size 1,000×1,000 pixels at 40× objective magnification.
68 papers · 2 benchmarks
DREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset.
68 papers · 2 benchmarks
FSOD (Few-Shot Object Detection Dataset)
Few-Shot Object Detection Dataset (FSOD) is a high-diverse dataset specifically designed for few-shot object detection and intrinsically designed to evaluate thegenerality of a model on novel categories.
68 papers · 0 benchmarks
HaluEval is a large-scale hallucination evaluation benchmark designed for Large Language Models (LLMs).
68 papers · 0 benchmarks
LIVE-VQC (LIVE Video Quality Challenge (VQC) Database)
The great variations of videographic skills in videography, camera designs, compression and processing protocols, communication and bandwidth environments, and displays leads to an enormous variety of video impairments.
68 papers · 1 benchmark
The PlantVillage dataset consists of 54303 healthy and unhealthy leaf images divided into 38 categories by species and disease.
68 papers · 1 benchmark
Recipe1M+ is a dataset which contains one million structured cooking recipes with 13M associated images.
68 papers · 3 benchmarks
HatEval (SemEval 2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter)
Hate Speech is commonly defined as any communication that disparages a person or a group on the basis of some characteristic such as race, color, ethnicity, gender, sexual orientation, nationality, religion, or other characteristics.
67 papers · 1 benchmark
HSOL is a dataset for hate speech detection.
67 papers · 0 benchmarks
InLoc is a dataset with reference 6DoF poses for large-scale indoor localization.
67 papers · 1 benchmark
MOT15 (Multiple Object Tracking 15)
MOT2015 is a dataset for multiple object tracking.
67 papers · 5 benchmarks
REAL275 is a benchmark for category-level pose estimation.
67 papers · 1 benchmark
T2I-CompBench is a comprehensive benchmark for open-world compositional text-to-image generation, consisting of 6,000 compositional textual prompts from 3 categories (attribute binding, object relationships, and complex compositions) and 6…
67 papers · 1 benchmark
The original ionosphere dataset from UCI machine learning repository is a binary classification dataset with dimensionality 34.
67 papers · 2 benchmarks
CFQ (Compositional Freebase Questions)
A large and realistic natural language question answering dataset.
66 papers · 1 benchmark
This work presents two new benchmark datasets (CIFAR-10N, CIFAR-100N), equipping the training dataset of CIFAR-10 and CIFAR-100 with human-annotated real-world noisy labels that we collect from Amazon Mechanical Turk.
66 papers · 1 benchmark
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
This project contains natural language data for human-robot interaction in home domain which we collected and annotated for evaluating NLU Services/platforms.
66 papers · 3 benchmarks
Contains 4,480 Wikipedia documents, 118,732 event mention instances, and 168 event types.
66 papers · 0 benchmarks
The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset.
66 papers · 5 benchmarks
NAB (Numenta Anomaly Benchmark)
The First Temporal Benchmark Designed to Evaluate Real-time Anomaly Detectors Benchmark The growth of the Internet of Things has created an abundance of streaming data.
66 papers · 1 benchmark
Natural-Instructions is a dataset of 61 distinct tasks, their human-authored instructions and 193k task instances.
66 papers · 0 benchmarks
This dataset focuses on heavily occluded human with comprehensive annotations including bounding-box, humans pose and instance mask.
66 papers · 6 benchmarks
Text corpus with almost one billion words of training data for statistical language modeling benchmarking.
66 papers · 0 benchmarks
ParaCrawl v.7.1 is a parallel dataset with 41 language pairs primarily aligned with English (39 out of 41) and mined using the parallel-data-crawling tool Bitextor which includes downloading documents, preprocessing and normalization,…
66 papers · 0 benchmarks
The Reuters-21578 dataset is a collection of documents with news articles.
66 papers · 5 benchmarks
Semantic3D is a point cloud dataset of scanned outdoor scenes with over 3 billion points.
66 papers · 1 benchmark
A knowledge-grounded human-human conversation dataset where the underlying knowledge spans 8 broad topics and conversation partners don’t have explicitly defined roles.
66 papers · 0 benchmarks
WLASL (Word-Level American Sign Language)
WLASL is a large video dataset for Word-Level American Sign Language (ASL) recognition, which features 2,000 common different words in ASL.
66 papers · 3 benchmarks
WikiHop is a multi-hop question-answering dataset.
66 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.