Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 21 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 961–1008 of 12,172
The UTD-MHAD dataset consists of 27 different actions performed by 8 subjects.
62 papers · 2 benchmarks
BTAD (beanTech Anomaly Detection)
The BTAD ( beanTech Anomaly Detection) dataset is a real-world industrial anomaly dataset.
61 papers · 2 benchmarks
CIFAR-10H is a new dataset of soft labels reflecting human perceptual uncertainty for the 10,000-image CIFAR-10 test set.
61 papers · 0 benchmarks
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
61 papers · 3 benchmarks
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
GAIA (a benchmark for general AI assistants)
We introduce GAIA, a benchmark for General AI Assistants that, if solved, would represent a milestone in AI research.
61 papers · 0 benchmarks
Game of 24 is a mathematical reasoning challenge, where the goal is to use 4 numbers and basic arithmetic operations (+-/) to obtain 24.
61 papers · 1 benchmark
The LIP (Look into Person) dataset is a large-scale dataset focusing on semantic understanding of a person.
61 papers · 1 benchmark
The MSVD-QA dataset is a Video Question Answering (VideoQA) dataset.
61 papers · 5 benchmarks
PanNuke is a semi automatically generated nuclei instance segmentation and classification dataset with exhaustive nuclei labels across 19 different tissue types.
61 papers · 4 benchmarks
SFEW (Static Facial Expression in the Wild)
The Static Facial Expressions in the Wild (SFEW) dataset is a dataset for facial expression recognition.
61 papers · 1 benchmark
The Salient Person dataset (SIP) contains 929 salient person samples with different poses and illumination conditions.
61 papers · 1 benchmark
The WebQuestionsSP dataset is released as part of our ACL-2016 paper “The Value of Semantic Parse Labeling for Knowledge Base Question Answering” [Yih, Richardson, Meek, Chang & Suh, 2016], in which we evaluated the value of gathering…
61 papers · 3 benchmarks
Brightkite was once a location-based social networking service provider where users shared their locations by checking-in.
60 papers · 0 benchmarks
GAP (GAP Benchmark Suite)
GAP is a graph processing benchmark suite with the goal of helping to standardize graph processing evaluations.
60 papers · 1 benchmark
Node classification on Penn94
60 papers · 2 benchmarks
SummScreen is a dataset for abstractive screenplay summarization.
60 papers · 1 benchmark
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
ToTTo is an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.
60 papers · 1 benchmark
Subset and preprocessed version of Chemical reactions from US patents (1976-Sep2016) by Daniel Lowe.
60 papers · 1 benchmark
VOCASET is a 4D face dataset with about 29 minutes of 4D scans captured at 60 fps and synchronized audio.
60 papers · 1 benchmark
CELEX database comprises three different searchable lexical databases, Dutch, English and German.
59 papers · 0 benchmarks
CHASEDB1 is a dataset for retinal vessel segmentation which contains 28 color retina images with the size of 999×960 pixels which are collected from both left and right eyes of 14 school children.
59 papers · 2 benchmarks
A collection of 10 pre-processed medical open datasets.
59 papers · 0 benchmarks
The Middlebury 2014 dataset contains a set of 23 high resolution stereo pairs for which known camera calibration parameters and ground truth disparity maps obtained with a structured light scanner are available.
59 papers · 2 benchmarks
The Oxford-IIIT Pet Dataset is a 37-category pet dataset with roughly 200 images for each class.
59 papers · 5 benchmarks
SParC (Semantic Parsing in Context)
SParC is a large-scale dataset for complex, cross-domain, and context-dependent (multi-turn) semantic parsing and text-to-SQL task (interactive natural language interfaces for relational databases).
59 papers · 2 benchmarks
Set11 is a dataset of 11 grayscale images.
59 papers · 1 benchmark
WOS (Web of Science Dataset)
Web of Science (WOS) is a document classification dataset that contains 46,985 documents with 134 categories which include 7 parents categories.
59 papers · 4 benchmarks
BLINK is a new benchmark for multimodal language models (LLMs) that focuses on core visual perception abilities not found in other evaluations¹².
58 papers · 1 benchmark
CANARD (A Dataset for Question-in-Context Rewriting)
CANARD is a dataset for question-in-context rewriting that consists of questions each given in a dialog context together with a context-independent rewriting of the question.
58 papers · 1 benchmark
CASIA-MFSD is a dataset for face anti-spoofing.
58 papers · 1 benchmark
CCNet is a dataset extracted from Common Crawl with a different filtering process than for OSCAR.
58 papers · 0 benchmarks
ETH is a dataset for pedestrian detection.
58 papers · 3 benchmarks
EmoryNLP comprises 97 episodes, 897 scenes, and 12,606 utterances, where each utterance is annotated with one of the seven emotions borrowed from the six primary emotions in the Willcox (1982)’s feeling wheel, sad, mad, scared, powerful,…
58 papers · 1 benchmark
ExDark (Exclusively Dark Image Dataset)
The Exclusively Dark (ExDARK) dataset is a collection of 7,363 low-light images from very low-light environments to twilight (i.e 10 different conditions) with 12 object classes (similar to PASCAL VOC) annotated on both image class level…
58 papers · 2 benchmarks
We introduce a dataset of 147 object categories containing over 6000 images that are suitable for the few-shot counting task.
58 papers · 4 benchmarks
The Implicit Hate corpus is a dataset for hate speech detection with fine-grained labels for each message and its implication.
58 papers · 0 benchmarks
LCSTS is a large corpus of Chinese short text summarization dataset constructed from the Chinese microblogging website Sina Weibo, which is released to the public.
58 papers · 2 benchmarks
The MultiTHUMOS dataset contains dense, multilabel, frame-level action annotations for 30 hours across 400 videos in the THUMOS'14 action detection dataset.
58 papers · 3 benchmarks
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
A novel large-scale corpus of manual annotations for the SoccerNet video dataset, along with open challenges to encourage more research in soccer understanding and broadcast production.
58 papers · 6 benchmarks
WenetSpeech is a multi-domain Mandarin corpus consisting of 10,000+ hours high-quality labeled speech, 2,400+ hours weakly labelled speech, and about 10,000 hours unlabeled speech, with 22,400+ hours in total.
58 papers · 1 benchmark
XD-Violence is a large-scale audio-visual dataset for violence detection in videos.
58 papers · 2 benchmarks
Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles.
57 papers · 4 benchmarks
Includes 4000 images; 200 from each of 20 categories covering different types of scenes such as Cartoons, Art, Objects, Low resolution images, Indoor, Outdoor, Jumbled, Random, and Line drawings.
57 papers · 2 benchmarks
Composition-1K is a large-scale image matting dataset including 49300 training images and 1000 testing images.
57 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.