Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 41 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1921–1968 of 12,172
CARRADA is a dataset of synchronized camera and radar recordings with range-angle-Doppler annotations.
22 papers · 0 benchmarks
COWC (Cars Overhead With Context)
The Cars Overhead With Context (COWC) data set is a large set of annotated cars from overhead.
22 papers · 0 benchmarks
A collection of single speaker speech datasets for ten languages.
22 papers · 0 benchmarks
CUFSF (CUHK Face Sketch FERET Database)
The CUHK Face Sketch FERET (CUFSF) is a dataset for research on face sketch synthesis and face sketch recognition.
22 papers · 1 benchmark
CxC (Crisscrossed Captions)
Crisscrossed Captions (CxC) contains 247,315 human-labeled annotations including positive and negative associations between image pairs, caption pairs and image-caption pairs.
22 papers · 1 benchmark
The Django dataset is a dataset for code generation comprising of 16000 training, 1000 development and 1805 test annotations.
22 papers · 1 benchmark
Do-Not-Answer is a dataset to evaluate safeguards in large language models, and deploy safer open-source LLMs at a low cost.
22 papers · 0 benchmarks
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
22 papers · 0 benchmarks
ETTh1 (96) (ETT (Electricity Transformer Temperature))
The Electricity Transformer Temperature (ETT) is a crucial indicator in the electric power long-term deployment.
22 papers · 2 benchmarks
Includes accurate pixel-wise motion masks, egomotion and ground truth depth.
22 papers · 0 benchmarks
A new benchmark dataset for cross-lingual and multilingual question answering for high school examinations.
22 papers · 0 benchmarks
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
The Fusion 360 Gallery Dataset contains rich 2D and 3D geometry data derived from parametric CAD models.
22 papers · 1 benchmark
Grounded SCAN poses a simple task, where an agent must execute action sequences based on a synthetic language instruction.
22 papers · 0 benchmarks
Gibson is an opensource perceptual and physics simulator to explore active and real-world perception.
22 papers · 0 benchmarks
This dataset contains card descriptions of the card game Hearthstone and the code that implements them.
22 papers · 0 benchmarks
HoME (Household Multimodal Environment)
HoME (Household Multimodal Environment) is a multimodal environment for artificial agents to learn from vision, audio, semantics, physics, and interaction with objects and other agents, all within a realistic context.
22 papers · 0 benchmarks
IMDb-Face is large-scale noise-controlled dataset for face recognition research.
22 papers · 0 benchmarks
IP102 contains more than 75,000 images belonging to 102 categories, which exhibit a natural long-tailed distribution.
22 papers · 0 benchmarks
(JHU-CROWD) a crowd counting dataset that contains 4,250 images with 1.11 million annotations.
22 papers · 0 benchmarks
KdConv (Knowledge-driven Conversation)
KdConv is a Chinese multi-domain Knowledge-driven Conversation dataset, grounding the topics in multi-turn conversations to knowledge graphs.
22 papers · 0 benchmarks
100 tasks from LIBERO-100 suite.
22 papers · 1 benchmark
The LOCATA dataset is a dataset for acoustic source localization.
22 papers · 0 benchmarks
Retrospectively collected medical data has the opportunity to improve patient care through knowledge discovery and algorithm development.
22 papers · 0 benchmarks
MINOS is a simulator designed to support the development of multisensory models for goal-directed navigation in complex indoor environments.
22 papers · 0 benchmarks
MOROCO (MOldavian and ROmanian Dialectal COrpus)
The MOldavian and ROmanian Dialectal COrpus (MOROCO) is a corpus that contains 33,564 samples of text (with over 10 million tokens) collected from the news domain.
22 papers · 0 benchmarks
The ManySStuBs4J corpus is a collection of simple fixes to Java bugs, designed for evaluating program repair techniques.
22 papers · 0 benchmarks
Market-1501-C is an evaluation set that consists of algorithmically generated corruptions applied to the Market-1501 test-set.
22 papers · 1 benchmark
MosMedData contains anonymised human lung computed tomography (CT) scans with COVID-19 related findings, as well as without such findings.
22 papers · 1 benchmark
Motion-X is a large-scale 3D expressive whole-body motion dataset, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes, meanwhile providing corresponding semantic…
22 papers · 1 benchmark
PIT (Paraphrase and Semantic Similarity in Twitter)
Paraphrase and Semantic Similarity in Twitter (PIT) presents a constructed Twitter Paraphrase Corpus that contains 18,762 sentence pairs.
22 papers · 1 benchmark
The ECGs in this collection were obtained using a non-commercial, PTB prototype recorder with the following specifications: 16 input channels, (14 for ECGs, 1 for respiration, 1 for line voltage) Input voltage: ±16 mV, compensated offset…
22 papers · 4 benchmarks
The PhotoShape dataset consists of photorealistic, relightable, 3D shapes produced by the work proposed in the work of Park et al.
22 papers · 1 benchmark
ReDWeb (Relative Depth from Web)
The ReDWeb dataset consists of 3600 RGB-RD image pairs collected from the Web.
22 papers · 0 benchmarks
The Sku110k dataset provides 11,762 images with more than 1.7 million annotated bounding boxes captured in densely packed scenarios, including 8,233 images for training, 588 images for validation, and 2,941 images for testing.
22 papers · 1 benchmark
SLUE (Spoken Language Understanding Evaluation)
Spoken Language Understanding Evaluation (SLUE) is a suite of benchmark tasks for spoken language understanding evaluation.
22 papers · 3 benchmarks
Satlas is a remote sensing dataset and benchmark that is large in both breadth, featuring all of the aforementioned applications and more, as well as scale, comprising 290M labels under 137 categories and 7 label modalities.
22 papers · 0 benchmarks
A simulation-based dataset featuring 20,000 stack configurations composed of a variety of elementary geometric primitives richly annotated regarding semantics and structural stability.
22 papers · 2 benchmarks
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from snopes.com.
22 papers · 1 benchmark
TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA.
22 papers · 1 benchmark
This dataset is aimed to study the existing reading comprehension models' capability to perform temporal reasoning, and see whether they are sensitive to the temporal description in the given question.
22 papers · 0 benchmarks
TopiOCQA (pronounced Tapioca) is an open-domain conversational dataset with topic switches on Wikipedia.
22 papers · 0 benchmarks
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain.
22 papers · 3 benchmarks
WIDER (Web Image Dataset for Event Recognition)
WIDER is a dataset for complex event recognition from static images.
22 papers · 1 benchmark
WISE, the first benchmark specifically designed for World Knowledge-Informed Semantic Evaluation.
22 papers · 2 benchmarks
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.