Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 71 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3361–3408 of 12,172
The BioDiv dataset includes manually labeled tables for CTA and CEA from the biodiversity domain.
8 papers · 2 benchmarks
BosphorusSign22k is a benchmark dataset for vision-based user-independent isolated Sign Language Recognition (SLR).
8 papers · 0 benchmarks
Building3D is an urban-scale dataset consisting of more than 160 thousands buildings along with corresponding point clouds, mesh and wireframe models, covering 16 cities in Estonia about 998 Km2.
8 papers · 0 benchmarks
CAD (Contextual Abuse Dataset)
Dataset of primarily English Reddit entries which addresses several limitations of prior work.
8 papers · 1 benchmark
CAMO++ is a dataset for camouflaged object segmentation.
8 papers · 0 benchmarks
CHEAT (CHatGPT-writtEn AbsTracts)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
8 papers · 0 benchmarks
CHIP-CDN (Clinical Diagnosis Normalization Dataset)
CHIP Clinical Diagnosis Normalization, a dataset that aims to standardize the terms from the final diagnoses of Chinese electronic medical records, is used for the CHIP-CDN task.
8 papers · 0 benchmarks
CHIP-STS (Semantic Textual Similarity Dataset)
CHIP Semantic Textual Similarity, a dataset for sentence similarity in the non-i.i.d.
8 papers · 1 benchmark
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
This dataset of motions is free for all uses.
8 papers · 0 benchmarks
Dataset [46 M] and readme: 42,306 movie plot summaries extracted from Wikipedia + aligned metadata extracted from Freebase, including: Movie box office revenue, genre, release date, runtime, and language Character names and aligned…
8 papers · 0 benchmarks
The COCO-MIG benchmark (Common Objects in Context Multi-Instance Generation) is a benchmark used to evaluate the generation capability of generators on text containing multiple attributes of multi-instance objects.
8 papers · 1 benchmark
COUCH is a large human-chair interaction dataset with clean annotations.
8 papers · 0 benchmarks
Under a close collaboration with an expert radiologist team of the Hospital Universitario San Cecilio, the COVIDGR-1.0 dataset of patients' anonymized X-ray images has been built.
8 papers · 2 benchmarks
We present CS-Campus3D, the first 3D aerial-ground cross-source dataset consisting of point cloud data from both aerial and ground LiDAR scans.
8 papers · 1 benchmark
CSD (Collaborative SLAM Dataset)
Comprises 4 different subsets - Flat, House, Priory and Lab - each containing a number of different sequences that can be successfully relocalised against each other.
8 papers · 1 benchmark
Collects shadow images for multiple scenarios and compiled a new dataset of 10,500 shadow images, each with labeled ground-truth mask, for supporting shadow detection in the complex world.
8 papers · 1 benchmark
The Cityscapes Panoptic Parts dataset introduces part-aware panoptic segmentation annotations for the Cityscapes dataset.
8 papers · 1 benchmark
Cityscapes-DVPS is derived from Cityscapes-VPS by adding re-computed depth maps from Cityscapes dataset.
8 papers · 0 benchmarks
File descriptions train - Training set.
8 papers · 0 benchmarks
CoNLL-2000 is a dataset for dividing text into syntactically related non-overlapping groups of words, so-called text chunking.
8 papers · 0 benchmarks
The task builds on the CoNLL-2008 task and extends it to multiple languages.
8 papers · 2 benchmarks
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 0 benchmarks
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 2 benchmarks
ColonINST is a large-scale instruction tuning dataset designed for multimodal analysis in colonoscopy.
8 papers · 2 benchmarks
A large commercial Ads Dataset includes 480K labeled query-ad pairwise data with structured information of image, title, seller, description, and so on.
8 papers · 1 benchmark
Cops-Ref is a dataset for visual reasoning in context of referring expression comprehension with two main features.
8 papers · 0 benchmarks
DISC21 is a benchmark for large-scale image similarity detection.
8 papers · 1 benchmark
DOTmark (Discrete Optimal Transport Benchmark)
DOTmark is a benchmark for discrete optimal transport, which is designed to serve as a neutral collection of problems, where discrete optimal transport methods can be tested, compared to one another, and brought to their limits on…
8 papers · 0 benchmarks
The DSSE-200 is a complex document layout dataset including various dataset styles.
8 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
DUO (Detecting Underwater Objects)
DUO is a dataset for Underwater object detection for robot picking.
8 papers · 1 benchmark
DarkTrack2021 is a challenging nighttime UAV tracking benchmark, which contains 110 challenging sequences with over 100 K frames in total.
8 papers · 0 benchmarks
DeepLoc is a large-scale urban outdoor localization dataset.
8 papers · 0 benchmarks
smac+ defense armored scenario with parallel episodic buffer
8 papers · 1 benchmark
smac+ defense outnumbered scenario with parallel episodic buffer
8 papers · 1 benchmark
A large scale of retina image dataset.
8 papers · 0 benchmarks
DocNLI is a large-scale dataset for document-level NLI.
8 papers · 0 benchmarks
Duke Breast Cancer MRI (Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations)
Breast MRI scans of 922 cancer patients from Duke University, with tumor bounding box annotations, clinical, imaging, and many other features, and more.
8 papers · 0 benchmarks
ERA (Event Recognition in Aerial videos)
Consists of 2,864 videos each with a label from 25 different classes corresponding to an event unfolding 5 seconds.
8 papers · 0 benchmarks
ETH SfM (ETH Structure-from-Motion)
The ETH SfM (structure-from-motion) dataset is a dataset for 3D Reconstruction.
8 papers · 0 benchmarks
The Electricity Transformer Temperature (ETT) is a crucial indicator in the electric power long-term deployment.
8 papers · 1 benchmark
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
Echocardiography, or cardiac ultrasound, is the most widely used and readily available imaging modality to assess cardiac function and structure.
8 papers · 0 benchmarks
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
Europarl-ASR (EN) is a 1300-hour English-language speech and text corpus of parliamentary debates for (streaming) Automatic Speech Recognition training and benchmarking, speech data filtering and speech data verbatimization, based on…
8 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.