Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 60 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2833–2880 of 12,172
WildScenes is a bi-modal benchmark dataset consisting of multiple large-scale, sequential traversals in natural environments, including semantic annotations in high-resolution 2D images and dense 3D LiDAR point clouds, and accurate 6-DoF…
12 papers · 2 benchmarks
YT-UGC is a large scale UGC (User Generated Content) dataset (1,500 20 sec video clips) sampled from millions of YouTube videos.
12 papers · 0 benchmarks
Constructed from over one million fashion images with a label space that includes 8 groups of 228 fine-grained attributes in total.
12 papers · 0 benchmarks
A prebuilt dataset for OpenAI's task for image-2-latex system.
12 papers · 1 benchmark
This dataset has 20 classes and each class has about 1000 documents.
11 papers · 1 benchmark
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
ACOS (Aspect Category Opinion Sentiment)
Most of the aspect based sentiment analysis research aims at identifying the sentiment polarities toward some explicit aspect terms while ignores implicit aspects in text.
11 papers · 1 benchmark
ACS PUMS stands for American Community Survey (ACS) Public Use Microdata Sample (PUMS) and has been used to construct several tabular datasets for studying fairness in machine learning: - ACSIncome: to predict whether an individual’s…
11 papers · 0 benchmarks
AKCES-GEC is a new dataset on grammatical error correction for Czech.
11 papers · 0 benchmarks
AMPS (Auxiliary Mathematics Problems and Solutions)
AMPS contains over 100,000 problems pulled from Khan Academy and approximately 5 million problems generated from manually designed Mathematica scripts.
11 papers · 0 benchmarks
ARCH is a computational pathology (CP) multiple instance captioning dataset to facilitate dense supervision of CP tasks.
11 papers · 0 benchmarks
ASAP-AES (Automated Student Assessment Prize)
There are eight essay sets.
11 papers · 1 benchmark
A set of 19 ASC datasets (reviews of 19 products) producing a sequence of 19 tasks.
11 papers · 1 benchmark
ASQP (Aspect Sentiment Quad Prediction)
Aspect-based sentiment analysis (ABSA) typically focuses on extracting aspects and predicting their sentiments on individual sentences such as customer reviews.
11 papers · 1 benchmark
AeBAD (Aero-engine Blade Anomaly Detection Dataset)
Unlike previous datasets that focus on detecting the diversity of defect categories (like MVTec AD and VisA), AeBAD is centered on the diversity of domains within the same data category.
11 papers · 3 benchmarks
The AmbigNQ dataset is a valuable resource for exploring ambiguity in open-domain question answering.
11 papers · 0 benchmarks
The Argoverse 2 Sensor Dataset is a collection of 1,000 scenarios with 3D object tracking annotations.
11 papers · 0 benchmarks
(1) provide financial news for each specific stock.
11 papers · 2 benchmarks
BCNB (Early Breast Cancer Core-Needle Biopsy WSI)
Breast cancer (BC) has become the greatest threat to women’s health worldwide.
11 papers · 0 benchmarks
BCN20000 is a dataset composed of 19,424 dermoscopic images of skin lesions captured from 2010 to 2016 in the facilities of the Hospital Clínic in Barcelona.
11 papers · 0 benchmarks
BIKED is a dataset comprised of 4500 individually designed bicycle models sourced from hundreds of designers.
11 papers · 0 benchmarks
BL30K is a synthetic dataset rendered using Blender with ShapeNet's data.
11 papers · 0 benchmarks
BSTC (Baidu Speech Translation Corpus)
BSTC (Baidu Speech Translation Corpus) is a large-scale dataset for automatic simultaneous interpretation.
11 papers · 0 benchmarks
This data set includes hourly air pollutants data from 12 nationally-controlled air-quality monitoring sites.
11 papers · 1 benchmark
BinaryCorp is built for binary similarity detection based on the ArchLinux official repositories and Arch User Repository.
11 papers · 0 benchmarks
BreakHis (Breast Cancer Histopathological Database)
The Breast Cancer Histopathological Image Classification (BreakHis) is composed of 9,109 microscopic images of breast tumor tissue collected from 82 patients using different magnifying factors (40X, 100X, 200X, and 400X).
11 papers · 5 benchmarks
BuildingNet is a large-scale dataset of 3D building models whose exteriors are consistently labeled.
11 papers · 0 benchmarks
The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English.
11 papers · 7 benchmarks
CATS (Color and Thermal Stereo Benchmark)
A dataset consisting of stereo thermal, stereo color, and cross-modality image pairs with high accuracy ground truth (< 2mm) generated from a LiDAR.
11 papers · 2 benchmarks
CMRC 2017 (Chinese Machine Reading Comprehension 2017)
Contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set.
11 papers · 0 benchmarks
COST (COCO Segmentation Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 0 benchmarks
Casia V1 is a dataset for forgery classification.
11 papers · 2 benchmarks
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
Cityscapes-Seq is a standard dataset for semantic urban scene understanding, featuring real-world videos from 50 cities in Germany and neighboring countries.
11 papers · 0 benchmarks
ClariQ is an extension of the Qulac dataset with additional new topics, questions, and answers in the training set.
11 papers · 0 benchmarks
Five classic grayscale images commonly used for image quality assessment tasks.
11 papers · 4 benchmarks
Complementary Commonsense (Com2Sense) is a dataset for benchmarking commonsense reasoning ability of NLP models.
11 papers · 0 benchmarks
The ConceptARC dataset is a benchmark for evaluating understanding and generalization in the Abstraction and Reasoning Corpus (ARC) domain.
11 papers · 0 benchmarks
Crello dataset consists of design templates obtained from online design service, crello.com.
11 papers · 0 benchmarks
Darpa is a dataset consisting of communications between source IPs and destination IPs.
11 papers · 0 benchmarks
DCASE 2013 is a dataset for sound event detection.
11 papers · 0 benchmarks
The DEAP dataset consists of two parts: - The ratings from an online self-assessment where 120 one-minute extracts of music videos were each rated by 14-16 volunteers based on arousal, valence and dominance.
11 papers · 1 benchmark
Forgery Diversity: DF40 comprises 40 distinct deepfake techniques (both representive and SOTA methods are included), facilitating the detection of nowadays' SOTA deepfakes and AIGCs.
11 papers · 0 benchmarks
Depth in the Wild is a dataset for single-image depth perception in the wild, i.e., recovering depth from a single image taken in unconstrained settings.
11 papers · 0 benchmarks
A new face annotation dataset with balanced distribution between genders and ethnic origins.
11 papers · 2 benchmarks
DocILE is a large dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition.
11 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.