Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 73 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3457–3504 of 12,172
The ModelNet40 zero-shot 3D classification performance of models pretrained on ShapeNet only.
8 papers · 0 benchmarks
X-ray images in this data set have been acquired from the tuberculosis control program of the Department of Health andHuman Services of Montgomery County, MD, USA.
8 papers · 1 benchmark
MultiBooked is a dataset for supervised aspect-level sentiment analysis in Basque and Catalan, both of which are under-resourced languages.
8 papers · 0 benchmarks
MyoPS is a dataset for myocardial pathology segmentation combining three-sequence cardiac magnetic resonance (CMR) images, which was first proposed in the MyoPS challenge, in conjunction with MICCAI 2020.
8 papers · 0 benchmarks
The NCT-CRC-HE-100K dataset is a set of 100,000 non-overlapping image patches extracted from 86 H&E stained human cancer tissue slides and normal tissue from the NCT biobank (National Center for Tumor Diseases) and the UMM pathology…
8 papers · 2 benchmarks
NIND (Natural Image Noise Dataset)
An open dataset of real photographs with real noise, from identical scenes captured with varying ISO values.
8 papers · 0 benchmarks
The NLC2CMD Competition hosted at NeurIPS 2020 aimed to bring the power of natural language processing to the command line.
8 papers · 1 benchmark
NorNE is a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank.
8 papers · 0 benchmarks
NuCLS (Nucleus Classification, Localization and Segmentation)
The NuCLS dataset contains over 220,000 labeled nuclei from breast cancer images from TCGA.
8 papers · 0 benchmarks
OPERAnet is a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors.
8 papers · 0 benchmarks
OPIEC (Open Information Extraction Corpus)
OPIEC is an Open Information Extraction (OIE) corpus, constructed from the entire English Wikipedia.
8 papers · 0 benchmarks
OPRA (Online Product Reviews for Affordances)
The OPRA Dataset was introduced in Demo2Vec: Reasoning Object Affordances From Online Videos (CVPR'18) for reasoning object affordances from online demonstration videos.
8 papers · 2 benchmarks
OSCD (Onera Satellite Change Detection)
The Onera Satellite Change Detection dataset addresses the issue of detecting changes between satellite images from different dates.
8 papers · 2 benchmarks
The OSM dataset, sourced from OpenStreetMap, is composed of the rasterized semantic maps and height fields of 80 cities worldwide, spanning an area of more than 6,000 km^2.
8 papers · 1 benchmark
The Oulu-NPU face presentation attack detection database consists of 4950 real access and attack videos.
8 papers · 1 benchmark
smac+ offensive hard scenario with 20 parallel episodic buffer.
8 papers · 1 benchmark
smac+ offensive scenario with 20 parallel episodic buffer.
8 papers · 1 benchmark
OpenFWI is a collection of large-scale open-source benchmark datasets for seismic full waveform inversion (FWI).
8 papers · 0 benchmarks
OpenForensics is a large-scale dataset posing a high level of challenges that is designed with face-wise rich annotations explicitly for face forgery detection and segmentation.
8 papers · 0 benchmarks
OpenMIC-2018 is an instrument recognition dataset containing 20,000 examples of Creative Commons-licensed music available on the Free Music Archive.
8 papers · 1 benchmark
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
OpenViDial is a large-scale open-domain dialogue dataset with visual contexts.
8 papers · 0 benchmarks
Source: BARThez: a Skilled Pretrained French Sequence-to-Sequence Model OrangeSum is a single-document extreme summarization dataset with two tasks: title and abstract.
8 papers · 1 benchmark
PAD (Purpose-driven Affordance Dataset)
PAD (Purpose-driven Affordance Dataset) is a dataset for affordance detection, which refers to identifying the potential action possibilities of objects in an image, which is an important ability for robot perception and manipulation.
8 papers · 0 benchmarks
Peyma is a Persian NER dataset to train and test NER systems.
8 papers · 0 benchmarks
PHINC is a parallel corpus of the 13,738 code-mixed English-Hindi sentences and their corresponding translation in English.
8 papers · 0 benchmarks
PSC (Polish Summaries Corpus)
The Polish Summaries Corpus is a resource created to support the development and evaluation of tools for automated single-document summarization of Polish.
8 papers · 0 benchmarks
Phee is a dataset for pharmacovigilance comprising over 5000 annotated events from medical case reports and biomedical literature.
8 papers · 0 benchmarks
The PolEmo2.0 is a dataset of online consumer reviews from four domains: medicine, hotels, products, and university.
8 papers · 0 benchmarks
PKU PosterLayout, which consists of 9,974 poster-layout pairs and 905 images, i.e., non-empty canvases.
8 papers · 0 benchmarks
ProteinKG25 is a large-scale KG dataset with aligned descriptions and protein sequences respectively to GO terms and proteins entities.
8 papers · 0 benchmarks
QA2D (Question to Declarative Sentence (QA2D) Dataset)
The Question to Declarative Sentence (QA2D) Dataset contains 86k question-answer pairs and their manual transformation into declarative sentences.
8 papers · 0 benchmarks
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
Consists of multiple sentences whose clues are arranged by difficulty (from obscure to obvious) and uniquely identify a well-known entity such as those found on Wikipedia.
8 papers · 1 benchmark
The goal of REFUGE2 challenge is to evaluate and compare automated algorithms for glaucoma detection and optic disc/cup segmentation on a standard dataset of retinal fundus images.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
To perform universal event stream segmentation, we collected a large-scale RGB-Event dataset for event-centric segmentation, from current available pixel-level aligned datasets (VisEvent, COESOT), namely RGBE-SEG.
8 papers · 1 benchmark
RIMES (Reconnaissance & Indexation de données Manuscrites et de fac similÉS / Recognition & Indexing of handwritten documents & faxes)
The RIMES database (Reconnaissance et Indexation de données Manuscrites et de fac similÉS / Recognition and Indexing of handwritten documents and faxes) was created to evaluate automatic systems of recognition and indexing of handwritten…
8 papers · 0 benchmarks
The RIT-18 dataset was built for the semantic segmentation of remote sensing imagery.
8 papers · 0 benchmarks
ROBUST-MIS (Robust Medical Instrument Segmentation Challenge 2019)
The ROBUST-MIS dataset was made available to support the Robust Medical Instrument Segmentation (ROBUST-MIS) Challenge 2019, part of the Endoscopic Vision Challenge associated with MICCAI.
8 papers · 1 benchmark
RONIN (Robust Neural Inertial Navigation)
RoNIN The RoNIN dataset contains over 40 hours of IMU sensor data from 100 human subjects with 3D ground-truth trajectories under natural human movements.
8 papers · 0 benchmarks
RRS (Restoration-200k for Response Selection)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
8 papers · 1 benchmark
RUSSE (Russian Words in Context (based on RUSSE))
WiC: The Word-in-Context Dataset A reliable benchmark for the evaluation of context-sensitive word embeddings.
8 papers · 1 benchmark
Over 1.5K images selected from the public Kaggle DR Detection dataset; Five DR grades (DR0 / DR1 / DR2 / DR3 / DR4), re-labeled by a panel of 45 experienced ophthalmologists; Eight retinal lesion classes, including microaneurysm,…
8 papers · 0 benchmarks
This dataset, called RodoSol-ALPR dataset, contains 20,000 images captured by static cameras located at pay tolls owned by the Rodovia do Sol (RodoSol) concessionaire, which operates 67.5 kilometers of a highway (ES-060) in the Brazilian…
8 papers · 0 benchmarks
The SAMM Long Videos dataset consists of 147 long videos with 343 macro-expressions and 159 micro-expressions.
8 papers · 0 benchmarks
The SB10k dataset is a valuable resource for sentiment analysis in German.
8 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.