Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 46 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2161–2208 of 12,172
CICIDS2017 (Intrusion Detection Evaluation Dataset (CIC-IDS2017))
Intrusion Detection Evaluation Dataset (CIC-IDS2017) Intrusion Detection Systems (IDSs) and Intrusion Prevention Systems (IPSs) are the most important defense tools against the sophisticated and ever-growing network attacks.
18 papers · 2 benchmarks
We collect a new dataset of human-posed free-form natural language questions about CLEVR images.
18 papers · 1 benchmark
CMB (Comprehensive Medical Benchmark in Chinese)
CMB is a comprehensive, multi-level Medical Benchmark in Chinese.
18 papers · 0 benchmarks
CoIR (Code Information Retrieval Benchmark)
CoIR (Code Information Retrieval) benchmark, is designed to evaluate code retrieval capabilities.
18 papers · 1 benchmark
CrossMoDA is a large and multi-class benchmark for unsupervised cross-modality Domain Adaptation.
18 papers · 0 benchmarks
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
A dataset with 2,437 dialogues and 10,917 QA pairs.
18 papers · 0 benchmarks
The DroneVehicle dataset consists of a total of 56,878 images collected by the drone, half of which are RGB images, and the resting are infrared images.
18 papers · 1 benchmark
The objective in extreme multi-label classification is to learn feature architectures and classifiers that can automatically tag a data point with the most relevant subset of labels from an extremely large label set.
18 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
The Endomapper dataset is the first collection of complete endoscopy sequences acquired during regular medical practice, including slow and careful screening explorations, making secondary use of medical data.
18 papers · 0 benchmarks
This dataset was collected and prepared by the CALO Project (A Cognitive Assistant that Learns and Organizes).
18 papers · 1 benchmark
FDST (Fudan-ShanghaiTech)
The Fudan-ShanghaiTech dataset (FDST) is a dataset for video crowd counting.
18 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
18 papers · 1 benchmark
Groningen Meaning Bank is a semantic resource that anyone can edit and that integrates various semantic phenomena, including predicate-argument structure, scope, tense, thematic roles, animacy, pronouns, and rhetorical relations.
18 papers · 0 benchmarks
The heavily occluded scene text (HOST) dataset is a dataset that contains images of text with occlusions.
18 papers · 1 benchmark
ICDAR2017 is a dataset for scene text detection.
18 papers · 1 benchmark
InfoTabS comprises of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.
18 papers · 0 benchmarks
Kennedy Space Center is a dataset for the classification of wetland vegetation at the Kennedy Space Center, Florida using hyperspectral imagery.
18 papers · 1 benchmark
KinFaceW-II Dataset consists of 1000 pairs of facial images of individuals with a kin relation.
18 papers · 1 benchmark
Subjective video quality assessment (VQA) strongly depends on semantics, context, and the types of visual distortions.
18 papers · 1 benchmark
KorNLI is a Korean Natural Language Inference (NLI) dataset.
18 papers · 0 benchmarks
The LIVE Public-Domain Subjective Image Quality Database is a resource developed by the Laboratory for Image and Video Engineering at the University of Texas at Austin.
18 papers · 7 benchmarks
LIVECell (Label-free In Vitro image Examples of Cells)
The LIVECell (Label-free In Vitro image Examples of Cells) dataset is a large-scale microscopic image dataset for instance-segmentation of individual cells in 2D cell cultures.
18 papers · 1 benchmark
A 3D facial landmark dataset of around 230,000 images.
18 papers · 1 benchmark
LSHTC is a dataset for large-scale text classification.
18 papers · 0 benchmarks
An annotated image memorability dataset to date (with 60,000 labeled images from a diverse array of sources).
18 papers · 0 benchmarks
MED (Monotonicity Entailment Dataset)
MED is a new evaluation dataset that covers a wide range of monotonicity reasoning that was created by crowdsourcing and collected from linguistics publications.
18 papers · 1 benchmark
The MMD (MultiModal Dialogs) dataset is a dataset for multimodal domain-aware conversations.
18 papers · 0 benchmarks
The dataset was created for video quality assessment problem.
18 papers · 2 benchmarks
This is a dataset for video frame interpolation task.
18 papers · 1 benchmark
Proposes three types of masked face detection dataset; namely, the Correctly Masked Face Dataset (CMFD), the Incorrectly Masked Face Dataset (IMFD) and their combination for the global masked face detection (MaskedFace-Net).
18 papers · 0 benchmarks
MuirBench is a benchmark containing 11,264 images and 2,600 multiple-choice questions, providing robust evaluation on 12 multi-image understanding tasks.
18 papers · 0 benchmarks
MultiBench, a systematic and unified large-scale benchmark for multimodal learning spanning 15 datasets, 10 modalities, 20 prediction tasks, and 6 research areas.
18 papers · 0 benchmarks
Nam (A holistic approach to cross-channel image noise modeling and its application to image denoising)
A holistic approach to cross-channel image noise modeling and its application to image denoising
18 papers · 1 benchmark
ODEX is an open-domain 📖, multilingual 🌍, execution-based 🛠 natural language to code generation 💻 data benchmark.
18 papers · 0 benchmarks
ORCAS is a click-based dataset.
18 papers · 0 benchmarks
PASTIS (Panoptic Segmentation of satellite image TImes Series)
PASTIS is a benchmark dataset for panoptic and semantic segmentation of agricultural parcels from satellite image time series.
18 papers · 2 benchmarks
We propose a large-scale benchmark here, which contains a total of 6,461 mirror images with ground truth annotations.
18 papers · 1 benchmark
PU1K is nearly 8 times larger than the largest publicly available dataset collected by PU-GAN.
18 papers · 0 benchmarks
Node classification on PubMed with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
RAFT (Realworld Annotated Few-shot Tasks)
The RAFT benchmark (Realworld Annotated Few-shot Tasks) focuses on naturally occurring tasks and uses an evaluation setup that mirrors deployment.
18 papers · 1 benchmark
The RSBlur dataset provides pairs of real and synthetic blurred images with ground truth sharp images.
18 papers · 2 benchmarks
We create a benchmark dataset named ReVOS.
18 papers · 1 benchmark
The Replay-Mobile Database for face spoofing consists of 1190 video clips of photo and video attack attempts to 40 clients, under different lighting conditions.
18 papers · 0 benchmarks
RepoEval is a benchmark specifically designed for evaluating repository-level code auto-completion systems.
18 papers · 0 benchmarks
SALAD-Bench (A Hierarchical and Comprehensive Safety Benchmark for Large Language Models)
In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount.
18 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.