Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 101 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4801–4848 of 12,172
The IS-A dataset is a dataset of relations extracted from a medical ontology.
4 papers · 0 benchmarks
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
Public ECG dataset of continuous raw signals for representation learning containing 11 thousand patients and 2 billion labelled beats.
4 papers · 0 benchmarks
A new language-guided image editing dataset that contains a large number of real image pairs with corresponding editing instructions.
4 papers · 0 benchmarks
The NINCO (No ImageNet Class Objects) dataset is introduced in the ICML 2023 paper In or Out?
4 papers · 1 benchmark
The ImageNet-50 dataset split as introduced in TEMI.
4 papers · 1 benchmark
ImageNet-D contains 4835 test images featuring diverse backgrounds (3,764), textures (498), and materials (573).
4 papers · 0 benchmarks
ImgEdit is a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.
4 papers · 0 benchmarks
Imp1k is a new dataset of designs annotated with importance information.
4 papers · 0 benchmarks
Incidents1M is a large-scale multi-label dataset for incident detection which contains 977,088 images, with 43 incident and 49 place categories.
4 papers · 0 benchmarks
IndoNLG is a benchmark to measure natural language generation (NLG) progress in three low-resource—yet widely spoken—languages of Indonesia: Indonesian, Javanese, and Sundanese.
4 papers · 0 benchmarks
IndoNLI is the first human-elicited NLI dataset for Indonesian consisting of nearly 18K sentence pairs annotated by crowd workers and experts.
4 papers · 0 benchmarks
InferWiki is a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns.
4 papers · 0 benchmarks
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality.
4 papers · 0 benchmarks
This is a large-scale dataset of tweets associated to thousands of news articles published on Italian disinformation websites in the context of 2019 European elections.
4 papers · 0 benchmarks
JaQuAD (Japanese Question Answering Dataset) is a question answering dataset in Japanese that consists of 39,696 extractive question-answer pairs on Japanese Wikipedia articles.
4 papers · 1 benchmark
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior.
4 papers · 1 benchmark
JobStack is a new corpus for de-identification of personal data in job vacancies on Stackoverflow.
4 papers · 0 benchmarks
K-hairstyle is a novel large-scale Korean hairstyle dataset with 256,679 high-resolution images.
4 papers · 0 benchmarks
KAIST-Lane (K-Lane) is the world’s first and the largest public urban road and highway lane dataset for Lidar.
4 papers · 1 benchmark
KRAUTS (Korpus of newspapeR Articles with Underlinded Temporal expressionS)
KRAUTS (Korpus of newspapeR Articles with Underlinded Temporal expressionS) is a German temporally annotated news corpus accompanied with TimeML annotation guidelines for German.
4 papers · 1 benchmark
The KTH-TIPS (Textures under varying Illumination, Pose and Scale) image database was created to extend the CUReT database in two directions, by providing variations in scale as well as pose and illumination, and by imaging other samples…
4 papers · 1 benchmark
Human Activity Recognition (HAR) refers to the capacity of machines to perceive human actions.
4 papers · 0 benchmarks
KazakhTTS is an open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide.
4 papers · 0 benchmarks
From my knowledge, the dataset used in the project is the largest crack segmentation dataset so far.
4 papers · 2 benchmarks
Kitchen Scenes is a multi-view RGB-D dataset of nine kitchen scenes, each containing several objects in realistic cluttered environments including a subset of objects from the BigBird dataset.
4 papers · 0 benchmarks
Kitsune Network Attack Dataset This is a collection of nine network attack datasets captured from a either an IP-based commercial surveillance system or a network full of IoT devices.
4 papers · 0 benchmarks
This is the dataset for knowledge editing.
4 papers · 0 benchmarks
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks.
4 papers · 0 benchmarks
Kosp2e (read as kospi'), is a corpus that allows Korean speech to be translated into English text in an end-to-end manner
4 papers · 0 benchmarks
Kuzushiji-Kanji is an imbalanced dataset of total 3832 Kanji characters (64x64 grayscale, 140,426 images), ranging from 1,766 examples to only a single example per class.
4 papers · 0 benchmarks
LAM(line-level) (The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition)
Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing.
4 papers · 1 benchmark
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
Comparative evaluation of virtual screening methods requires a rigorous benchmarking procedure on diverse, realistic, and unbiased data sets.
4 papers · 1 benchmark
Comparative evaluation of virtual screening methods requires a rigorous benchmarking procedure on diverse, realistic, and unbiased data sets.
4 papers · 1 benchmark
Comparative evaluation of virtual screening methods requires a rigorous benchmarking procedure on diverse, realistic, and unbiased data sets.
4 papers · 1 benchmark
LLM-Seg40K dataset contains 14K images in total.
4 papers · 0 benchmarks
Laptop-ACOS is a brand new Laptop dataset collected from the Amazon platform in the years 2017 and 2018 (covering ten types of laptops under six brands such as ASUS, Acer, Samsung, Lenovo, MBP, MSI, and so on).
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LeetCode-Hard is a benchmark dataset for code generation, consisting of 40 challenging LeetCode "hard-level" questions across 19 programming languages.
4 papers · 0 benchmarks
LegalNERo (Romanian Named Entity Recognition in the Legal domain)
LegalNERo is a manually annotated corpus for named entity recognition in the Romanian legal domain.
4 papers · 1 benchmark
LemgoRL is an open-source benchmark tool for traffic signal control designed to train reinforcement learning agents in a highly realistic simulation scenario with the aim to reduce Sim2Real gap.
4 papers · 0 benchmarks
This is a 4D light-field dataset of materials.
4 papers · 0 benchmarks
LineCap is a dataset of line charts scraped from scientific papers each accompanied with crowd-sourced captions describing the trends of individual lines in the figure and the figure as a whole.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.