Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 115 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5473–5520 of 12,172
A large scale dataset of scientific records from PubMed for scientific keyphrase generation.
3 papers · 0 benchmarks
The Kenyan Food Type Dataset (KenyanFood13) is an image classification dataset for Kenyan food.
3 papers · 0 benchmarks
Collected by cleaning data from knowledge-intensive websites like Wikipedia and science and technology reports, and processing it using reverse engineering techniques.
3 papers · 0 benchmarks
Konzil dataset was created by specialists of the University of Greifswald.
3 papers · 0 benchmarks
Presents 9.4K manually labeled entertainment news comments for identifying Korean toxic speech, collected from a widely used online news platform in Korea.
3 papers · 0 benchmarks
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
The dataset contains a Video capsule endoscopy dataset for polyp segmentation.
3 papers · 1 benchmark
LAGENDA (Layer Age and Gender Dataset)
The LAGENDA dataset is a large-scale dataset with age and gender annotations for face and body bounding boxes.
3 papers · 4 benchmarks
LAION-COCO is the world’s largest dataset of 600M generated high-quality captions for publicly available web-images.
3 papers · 1 benchmark
A subset of the LAION 5B samples with English captions, obtained using LAION-AestheticsPredictor V2 625K image-text pairs with predicted aesthetics scores of 6.5 or higher available at…
3 papers · 0 benchmarks
LARC (Language-annotated Abstraction and Reasoning)
LARC is a dataset built from ARC (Abstraction and Reasoning Corpus).
3 papers · 0 benchmarks
A Large Dataset for Remote Sensing Image Change Captioning.
3 papers · 0 benchmarks
LIDDI (LInked Drug-Drug Interactions)
LInked Drug-Drug Interactions (LIDDI) is a public nanopublication-based RDF dataset with trusty URIs that encompasses some of the most cited prediction methods and sources to provide researchers a resource for leveraging the work of others…
3 papers · 0 benchmarks
Comparative evaluation of virtual screening methods requires a rigorous benchmarking procedure on diverse, realistic, and unbiased data sets.
3 papers · 1 benchmark
The LITIS-Rouen dataset is a dataset for audio scenes.
3 papers · 0 benchmarks
LKS (Liver Kidney Stomach)
LKS is a dataset of 684 Liver-Kidney-Stomach immunofluorescence whole slide images (WSIs) used in the investigation of autoimmune liver disease.
3 papers · 0 benchmarks
The dataset was proposed in LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
A large-scale logo image database for logo detection and brand recognition from real-world product images.
3 papers · 0 benchmarks
LOOK is a large-scale dataset for eye contact detection in the wild, which focuses on diverse and unconstrained scenarios for real-world generalization.
3 papers · 0 benchmarks
LSA16 (Lengua de Señas Argentina - 16 Handshapes classes)
This database contains images of 16 handshapes of the Argentinian Sign Language (LSA), each performed 5 times by 10 different subjects, for a total of 800 images.
3 papers · 1 benchmark
Large Age-Gap (LAG) is a dataset for face verification, The dataset contains 3,828 images of 1,010 celebrities.
3 papers · 0 benchmarks
We introduce here our Large Time Lags Location (LTLL) dataset containing pictures of 25 locations captured over a range of more than 150 years.
3 papers · 1 benchmark
Got "pubchemsmilescanonical.zip" from https://ibm.ent.box.com/v/MoLFormer-data
3 papers · 0 benchmarks
LayoutBench is a diagnostic benchmark that examines 4 spatial control skills (number, position, size, shape), where each skill consists of 2 OOD layout splits, i.e., in total 8 tasks = 4 skills x 2 splits.
3 papers · 1 benchmark
LeNER-Br is a dataset for named entity recognition (NER) in Brazilian Legal Text.
3 papers · 2 benchmarks
LibriVoxDeEn is a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks.
3 papers · 0 benchmarks
Long-term visual localization provides a benchmark datasets aimed at evaluating 6 DoF pose estimation accuracy over large appearance variations caused by changes in seasonal (summer, winter, spring, etc.) and illumination (dawn, day,…
3 papers · 0 benchmarks
The dataset contains the annotations of characters' visual appearances, in the form of tracks of face bounding boxes, and the associations with characters' textual mentions, when available.
3 papers · 1 benchmark
MCVQA (Multilingual and Code-mixed Visual Question Answering)
The MCVQA dataset consists of 248, 349 training questions and 121, 512 validation questions for real images in Hindi and Code-mixed.
3 papers · 0 benchmarks
MDIA is a large-scale multilingual benchmark for dialogue generation.
3 papers · 0 benchmarks
MDID (Multimodal Document Intent Dataset)
The Multimodal Document Intent Dataset (MDID) is a dataset for computing author intent from multimodal data from Instagram.
3 papers · 0 benchmarks
MED-NODE (Dermatology database used in MED-NODE)
"Our dataset consists of 70 melanoma and 100 naevus images from the digital image archive of the Department of Dermatology of the University Medical Center Groningen (UMCG) used for the development and testing of the MED-NODE system for…
3 papers · 0 benchmarks
MERL Shopping is a dataset for training and testing action detection algorithms.
3 papers · 0 benchmarks
The original dataset from Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting contains traffic readings collected from 207 loop detectors on highways in Los Angeles County, aggregated in 5 minutes intervals…
3 papers · 1 benchmark
A multi-sensor, multi-modal dataset, implemented to benchmark Human Activity Recognition(HAR) and Multi-modal Fusion algorithms.
3 papers · 0 benchmarks
MFAQ is a multilingual FAQ dataset publicly available.
3 papers · 0 benchmarks
The MG-ShopDial dataset contains English conversations that mix different conversational goals, including search, recommendation, and question answering in the domain of e-commerce.
3 papers · 0 benchmarks
MIMIC-IV ICD-10 contains 122,279 discharge summaries—free-text medical documents—annotated with ICD-10 diagnosis and procedure codes.
3 papers · 1 benchmark
MIMIC-IV-ECG (MIMIC-IV-ECG: Diagnostic Electrocardiogram Matched Subset)
The MIMIC-IV-ECG module contains approximately 800,000 diagnostic electrocardiograms across nearly 160,000 unique patients.
3 papers · 0 benchmarks
The MIMIC-IV-ICD10-full dataset, including occurring labels.
3 papers · 1 benchmark
The MIMIC-IV-ICD9 dataset, including all occurring labels.
3 papers · 1 benchmark
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
MISP2021 (Multimodal Information Based Speech Processing 2021)
The MISP2021 challenge dataset is a collection of audio-visual conversational data recorded in a home TV scenario using distant multi-microphones.
3 papers · 0 benchmarks
This database includes 25 long-term ECG recordings of human subjects with atrial fibrillation (mostly paroxysmal).
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.