Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 91 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4321–4368 of 12,172
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
InsPLAD (Inspection Power Line Asset Dataset)
InsPLAD is a Dataset for Power Line Asset Inspection containing 10,607 high-resolution Unmanned Aerial Vehicles colour images.
5 papers · 1 benchmark
This discourse treebank includes annotated instructional texts originally assembled at the Information Technology Research Institute, University of Brighton.
5 papers · 1 benchmark
Context This is image data of Natural Scenes around the world.
5 papers · 2 benchmarks
IoT Inspector is a large dataset of labeled network traffic from smart home devices from within real-world home networks.
5 papers · 0 benchmarks
IoT devices captures - Samuel Marchal (Creator) Description This dataset represents the traffic emitted during the setup of 31 smart home IoT devices of 27 different types (4 types are represented by 2 devices each).
5 papers · 0 benchmarks
IowaRain is a dataset of rainfall events for the state of Iowa (2016-2019) acquired from the National Weather Service Next Generation Weather Radar (NEXRAD) system and processed by a quantitative precipitation estimation system.
5 papers · 0 benchmarks
ItaCoLA is a corpus for monolingual and cross-lingual acceptability judgments which contains almost 10,000 sentences with acceptability judgments.
5 papers · 1 benchmark
JRDB-Act is an extension of the JRDB dataset to create a large-scale multi-modal dataset for spatio-temporal action, social group and activity detection.
5 papers · 0 benchmarks
Dataset for lyrics alignment and transcription evaluation.
5 papers · 0 benchmarks
6.5 million anonymous ratings of jokes by users of the Jester Joke Recommender System.
5 papers · 0 benchmarks
This is the data set used for The Third International Knowledge Discovery and Data Mining Tools Competition, which was held in conjunction with KDD-99 The Fifth International Conference on Knowledge Discovery and Data Mining.
5 papers · 1 benchmark
A clickthrough prediction dataset, for more information please see the Kaggle page
5 papers · 1 benchmark
KETOD (Knowledge-Enriched Task-Oriented Dialogue)
KETOD (Knowledge-Enriched Task-Oriented Dialogue) is a dataset containing system responses designed for enriching task-oriented dialogues with chit-chat based on relevant entity knowledge.
5 papers · 0 benchmarks
KG20C (A scholarly knowledge graph benchmark dataset)
KG20C is a Knowledge Graph about high quality papers from 20 top computer science Conferences.
5 papers · 1 benchmark
Dataset Description: Toward making use of the complementary information captured by the various bioactivity types, including IC50, K(i), and K(d), Tang et al.
5 papers · 1 benchmark
The KITTI360Pose dataset encompasses a total area of 15.51 square kilometers across nine urban regions, consisting of 43,381 point cloud- text pairs.
5 papers · 1 benchmark
KiloGram is a resource for studying abstract visual reasoning in humans and machines.
5 papers · 0 benchmarks
Klexikon (Klexikon: A German Dataset for Joint Summarization and Simplification)
The dataset introduces document alignments between German Wikipedia and the children's lexicon Klexikon.
5 papers · 1 benchmark
The Kvasir-SEG dataset includes 196 polyps smaller than 10 mm classified as Paris class 1 sessile or Paris class IIa.
5 papers · 0 benchmarks
LINNAEUS is a general-purpose dictionary matching software, capable of processing multiple types of document formats in the biomedical domain (MEDLINE, PMC, BMC, OTMI, text, etc.).
5 papers · 1 benchmark
LIVE Livestream is a database for Video Quality Assessment (VQA), specifically designed for live streaming VQA research.
5 papers · 1 benchmark
LUDB (Lobachevsky University Electrocardiography Database)
Abstract Lobachevsky University Electrocardiography Database (LUDB) is an ECG signal database with marked boundaries and peaks of P, T waves and QRS complexes.
5 papers · 1 benchmark
LaRS (Lakes, Rivers and Seas Dataset)
LaRS is the largest and most diverse panoptic maritime obstacle detection dataset.
5 papers · 3 benchmarks
LabPics (LabPics Dataset for computer vision for autonomous chemistry labs and medical labs)
LabPics Chemistry Dataset Dataset for computer vision for materials segmentation and classification in chemistry labs, medical labs, and any setting where materials are handled inside containers.
5 papers · 0 benchmarks
Lesion Boundary Segmentation Dataset is a dataset for lesion segmentation from the ISIC2018 challenge.
5 papers · 0 benchmarks
LiDiRus (Linguistic Diagnostic for Russian)
LiDiRus is a diagnostic dataset that covers a large volume of linguistic phenomena, while allowing you to evaluate information systems on a simple test of textual entailment recognition.
5 papers · 1 benchmark
Libri-Adapt aims to support unsupervised domain adaptation research on speech recognition models.
5 papers · 0 benchmarks
A large-scale Indonesian summarization dataset consisting of harvested articles from Liputan6.com, an online news portal, resulting in 215,827 document-summary pairs.
5 papers · 0 benchmarks
LoLi-Phone is a large-scale low-light image and video dataset for Low-light image enhancement (LLIE).
5 papers · 0 benchmarks
We propose a novel long-context benchmark, 🐉 Loong, aligning with realistic scenarios through extended multi-document question answering (QA).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MAG-Scholar-C is constructed by Bojchevski et al.
5 papers · 1 benchmark
MATINF (Maternal and Infant Dataset)
Maternal and Infant (MATINF) Dataset is a large-scale dataset jointly labeled for classification, question answering and summarization in the domain of maternity and baby caring in Chinese.
5 papers · 0 benchmarks
MEDIC is a large social media image classification dataset for humanitarian response consisting of 71,198 images to address four different tasks in a multi-task learning setup.
5 papers · 0 benchmarks
The Messidor database has been established to facilitate studies on computer-assisted diagnoses of diabetic retinopathy.
5 papers · 0 benchmarks
This dataset is a sound dataset for malfunctioning industrial machine investigation and inspection with domain shifts due to changes in operational and environmental conditions (MIMII DUE).
5 papers · 0 benchmarks
A new dataset on the baseball domain.
5 papers · 1 benchmark
The MLB-YouTube dataset is a new, large-scale dataset consisting of 20 baseball games from the 2017 MLB post-season available on YouTube with over 42 hours of video footage.
5 papers · 0 benchmarks
MLQE (MultiLingual Quality Estimation)
The MLQE dataset is a dataset for sentence-level Machine Translation Quality Estimation.
5 papers · 0 benchmarks
- A large scale Chinese multi-modal dialogue corpus (120.84K dialogues and 198.82 K images).
5 papers · 0 benchmarks
MMFlood is remote sensing dataset derived from Sentinel-1 (VV-VH), MapZen (DEM) and OpenStreetMap (Hydrography).
5 papers · 1 benchmark
Contains 40K human judgement scores on model outputs from 6 diverse question answering datasets and an additional set of minimal pairs for evaluation.
5 papers · 0 benchmarks
A dataset which provides detailed annotations for activity recognition.
5 papers · 1 benchmark
MSAW (Multi-Sensor All Weather Mapping)
Multi-Sensor All Weather Mapping (MSAW) is a dataset and challenge, which features two collection modalities (both SAR and optical).
5 papers · 1 benchmark
We construct the first large-scale mirror dataset, named MSD.
5 papers · 1 benchmark
MSMT17-C is an evaluation set that consists of algorithmically generated corruptions applied to the MSMT17 test-set.
5 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.