Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 189 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9025–9072 of 12,172
Studying how human drivers react differently when following autonomous vehicles (AV) vs.
1 paper · 0 benchmarks
It is composed of around 770k of color 256x256 RGB images extracted from the European Union Intellectual Property Office (EUIPO) open registry.
1 paper · 1 benchmark
This dataset contains two types of audio recordings.
1 paper · 0 benchmarks
A corpus of 21,570 newspaper headlines written in European Spanish annotated with emergent anglicisms.
1 paper · 0 benchmarks
LeQua2022 (Learning to Quantify Dataset 2024)
This is the dataset used in the 1st data challenge on Learning to Quantify.
1 paper · 0 benchmarks
LeT-Mi (Levantine Twitter dataset for Misogynistic language)
Levantine Twitter dataset for Misogynistic language (LeT-Mi) is an Arabic Levantine Twitter dataset for misogynistic language to be the first benchmark dataset for Arabic misogyny.
1 paper · 0 benchmarks
Dataset containing 9372 RGB images of weeds with the number of leaves counted.
1 paper · 0 benchmarks
LeafNet (LeafNet: A large-scale dataset for training image-text models in leaf disease identification)
The PlantVillage dataset, with over 54,000 images spanning 14 plant species and 26 disease types, has been widely used for leaf disease classification.
1 paper · 1 benchmark
Contains 80 questions of LeetCode weekly and bi-weekly contests released after March 2024.
1 paper · 0 benchmarks
Dataset Summary New dataset introduced in Parameter-Efficient Legal Domain Adaptation (Li et al., 2022) from the Legal Advice Reddit community (known as "/r/legaldvice"), sourcing the Reddit posts from the Pushshift Reddit dataset.
1 paper · 0 benchmarks
Recognizing events and their coreferential men- tions in a document is essential for understand- ing semantic meanings of text.
1 paper · 0 benchmarks
This dataset includes sharp-blur pairs of Leishmania image, which is a protozoan parasite microscopy image dataset of Leishmania, obtained from the preserved slides stained with Giemsa.
1 paper · 0 benchmarks
The Lens Flare dataset is an internal dataset for Flare Spot detection used in the paper "Automatic Flare Spot Artifact Detection and Removal in Photographs" by Patricia Vitoria and Coloma Ballester.
1 paper · 0 benchmarks
The Lenta Short Sentences dataset is a text dataset for language modelling for the Russian language.
1 paper · 0 benchmarks
LfGP Data (Learning from Guided Play Expert Data and Trained Models)
The expert data and trained models used for our Learning from Guided Play paper.
1 paper · 0 benchmarks
LiDAR-CS is a dataset for 3D object detection in real traffic.
1 paper · 0 benchmarks
LiPC (LiDAR Point Cloud Clustering Benchmark Suite)
LiPC (LiDAR Point Cloud Clustering Benchmark Suite) is a benchmark suite for point cloud clustering algorithms based on open-source software and open datasets.
1 paper · 0 benchmarks
LibriS2S is a Speech to Speech Translation (S2ST) dataset build further upon existing resources.
1 paper · 0 benchmarks
These are processed versions of the monthly lichess database release.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LigBoundConf is a set of high-quality drug-like bound ligand structures in the Protein Data Bank (PDB).
1 paper · 0 benchmarks
We created this robust and custom light field dataset in order to assist light field researchers in using SOTA machine learning algorithms for a variety of light field tasks such as depth estimation, synthetic aperture imaging, and more.
1 paper · 0 benchmarks
LimeSoDa (Precision Liming Soil Datasets)
Precision Liming Soil Datasets (LimeSoDa) is a collection of 31 datasets from a field- and farm-scale soil mapping context.
1 paper · 0 benchmarks
The Lincolnbeet dataset is an object detection dataset designed to encourage research in the identification of items in environments with high levels of occlusion, and in the development of better approaches to evaluate object detection…
1 paper · 0 benchmarks
The Linguistic Benchmark (JSON), consisting of 30 questions was developed to be easy for human adults to answer but challenging for LLMs.
1 paper · 0 benchmarks
The Linked Wikitext-2 language modeling dataset contains over 2 million tokens from Wikipedia articles, along with annotations linking mentions to their corresponding entities and relations in Wikidata.
1 paper · 0 benchmarks
The LinkedResults dataset contains around 1,600 results capturing performance of machine learning models from tables of 239 papers.
1 paper · 0 benchmarks
This is a dataset of 3 English books which do not contain the letter "e" in them.
1 paper · 1 benchmark
CSV file with a list of all examined OWL reasoners.
1 paper · 0 benchmarks
A listwise multi-response dataset for human preferences alignment.
1 paper · 0 benchmarks
An open-source online generative dictionary that takes a word and context containing the word as input and automatically generates a definition as output.
1 paper · 0 benchmarks
Liver-US (Liver Ultrasound Dataset for Medical Image Classification)
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases.
1 paper · 1 benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models LoRA-WiSE spans various dataset sizes, backbones, ranks, and…
1 paper · 0 benchmarks
This is a large-scale RF fingerprinting dataset, collected from 25 different LoRa-enabled IoT transmitting devices using USRP B210 receivers.
1 paper · 0 benchmarks
This dataset was collected during a LoRaWAN measurement campaign in a multi-room indoor office environment at the University of Siegen, Germany.
1 paper · 0 benchmarks
The LoRA Weight Recovery Attack (LoWRA) Bench is a comprehensive benchmark designed to evaluate Pre-Fine-Tuning (Pre-FT) weight recovery methods as presented in the "Recovering the Pre-Fine-Tuning Weights of Generative Models" paper.
1 paper · 0 benchmarks
LocoVR (LocoVR: Multiuser Indoor Locomotion Dataset in Virtual Reality)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Loucount is a retail object detection and and counting dataset with rich annotations in retail stores, which consists of 50, 394 images with more than 1.9 million object instances in 140 categories
1 paper · 0 benchmarks
A synthetic dataset of procedurally generated environments, dynamically simulated crowd flows, and statically derived “proxy” crowd flows (which have more error but are more efficient to compute), for model training and evaluation.
1 paper · 0 benchmarks
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
1 paper · 0 benchmarks
A 160B bilingual long-text dataset with 3 categories: holistic, aggregated and chaotic long texts.
1 paper · 0 benchmarks
LoT-insts contains over 25k classes whose frequencies are naturally long-tail distributed.
1 paper · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset reports the lower-limb kinematics and kinetics of ten able-bodied subjects walking at multiple inclines (± 0°, 5°, and 10°) and speeds (0.8 m/s, 1 m/s, and 1.2 m/s), running over level-ground at multiple speeds (1.8 m/s, 2…
1 paper · 0 benchmarks
This dataset contains simulated synthetic particle decays, simulated using the PhaseSpace library.
1 paper · 0 benchmarks
This is the supporting dataset for the ECCV 2024 paper "MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain".
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.