Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 135 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6433–6480 of 12,172
JRDB-Pose is a large-scale dataset and benchmark for multi-person pose estimation and tracking using videos captured from a social navigation robot.
2 papers · 0 benchmarks
JVS-MuSiC is a Japanese multispeaker singing-voice corpus called "JVS-MuSiC" with the aim to analyze and synthesize a variety of voices.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
This dataset contains information about Japanese word similarity including rare words.
2 papers · 0 benchmarks
Dataset for 'Jet Flavor Classification in High-Energy Physics with Deep Neural Networks'
2 papers · 0 benchmarks
K-SportsSum is a sports game summarization dataset with two characteristics: (1) K-SportsSum collects a large amount of data from massive games.
2 papers · 0 benchmarks
KANFace consists of 40K still images and 44K sequences (14.5M video frames in total) captured in unconstrained, real-world conditions from 1,045 subjects.
2 papers · 1 benchmark
KGRC-RDF-star is an RDF-star dataset converted from KGRC-RDF, which is a Knowledge graph dataset of novel stories.
2 papers · 0 benchmarks
We introduce KPI-EDGAR, a novel dataset for Joint Named Entity Recognition and Relation Extraction building on financial reports uploaded to the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, where the main objective is…
2 papers · 1 benchmark
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
KanHope (Kannada Hope speech dataset)
KanHope is a code mixed hope speech dataset for equality, diversity, and inclusion in Kannada, an under-resourced Dravidian language.
2 papers · 1 benchmark
KazNERD is a dataset for Kazakh named entity recognition.
2 papers · 0 benchmarks
This is a subset of Kinetics-400, introduced in Look, Listen and Learn by Relja Arandjelovic and Andrew Zisserman.
2 papers · 0 benchmarks
The Kinships dataset describes relationships between members of the Australian tribe Alyawarra and consists of 10,686 triples.
2 papers · 0 benchmarks
Korean tabular dataset is a collection of 1.4M tables with corresponding descriptions for unsupervised pre-training language models.
2 papers · 0 benchmarks
General Corpora for the Maltese Language.
2 papers · 0 benchmarks
A dataset for benchmarking keyphrase extraction and generation techniques from long document English scientific papers.
2 papers · 1 benchmark
Kvasir-Capsule dataset is the largest publicly released VCE dataset.
2 papers · 0 benchmarks
The L-Bird (Large-Bird) dataset contains nearly 4.8 million images which are obtained by searching images of a total of 10,982 bird species from the Internet.
2 papers · 0 benchmarks
LAW (The Laboratory for Web Algorithmics)
The Laboratory for Web Algorithmics (LAW) was established in 2002 at the Dipartimento di Scienze dell'Informazione (now merged in the Computer Science Department) of the Università degli studi di Milano.
2 papers · 0 benchmarks
Nouns extracted automatically from Bible translations across 1580 languages.
2 papers · 0 benchmarks
LEPISZCZE is an open-source comprehensive benchmark for Polish NLP and a continuous-submission leaderboard, concentrating public Polish datasets (existing and new) in specific tasks.
2 papers · 0 benchmarks
This dataset is based on the LFM-1b [ and the Cultural LFM-1b [2] datasets.
2 papers · 0 benchmarks
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
LPR4M (Livestreaming Product Recognition 4M)
LPR4M is a large-scale live commerce dataset, offering a significantly broader coverage of categories and diverse modalities such as video, image, and text.
2 papers · 0 benchmarks
LUMA (Learning from Uncertain and Multimodal Data)
LUMA is a multimodal dataset that consists of audio, image, and text modalities.
2 papers · 0 benchmarks
Dataset of validated OCT and Chest X-Ray images described and analyzed in "Deep learning-based classification and referral of treatable human diseases".
2 papers · 0 benchmarks
LaboroTVSpeech is a large-scale Japanese speech corpus built from broadcast TV recordings and their subtitles.
2 papers · 0 benchmarks
The LanguageNet (English) is a collection of sentence level paraphrases from Twitter by linking tweets through shared URLs.
2 papers · 0 benchmarks
"We built a large lung CT scan dataset for COVID-19 by curating data from 7 public datasets listed in the acknowledgements.
2 papers · 2 benchmarks
The Large-Scale CLIR Dataset is a retrieval dataset built for Cross-Language Information Retrieval (CLIR).
2 papers · 0 benchmarks
Large-scale Anomaly Detection (LAD) is a database to benchmark anomaly detection in video sequences, which is featured in two aspects.
2 papers · 0 benchmarks
This dataset presents a set of large-scale ridesharing Dial-a-Ride Problem (DARP) instances.
2 papers · 0 benchmarks
A large-scale video database for rain removal (LasVR), which consists of 316 rain videos.
2 papers · 0 benchmarks
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
Description Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation" (Li et al., 2022).
2 papers · 0 benchmarks
This dataset contains 510 focal stacks (49 different focal distances) from in-the-wild scenes with calculated depth from SFM.
2 papers · 0 benchmarks
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources.
2 papers · 0 benchmarks
The LeukemiaAttri dataset is a large-scale, multi-domain collection of microscopy images derived from leukemia patient samples, enriched with detailed morphological information.
2 papers · 2 benchmarks
LibriCount is a synthetic dataset for speaker count estimation.
2 papers · 0 benchmarks
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
2 papers · 1 benchmark
This data set contains 775 video sequences, captured in the wildlife park Lindenthal (Cologne, Germany) as part of the AMMOD project, using an Intel RealSense D435 stereo camera.
2 papers · 0 benchmarks
The dataset contains road networks taken from 50 most populous cities in the world.
2 papers · 0 benchmarks
An RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications.
2 papers · 0 benchmarks
The Live Comment Dataset is a large-scale dataset with 2,361 videos and 895,929 live comments that were written while the videos were streamed.
2 papers · 0 benchmarks
LoED (LoRaWAN at the Edge Dataset) is a dataset from nine LoRaWAN gateways collected in an urban environment.
2 papers · 0 benchmarks
LoTE-Animal (LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding)
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species.
2 papers · 1 benchmark
The LogiEval dataset is a benchmark suite designed for evaluating the logical reasoning abilities of prompt-based language models, particularly instruct-prompt large language models.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.