Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 56 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2641–2688 of 12,172
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
GenAI-Bench (GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation)
GenAI-Bench is a benchmarking framework designed to evaluate and improve compositional text-to-visual generation models.
13 papers · 0 benchmarks
The Ghostbusters dataset leverages the GPT-3.5-turbo model for generating texts in the domains of creative writing, news, and student essays, providing 2,000 texts in the first two domains and 1,994 in the latter.
13 papers · 1 benchmark
GooAQ is a large-scale dataset with a variety of answer types.
13 papers · 0 benchmarks
HEIM (Holistic Evaluation of Text-to-Image Models)
HEIM stands for Holistic Evaluation of Text-To-Image Models.
13 papers · 0 benchmarks
HINT3 is a dataset for intent detection.
13 papers · 0 benchmarks
HLW (Horizon Lines in the Wild)
We introduce Horizon Lines in the Wild (HLW), a large dataset of real-world images with labeled horizon lines, captured in a diverse set of environments.
13 papers · 1 benchmark
The Habitat-Matterport 3D Semantics Dataset (HM3DSem) is the largest-ever dataset of 3D real-world and indoor spaces with densely annotated semantics that is available to the academic community.
13 papers · 0 benchmarks
HopeEDI (HopeEDI: A Multilingual Hope Speech Detection Dataset for Equality, Diversity, and Inclusion)
Over the past few years, systems have been developed to control online content and eliminate abusive, offensive or hate speech content.
13 papers · 4 benchmarks
The IDiff-Face dataset was proposed in the paper "IDiff-Face: Synthetic-based Face Recognition through Fizzy Identity-Conditioned Diffusion Models".
13 papers · 0 benchmarks
A dataset of ~19K questions that are elicited while a person is reading through a document.
13 papers · 0 benchmarks
The goal for ISIC 2019 is classify dermoscopic images among nine different diagnostic categories.25,331 images are available for training across 8 different categories.
13 papers · 3 benchmarks
A benchmark which bridges the gap between freely available, documented, and motivated artificial benchmarks and properties of real industrial problems.
13 papers · 0 benchmarks
A video dataset for benchmarking upsampling methods.
13 papers · 0 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
JVS is a Japanese multi-speaker voice corpus which contains voice data of 100 speakers in three styles (normal, whisper, and falsetto).
13 papers · 0 benchmarks
The KLEJ benchmark (Kompleksowa Lista Ewaluacji Językowych) is a set of nine evaluation tasks for the Polish language understanding task.
13 papers · 0 benchmarks
KorQuAD (The Korean Question Answering Dataset)
KorQuAD is a large-scale question-and-answer dataset constructed for Korean machine reading comprehension, and investigate the dataset to understand the distribution of answers and the types of reasoning required to answer the question.
13 papers · 0 benchmarks
KorSTS is a dataset for semantic textural similarity (STS) in Korean.
13 papers · 0 benchmarks
L3DAS22: MACHINE LEARNING FOR 3D AUDIO SIGNAL PROCESSING This dataset supports the L3DAS22 IEEE ICASSP Gand Challenge.
13 papers · 0 benchmarks
The real captured dataset of LOL contains 500 low/normallight image pairs.
13 papers · 1 benchmark
Logo-2K+:A Large-Scale Logo Dataset for Scalable Logo Classification The Logo-2K+ dataset contains a diverse range of logo classes from real-world logo images.
13 papers · 0 benchmarks
The first summarization collection containing question-driven summaries of answers to consumer health questions.
13 papers · 0 benchmarks
500 video clips for 50 different identity document types with ground truth.
13 papers · 0 benchmarks
Contains 300 scans of 10 people in a wide range of poses together with an evaluation methodology.
13 papers · 0 benchmarks
This is a dataset for a super-resolution task.
13 papers · 1 benchmark
MedConceptsQA - Open Source Medical Concepts QA Benchmark The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA
13 papers · 2 benchmarks
Provides detailed, graph-based annotations of social situations depicted in movie clips.
13 papers · 0 benchmarks
The Multilingual Reuters Collection dataset comprises over 11,000 articles from six classes in five languages, i.e., English (E), French (F), German (G), Italian (I), and Spanish (S).
13 papers · 0 benchmarks
Mutagenicity is a chemical compound dataset of drugs, which can be categorized into two classes: mutagen and non-mutagen.
13 papers · 1 benchmark
NTU4DRadLM is a novel 4D radar dataset specifically proposed for research on robust SLAM, based on 4D radar, thermal camera, and IMU.
13 papers · 0 benchmarks
The NaturalProofs Dataset is a large-scale dataset for studying mathematical reasoning in natural language.
13 papers · 0 benchmarks
It contains 15K triplets of essay problem statements, student-written, and LLM-generated essays.
13 papers · 0 benchmarks
PARANMT-50M is a dataset for training paraphrastic sentence embeddings.
13 papers · 0 benchmarks
PISC (People in Social Context)
The People in Social Context (PISC) dataset is a dataset that focuses on social relationships.
13 papers · 1 benchmark
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
PreSIL (Precise Synthetic Image and LiDAR)
Consists of over 50,000 frames and includes high-definition images with full resolution depth information, semantic segmentation (images), point-wise segmentation (point clouds), and detailed annotations for all vehicles and people.
13 papers · 0 benchmarks
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.
13 papers · 1 benchmark
RGB-Stacking is a benchmark for vision-based robotic manipulation.
13 papers · 3 benchmarks
ROSCOE is a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics.
13 papers · 0 benchmarks
Real HSI (End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention)
End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention
13 papers · 1 benchmark
The goal of the Robust track is to improve the consistency of retrieval technology by focusing on poorly performing topics.
13 papers · 1 benchmark
The SAP benchmark is a significant development in the realm of attack prompt generation for red teaming and defending large language models (LLMs).
13 papers · 0 benchmarks
SARA (StAtutory Reasoning Assessment)
A dataset for statutory reasoning in tax law entailment and question answering.
13 papers · 0 benchmarks
The SARDet-100K dataset encompasses a total of 116,598 images, and 245,653 instances distributed across six categories: Aircraft, Ship, Car, Bridge, Tank, and Harbor.
13 papers · 1 benchmark
SEN12MS-CR is a multi-modal and mono-temporal data set for cloud removal.
13 papers · 1 benchmark
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.
13 papers · 2 benchmarks
SSCBench establishes a large-scale SSC benchmark in street views that facilitates the training of robust and generalizable SSC models.
13 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.