Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 43 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2017–2064 of 12,172
Consists of 1106 action samples from seven actions with quality scores as measured by expert human judges.
20 papers · 1 benchmark
AVSBench is a pixel-level audio-visual segmentation benchmark that provides ground truth labels for sounding objects.
20 papers · 0 benchmarks
The Amazon-Google dataset for entity resolution derives from the online retailers Amazon.com and the product search service of Google accessible through the Google Base Data API.
20 papers · 2 benchmarks
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
Node classification on Chameleon with the fixed 48%/32%/20% splits provided by Geom-GCN.
20 papers · 2 benchmarks
Common Objects in 3D is a large-scale dataset with real multi-view images of object categories annotated with camera poses and ground truth 3D point clouds.
20 papers · 1 benchmark
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
20 papers · 2 benchmarks
The DiCOVA Challenge dataset is derived from the Coswara dataset, a crowd-sourced dataset of sound recordings from COVID-19 positive and non-COVID-19 individuals.
20 papers · 1 benchmark
EORSSD (Extended Optical Remote Sensing Saliency Detection)
The Extended Optical Remote Sensing Saliency Detection (EORSSD) dataset is an extension of the ORSSD dataset.
20 papers · 0 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
This dataset has 1,842 images with pixel-level DR-related lesion annotations, and 1,000 images with image-level labels graded by six board-certified ophthalmologists with intra-rater consistency.
20 papers · 0 benchmarks
FMB Dataset (Full-time Multi-modality Benchmark Dataset)
FMB contains 1500 well-registered infrared and visible image pairs with 14 annotated pixel-level categories.
20 papers · 1 benchmark
GEOM-DRUGS is a dataset of 430,000 large organic molecules of up to 180 atoms from Axelrod and Gómez-Bombarelli, Nature Scientific Data, 2022.
20 papers · 1 benchmark
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
The HIV dataset was introduced by the Drug Therapeutics Program (DTP) AIDS Antiviral Screen, which tested the ability to inhibit HIV replication for over 40,000 compounds.
20 papers · 5 benchmarks
Harm-P contains around 3,000 memes related to US politics.
20 papers · 1 benchmark
JNLPBA is a biomedical dataset that comes from the GENIA version 3.02 corpus (Kim et al., 2003).
20 papers · 2 benchmarks
Using conservation of energy -- a fundamental property of closed classical and quantum mechanical systems -- we develop an efficient gradient-domain machine learning (GDML) approach to construct accurate molecular force fields using a…
20 papers · 0 benchmarks
MINTAKA is a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models.
20 papers · 0 benchmarks
MLQE-PE (Multilingual Quality Estimation and Automatic Post-editing Dataset)
The Multilingual Quality Estimation and Automatic Post-editing (MLQE-PE) Dataset is a dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE).
20 papers · 0 benchmarks
The dataset was created for video quality assessment problem.
20 papers · 2 benchmarks
Spatio-temporal action detection is an important and challenging problem in video understanding.
20 papers · 2 benchmarks
Obstacle Tower is a high fidelity, 3D, 3rd person, procedurally generated environment for reinforcement learning.
20 papers · 6 benchmarks
PASCAL VOC 2011 is an image segmentation dataset.
20 papers · 2 benchmarks
Paralex learns from a collection of 18 million question-paraphrase pairs scraped from WikiAnswers.
20 papers · 1 benchmark
The dataset refers to the traffic speed data in San Francisco Bay Area, containing 307 sensors on 29 roads.
20 papers · 1 benchmark
PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging.
20 papers · 2 benchmarks
PlantDoc is a dataset for visual plant disease detection.
20 papers · 1 benchmark
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from politifact.com.
20 papers · 1 benchmark
A large-scale English dataset for coreference resolution.
20 papers · 1 benchmark
PsyQA is a Chinese Dataset for generating long counseling text for mental health support.
20 papers · 0 benchmarks
PubMed 200k RCT is new dataset based on PubMed for sequential sentence classification.
20 papers · 0 benchmarks
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
20 papers · 0 benchmarks
RCTW-17 (Reading Chinese Text in the Wild)
Features a large-scale dataset with 12,263 annotated images.
20 papers · 0 benchmarks
he RSSCN7 dataset contains satellite images acquired from Google Earth, which is originally collected for remote sensing scene classification.
20 papers · 1 benchmark
20 real low-resolution images selected from existing datasets or downloaded from internet
20 papers · 0 benchmarks
RoboCup is an initiative in which research groups compete by enabling their robots to play football matches.
20 papers · 0 benchmarks
SF-XL (San Francisco eXtra Large)
Large scale dataset for visual geo-localization / visual place recognition.
20 papers · 0 benchmarks
SHREC'19 (SHREC'19 track Matching Humans with Different Connectivity)
Shape matching plays an important role in geometry processing and shape analysis.
20 papers · 1 benchmark
SUBJ (Subjectivity dataset)
Available are collections of movie-review documents labeled with respect to their overall sentiment polarity (positive or negative) or subjective rating (e.g., "two and a half stars") and sentences labeled with respect to their…
20 papers · 1 benchmark
Large-scale manually-annotated corpus for 1,000 scientific papers (on computational linguistics) for automatic summarization.
20 papers · 0 benchmarks
ShipsEar (ShipsEar: An underwater vessel noise database)
This contribution presents a database of underwater sounds produced by vessels of various types.
20 papers · 0 benchmarks
The Social Chemistry 101 dataset is a collection of over 290,000 rules of thumb (ROTs) for evaluating people’s behavior in everyday social situations.
20 papers · 0 benchmarks
StreetHazards is a synthetic dataset for anomaly detection, created by inserting a diverse array of foreign objects into driving scenes and re-render the scenes with these novel objects.
20 papers · 1 benchmark
SynLiDAR is a large-scale synthetic LiDAR sequential point cloud dataset with point-wise annotations.
20 papers · 1 benchmark
Briefly describe the dataset.
20 papers · 2 benchmarks
VIDIT (Virtual Image Dataset for Illumination Transfer)
VIDIT is a reference evaluation benchmark and to push forward the development of illumination manipulation methods.
20 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.