Home › Datasets › task › Self-Supervised Learning
Self-Supervised Learning datasets
archive 2025-07-28
47 datasets carry the task tag "Self-Supervised Learning" (the task itself: Self-Supervised Learning), ordered by the archive's paper count. Page 1 of 1: 47 shown of 47. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Self-Supervised Learning datasets 1–47 of 47
description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The CIFAR-100 dataset (Canadian Institute for Advanced Research, 100 classes) is a subset of the Tiny Images dataset and consists of 60000 32x32 color images.
9,045 papers · 51 benchmarks
Tiny ImageNet contains 100000 images of 200 classes (500 for each class) downsized to 64×64 colored images.
1,232 papers · 7 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations.
147 papers · 6 benchmarks
NSynth is a dataset of one shot instrumental notes, containing 305,979 musical notes with unique pitch, timbre and envelope.
138 papers · 2 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
Rendered synthetically using a library of standard 3D objects, and tests the ability to recognize compositions of object movements that require long-term reasoning.
51 papers · 3 benchmarks
This dataset includes time-series data generated by accelerometer and gyroscope sensors (attitude, gravity, userAcceleration, and rotationRate).
35 papers · 0 benchmarks
Semi-Supervised Object Detection on COCO 10% labeled data
28 papers · 2 benchmarks
Contains 349 COVID-19 CT images from 216 patients and 463 non-COVID-19 CTs.
28 papers · 0 benchmarks
CREMA-D is an emotional multimodal actor data set of 7,442 original clips from 91 actors.
28 papers · 7 benchmarks
MVSEC (Multi Vehicle Stereo Event Camera)
The Multi Vehicle Stereo Event Camera (MVSEC) dataset is a collection of data designed for the development of novel 3D perception algorithms for event based cameras.
28 papers · 2 benchmarks
Contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible.
22 papers · 1 benchmark
Gibson is an opensource perceptual and physics simulator to explore active and real-world perception.
22 papers · 0 benchmarks
MosMedData contains anonymised human lung computed tomography (CT) scans with COVID-19 related findings, as well as without such findings.
22 papers · 1 benchmark
Consists of 20k English biomedical entity mentions from Reddit expert-annotated with links to SNOMED CT, a widely-used medical knowledge graph.
21 papers · 0 benchmarks
The INRIA Aerial Image Labeling dataset is comprised of 360 RGB tiles of 5000×5000px with a spatial resolution of 30cm/px on 10 cities across the globe.
21 papers · 1 benchmark
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
This split was introduced in TEMI (BMVC 2023) Adaloglou, Nikolas, Felix Michels, Hamza Kalisch, and Markus Kollmann.
12 papers · 4 benchmarks
The Argoverse 2 Sensor Dataset is a collection of 1,000 scenarios with 3D object tracking annotations.
11 papers · 0 benchmarks
Source: BARThez: a Skilled Pretrained French Sequence-to-Sequence Model OrangeSum is a single-document extreme summarization dataset with two tasks: title and abstract.
8 papers · 1 benchmark
ACAV100M (Automatically Curated Audio-Visual)
ACAV100M processes 140 million full-length videos (total duration 1,030 years) which are used to produce a dataset of 100 million 10-second clips (31 years) with high audio-visual correspondence.
7 papers · 0 benchmarks
YUD+ (Additional Vanishing Point Labels for the York Urban Database)
YUD+ is a dataset containing additional Vanishing Point Labels for the York Urban Database.
7 papers · 0 benchmarks
The 3DSeg-8 is a collection of several publicly available 3D segmentation datasets from different medical imaging modalities, e.g.
6 papers · 0 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
Wild-Time is a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification.
6 papers · 0 benchmarks
NYU-VP is a new dataset for multi-model fitting, vanishing point (VP) estimation in this case.
5 papers · 0 benchmarks
StreetStyle is a large-scale dataset of photos of people annotated with clothing attributes, and use this dataset to train attribute classifiers via deep learning.
5 papers · 0 benchmarks
The Argoverse 2 Lidar Dataset is a collection of 20,000 scenarios with lidar sensor data, HD maps, and ego-vehicle pose.
4 papers · 0 benchmarks
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Unsupervised Domain Adaptation demonstrates great potential to mitigate domain shifts by transferring models from labeled source domains to unlabeled target domains.
4 papers · 3 benchmarks
DABS (Domain-Agnostic Benchmark for Self-supervised learning)
DABS is a domain-agnostic benchmark for self-supervised learning to encourage research and progress towards domain-agnostic methods.
4 papers · 1 benchmark
LIMUC (Labeled Images for Ulcerative Colitis)
The LIMUC dataset is the largest publicly available labeled ulcerative colitis dataset that compromises 11276 images from 564 patients and 1043 colonoscopy procedures.
4 papers · 1 benchmark
DCASE2014 is an audio classification benchmark.
3 papers · 0 benchmarks
The Argoverse 2 Map Change Dataset is a collection of 1,000 scenarios with ring camera imagery, lidar, and HD maps.
2 papers · 0 benchmarks
Extended Agriculture-Vision dataset comprises two parts: 1.
2 papers · 0 benchmarks
OADAT (OADAT: Experimental and Synthetic Clinical Optoacoustic Data for Standardized Image Processing)
An experimental and synthetic (simulated) OA raw signals and reconstructed image domain datasets rendered with different experimental parameters and tomographic acquisition geometries.
2 papers · 0 benchmarks
A classification dataset of radar spectrograms in i "ground surveillance" setting recorded with the Open Radar Initiative.
2 papers · 0 benchmarks
SSL4EO-S12 is a large-scale, global, multimodal, and multi-seasonal corpus of satellite imagery from the ESA Sentinel-1 & -2 satellite missions.
2 papers · 0 benchmarks
The Unified SSL Benchmark (USB) consists of 15 diverse, challenging, and comprehensive tasks from CV, natural language processing (NLP), and audio processing (Audio) to evaluate self-supervised learning (SSL) methods.
2 papers · 0 benchmarks
Contains a large number of online videos and subtitles.
1 paper · 0 benchmarks
GeoJEPAD is a multimodal dataset combining OpenStreetMap (OSM) data (attributes and geometries) with high-resolution aerial imagery from diverse urban areas.
1 paper · 0 benchmarks
The Sentinel-2 satellite carries 12 CMOS detectors for the VNIR bands, with adjacent detectors having overlapping fields of view that result in overlapping regions in level-1 B (L1B) images.
1 paper · 0 benchmarks
MAX-60K (Masked Autoencoder for X-ray Fluorescence 60K Dataset)
The dataset for masked autoencoder for X-ray fluorescence (XRF) is a following development after the dataset (Chao et al., 2022).
1 paper · 0 benchmarks
The scales of the data accessible through internet search engines can reach hundreds of millions, or even billions.
1 paper · 0 benchmarks
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.