Home › Datasets › task › Active Learning

Active Learning datasets

archive 2025-07-28

16 datasets carry the task tag "Active Learning" (the task itself: Active Learning), ordered by the archive's paper count. Page 1 of 1: 16 shown of 16. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Active Learning datasets 1–16 of 16

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
MNIST-8M (Infinite MNIST)
MNIST8M is derived from the MNIST dataset by applying random deformations and translations to the dataset.
26 papers · 0 benchmarks
The DeepWeeds dataset consists of 17,509 images capturing eight different weed species native to Australia in situ with neighbouring flora.
23 papers · 0 benchmarks
FDST (Fudan-ShanghaiTech)
The Fudan-ShanghaiTech dataset (FDST) is a dataset for video crowd counting.
18 papers · 0 benchmarks
DialoGLUE is a natural language understanding benchmark for task-oriented dialogue designed to encourage dialogue research in representation-based transfer, domain adaptation, and sample-efficient task learning.
17 papers · 2 benchmarks
Groove (Groove MIDI Dataset)
The Groove MIDI Dataset (GMD) is composed of 13.6 hours of aligned MIDI and (synthesized) audio of human-performed, tempo-aligned expressive drumming.
16 papers · 2 benchmarks
A benchmark which bridges the gap between freely available, documented, and motivated artificial benchmarks and properties of real industrial problems.
13 papers · 0 benchmarks
HJDataset is a large dataset of Historical Japanese Documents with Complex Layouts.
5 papers · 0 benchmarks
Illness-dataset (Illness multi-domain textual dataset)
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
Specially designed to evaluate active learning for video object detection in road scenes.
5 papers · 0 benchmarks
COMP6 (COmprehensive Machine-learning Potential)
COMP6 is a benchmark for evaluating the extensibility of machine-learning based molecular potentials.
4 papers · 0 benchmarks
Arxiv GR-QC (General Relativity and Quantum Cosmology collaboration network)
Arxiv GR-QC (General Relativity and Quantum Cosmology) collaboration network is from the e-print arXiv and covers scientific collaborations between authors papers submitted to General Relativity and Quantum Cosmology category.
3 papers · 0 benchmarks
Goldfinch (GOogLe image-search Dataset)
Goldfinch is a dataset for fine-grained recognition challenges.
3 papers · 0 benchmarks
A benchmark for molecular machine learning where improvements in model performance can be immediately observed in the throughput of promising molecules synthesized in the lab.
3 papers · 0 benchmarks
L-Bird (Large-Bird)
The L-Bird (Large-Bird) dataset contains nearly 4.8 million images which are obtained by searching images of a total of 10,982 bird species from the Internet.
2 papers · 0 benchmarks
Annotating data is a time-consuming and costly task, but it is inherently required for supervised machine learning.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.