Home › Datasets › task › Dimensionality Reduction

Dimensionality Reduction datasets

archive 2025-07-28

11 datasets carry the task tag "Dimensionality Reduction" (the task itself: Dimensionality Reduction), ordered by the archive's paper count. Page 1 of 1: 11 shown of 11. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Dimensionality Reduction datasets 1–11 of 11

EMNIST (Extended MNIST)
EMNIST (extended MNIST) has 4 times more data than MNIST.
264 papers · 10 benchmarks
CN-Celeb is a large-scale speaker recognition dataset collected in the wild'.
68 papers · 1 benchmark
Oxford105k is the combination of the Oxford5k dataset and 99782 negative images crawled from Flickr using 145 most popular tags.
44 papers · 0 benchmarks
STRING is a collection of protein-protein interaction (PPI) networks.
35 papers · 0 benchmarks
HolStep is a dataset based on higher-order logic (HOL) proofs, for the purpose of developing new machine learning-based theorem-proving strategies.
10 papers · 2 benchmarks
The Oxford-Affine dataset is a small dataset containing 8 scenes with sequence of 6 images per scene.
7 papers · 0 benchmarks
AtariARI (Atari Annotated RAM Interface)
The AtariARI (Atari Annotated RAM Interface) is an environment for representation learning.
6 papers · 0 benchmarks
GoodSounds dataset contains around 28 hours of recordings of single notes and scales played by 15 different professional musicians, all of them holding a music degree and having some expertise in teaching.
4 papers · 0 benchmarks
Deep Fakes Dataset (inamibora)
The Deep Fakes Dataset is a collection of "in the wild" portrait videos for deepfake detection.
3 papers · 0 benchmarks
SoF (Specs on Faces)
The Specs on Faces (SoF) dataset, a collection of 42,592 (2,662×16) images for 112 persons (66 males and 46 females) who wear glasses under different illumination conditions.
3 papers · 0 benchmarks
ALGAD (Andy Lomas Generative Art Dataset)
Repository of a generative art dataset by computer artist Andy Lomas.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.