Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 106 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5041–5088 of 12,172
THUCNews (THU Chinese Text Classification)
The THUCNews Chinese text dataset is a large-scale Chinese text classification dataset.
4 papers · 0 benchmarks
TOSCA (TOSCA high-resolution)
Hi-resolution three-dimensional nonrigid shapes in a variety of poses for non-rigid shape similarity and correspondence experiments.
4 papers · 0 benchmarks
TRansPose is a large-scale multispectral dataset that combines stereo RGB-D, TIR (TIR) images, and object poses to promote transparent object research.
4 papers · 0 benchmarks
Our goal is to enable deep learning research in neuroscience by releasing the largest publicly available unencumbered database of EEG recordings.
4 papers · 1 benchmark
TemporalWiki is a lifelong benchmark for ever-evolving LMs that utilizes the difference between consecutive snapshots of English Wikipedia and English Wikidata for training and evaluation, respectively.
4 papers · 0 benchmarks
TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century.
4 papers · 4 benchmarks
> The data we use include 366 monthly series, 427 quarterly series and 518 yearly series.
4 papers · 0 benchmarks
The TimberSeg 1.0 dataset is composed of 220 images showing wood logs in various environments and conditions in Canada.
4 papers · 0 benchmarks
TimeBankPT is a corpus of Portuguese text with annotations about time.
4 papers · 1 benchmark
TimeHetNet (Meta Dataset for Time Series with heterogeneous networks)
This meta-dataset is composed of previously known datasets.
4 papers · 0 benchmarks
TinyVIRAT contains natural low-resolution activities.
4 papers · 0 benchmarks
Topo-boundary is a new benchmark dataset, named \textit{Topo-boundary}, for off-line topological road-boundary detection.
4 papers · 0 benchmarks
Contains 140 videos with multiple human created summaries, which were acquired in a controlled experiment.
4 papers · 0 benchmarks
Specifically designed for the evaluation of change point detection algorithms, consisting of 37 time series from various domains.
4 papers · 0 benchmarks
This is an entity-level Twitter Sentiment Analysis dataset.
4 papers · 1 benchmark
UIT-VSMEC (Vietnamese Social Media Emotion Corpus)
Emotion recognition is a higher approach or special case of sentiment analysis.
4 papers · 0 benchmarks
USIS10K (Large-scale Underwater Salient Instance Segmentation Dataset)
We construct the first large-scale dataset, USIS10K, for the underwater salient instance segmentation task, which contains 10,632 images and pixel-level annotations of 7 categories.
4 papers · 0 benchmarks
UTA-RLDD (University of Texas at Arlington Real-Life Drowsiness Dataset)
Consists of around 30 hours of video, with contents ranging from subtle signs of drowsiness to more obvious ones.
4 papers · 0 benchmarks
The ukiyo-e faces dataset comprises of 5209 images of faces from ukiyo-e prints.
4 papers · 0 benchmarks
Includes 11,771 samples of both human activities and falls performed by 30 subjects of ages ranging from 18 to 60 years.
4 papers · 0 benchmarks
UniProtQA consists of proteins and textual queries about their functions and properties.
4 papers · 1 benchmark
Us Vs. Them (Us vs. Them: A Dataset of Populist Attitudes, News Bias and Emotions)
Us vs.
4 papers · 0 benchmarks
VANiLLa is a dataset for Question Answering over Knowledge Graphs (KGQA) offering answers in natural language sentences.
4 papers · 0 benchmarks
VASR (Visual Analogies of Situation Recognition)
Visual Analogies of Situation Recognition (VASR) is a dataset for visual analogical mapping, adapting the classical word-analogy task into the visual domain.
4 papers · 1 benchmark
Covers 5 generic driving scenarios, with a total of 25 distinct action classes.
4 papers · 0 benchmarks
Visuelle 2.0 is a dataset containing real data for 5355 clothing products of the retail fast-fashion Italian company, Nuna Lie.
4 papers · 2 benchmarks
VIVA (Vision for Intelligent Vehicles and Applications)
The VIVA challenge’s dataset is a multimodal dynamic hand gesture dataset specifically designed with difficult settings of cluttered background, volatile illumination, and frequent occlusion for studying natural human activities in…
4 papers · 2 benchmarks
A Large Vision-Language Model Knowledge Editing Benchmark
4 papers · 0 benchmarks
A large collection of interaction-rich video data which are annotated and analyzed.
4 papers · 0 benchmarks
VOT2019 is a Visual Object Tracking benchmark for short-term tracking in RGB.
4 papers · 1 benchmark
Verse is a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.
4 papers · 0 benchmarks
ViMMRC (Vietnamese Multiple-choice Machine Reading Comprehension Corpus)
A challenging machine comprehension corpus with multiple-choice questions, intended for research on the machine comprehension of Vietnamese text.
4 papers · 0 benchmarks
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
VirtualHome2KG is a system for constructing and augmenting knowledge graphs (KGs) of daily living activities using virtual space.
4 papers · 0 benchmarks
VizNet-Sato is a dataset from the authors of Sato and is based on the VizNet dataset.
4 papers · 2 benchmarks
The VizWiz-VQA-Grounding dataset is a dataset that visually grounds answers to visual questions asked by people with visual impairments.
4 papers · 0 benchmarks
This is a subset of the TREC 2005 enterprise track data, and consists of 48 topics and 200 candidates per topic, with each candidate labeled as an expert or non-expert for the topic.
4 papers · 0 benchmarks
WEC-eng is a cross-document event coreference resolution dataset extracted from English Wikipedia.
4 papers · 0 benchmarks
WHU-Hi (Wuhan UAV-borne hyperspectral image)
WHU-Hi dataset (Wuhan UAV-borne hyperspectral image) is collected and shared by the RSIDEA research group of Wuhan University, and it could serve as a benchmark dataset for precise crop classification and hyperspectral image classification…
4 papers · 0 benchmarks
WHU-RS19 is a set of satellite images exported from Google Earth, which provides high-resolution satellite images up to 0.5 m.
4 papers · 0 benchmarks
WIKIPerson is a high-quality human-annotated visual person linking dataset based on Wikipedia.
4 papers · 0 benchmarks
WISDOM (Warehouse Instance Segmentation Dataset for Object Manipulation)
Synthetic training dataset of 50,000 depth images and 320,000 object masks using simulated heaps of 3D CAD models.
4 papers · 1 benchmark
This HCP data release includes high-resolution 3T MR scans from young healthy adult twins and non-twin siblings (ages 22-35) using four imaging modalities: structural images (T1w and T2w), resting-state fMRI (rfMRI), task-fMRI (tfMRI), and…
4 papers · 0 benchmarks
Warblr is a dataset for the acoustic detection of birds.
4 papers · 0 benchmarks
A large-scale dataset for multi-domain aspect-based summarization that attempts to spur research in the direction of open-domain aspect-based summarization.
4 papers · 0 benchmarks
WikiCLIR is a large-scale (German-English) retrieval data set for Cross-Language Information Retrieval (CLIR).
4 papers · 0 benchmarks
WikiGraphs is a dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.