Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 95 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4513–4560 of 12,172
The goal of this dataset is to probe video-language models for understanding of simple temporal relations like "before" and "after".
5 papers · 1 benchmark
A Dense-text Image Benchmark to evaluate large generation model's ability on text generation.
5 papers · 1 benchmark
Thyroid is a dataset for detection of thyroid diseases, in which patients diagnosed with hypothyroid or subnormal are anomalies against normal patients.
5 papers · 1 benchmark
Timers and Such is an open source dataset of spoken English commands for common voice control use cases involving numbers.
5 papers · 1 benchmark
ToLD-Br (Toxic Language Detection for Brazilian Portuguese)
The Toxic Language Detection for Brazilian Portuguese (ToLD-Br) is a dataset with tweets in Brazilian Portuguese annotated according to different toxic aspects.
5 papers · 1 benchmark
ToT is a benchmark for evaluating LLMs on temporal reasoning.
5 papers · 0 benchmarks
The Tobacco-3482 dataset consists of document images belonging to 10 classes such as letter, form, email, resume, memo, etc.
5 papers · 2 benchmarks
The ToolE dataset encompasses various types of user queries in the form of prompts.
5 papers · 0 benchmarks
ToyADMOS2 is a dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.
5 papers · 0 benchmarks
A collection of 840 pretrained VQA models which may be regular “clean” models or malicious “backdoored” models which have been trained to include a secret backdoor trigger and behavior.
5 papers · 0 benchmarks
TutorialVQA is a new type of dataset used to find answer spans in tutorial videos.
5 papers · 0 benchmarks
80k tweets annotated concerning Inappropriate Speech (more particularly in matters of Abusive and Hateful speech) as well as Normal and Spam.
5 papers · 0 benchmarks
Twitter-MEL is a multimodal entity linking (MEL) dataset built from Twitter.
5 papers · 0 benchmarks
Twitter100k is a large-scale dataset for weakly supervised cross-media retrieval.
5 papers · 0 benchmarks
The UCLA Aerial Event Dataest has been captured by a low-cost hex-rotor with a GoPro camera, which is able to eliminate the high frequency vibration of the camera and hold in air autonomously through a GPS and a barometer.
5 papers · 0 benchmarks
UESTC RGB-D Varying-view action database contains 40 categories of aerobic exercise.
5 papers · 1 benchmark
UIT-ViNewsQA is a new corpus for the Vietnamese language to evaluate healthcare reading comprehension models.
5 papers · 0 benchmarks
UPLight is an underwater RGB-Polarization multimodal semantic segmentation dataset with 12 typical underwater semantic classes.
5 papers · 1 benchmark
UruDendro (UruDendro, a public dataset of cross-section images of pinus taeda)
UruDendro is a database of wood cross section images of commercially grown Pinus taeda trees from northern Uruguay.
5 papers · 1 benchmark
VCSL (Video Copy Segment Localization)
VCSL (Video Copy Segment Localization) is a new comprehensive segment-level annotated video copy dataset.
5 papers · 0 benchmarks
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
We present the VIS30K dataset, a collection of 29,689 images that represents 30 years of figures and tables from each track of the IEEE Visualization conference series (Vis, SciVis, InfoVis, VAST).
5 papers · 0 benchmarks
VNAT (VPN/NONVPN NETWORK APPLICATION TRAFFIC DATASET)
This dataset is a collection of labelled PCAP files, both encrypted and unencrypted, across 10 applications, as well as a pandas dataframe in HDF5 format containing detailed metadata summarizing the connections from those files.
5 papers · 0 benchmarks
VQA-VS (a new VQA benchmark considering Varying Shortcuts)
The current OOD benchmark VQA-CP v2 only considers one type of shortcut (from question type to answer) and thus still cannot guarantee that the modelrelies on the intended solution rather than a solution specific to this shortcut.
5 papers · 0 benchmarks
The Vent dataset is a large annotated dataset of text, emotions, and social connections.
5 papers · 0 benchmarks
A large-scale multi-modal dataset to facilitate research and studies that concentrate on vision-wireless systems.
5 papers · 1 benchmark
VidChapters-7M is a dataset of 817K user-chaptered videos including 7M chapters in total.
5 papers · 4 benchmarks
VidOR (Video Object Relation) dataset contains 10,000 videos (98.6 hours) from YFCC100M collection together with a large amount of fine-grained annotations for relation understanding.
5 papers · 1 benchmark
Due to the lack of training data for video waterdrop removal, we propose a large-scale synthetic dataset with simulated waterdrops in complex driving scenes on rainy days.
5 papers · 1 benchmark
VinDr-RibCXR is a benchmark dataset for automatic segmentation and labeling of individual ribs from chest X-ray (CXR) scans.
5 papers · 0 benchmarks
A large-scale corpus for phonetic typology, with aligned segments and estimated phoneme-level labels in 690 readings spanning 635 languages, along with acoustic-phonetic measures of vowels and sibilants.
5 papers · 0 benchmarks
WDBC (Breast Cancer Wisconsin (Diagnostic))
Features are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass.
5 papers · 0 benchmarks
WEAR (WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition)
WEAR is an outdoor sports dataset for both vision- and inertial-based human activity recognition (HAR).
5 papers · 0 benchmarks
WGISD (Embrapa Wine Grape Instance Segmentation Dataset)
Embrapa Wine Grape Instance Segmentation Dataset (WGISD) contains grape clusters properly annotated in 300 images and a novel annotation methodology for segmentation of complex objects in natural images.
5 papers · 0 benchmarks
🤖 Robo3D - The WOD-C Benchmark WOD-C is an evaluation benchmark heading toward robust and reliable 3D perception in autonomous driving.
5 papers · 1 benchmark
WWW Crowd provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a comprehensive dataset for the area of crowd understanding.
5 papers · 0 benchmarks
WiGesture (Wireless Sensing Dataset for Gesture Recognition and People ID Identification with ESP32)
WiGesture dataset contains data related to gesture recognition and people id identification in a meeting room scenario.
5 papers · 2 benchmarks
WikiHowQA is a Community-based Question Answering dataset, which can be used for both answer selection and abstractive summarization tasks.
5 papers · 0 benchmarks
This dataset gathers 428,748 person and 12,236 animal infobox with descriptions based on Wikipedia dump (2018/04/01) and Wikidata (2018/04/12).
5 papers · 3 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
The WorldKG knowledge graph is a comprehensive large-scale geospatial knowledge graph based on OpenStreetMap that provides a semantic representation of geographic entities from over 188 countries.
5 papers · 0 benchmarks
YACLC (Yet Another Chinese Learner Corpus)
YACLC is a large scale, multidimensional annotated Chinese learner corpus.
5 papers · 0 benchmarks
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
5 papers · 1 benchmark
ZeroWaste is a dataset for automatic waste detection and segmentation.
5 papers · 0 benchmarks
The aGender corpus contains audio recordings of predefined utterances and free speech produced by humans of different age and gender.
5 papers · 0 benchmarks
aiMotive dataset is a multimodal dataset for robust autonomous driving with long-range perception.
5 papers · 1 benchmark
australian (Statlog (Australian Credit Approval) Data Set)
Data Set Information: This file concerns credit card applications.
5 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.