Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 228 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10897–10944 of 12,172
WUSTL_EHMS_2020 (WUSTL EHMS 2020 Dataset for Internet of Medical Things (IoMT) Cybersecurity Research)
The WUSTL-EHMS-2020 dataset was created using a real-time Enhanced Healthcare Monitoring System (EHMS) testbed [1].
1 paper · 0 benchmarks
The provided dataset consists of high-quality realistic head models and combined EEG/MEG data which can be used for state-of-the-art methods in brain research, such as modern finite element methods (FEM) to compute the EEG/MEG forward…
1 paper · 0 benchmarks
WYWEB (https://github.com/baudzhou/WYWEB)
An evaluation bentchmark for classical Chinese.
1 paper · 0 benchmarks
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering.
1 paper · 0 benchmarks
This repository contains all the ensemble datasets (along with their meta-data) used in the manuscript "Wasserstein Distances, Geodesics and Barycenters of Merge Trees".
1 paper · 0 benchmarks
PROBLEM Waste management is a big problem in our country.
1 paper · 0 benchmarks
The Watch Your Mouth dataset is a custom silent speech dataset consisting of depth-only recordings of users silently mouthing full English sentences, captured using consumer-grade depth cameras such as the iPhone TrueDepth sensor.
1 paper · 0 benchmarks
It contains data from two different realities: Food.com, a well-known American recipe site, and Planeat, an Italian site that allows you to plan recipes to save food waste.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
About Dataset Context With increasing unrest, security cams should be armed with advance technology.
1 paper · 0 benchmarks
Wearanize+ includes overnight sleep data from 130 participants (one night each) using three different wearable devices: Zmax headband, Empatica E4 wristband, and ActivPAL leg patch, alongside full-scale PSG recorded with SomnoScreen Plus…
1 paper · 0 benchmarks
Test-driven benchmark to challenge LLMs to write long JavaScript React application GitHub Script
1 paper · 1 benchmark
WebBrain-Raw is a large-scale dataset built from English Wikipedia articles and their crawlable Wikipedia references.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on WebNLG dataset.
1 paper · 1 benchmark
WebGen-Bench WebGen-Bench is created to benchmark LLM-based agent's ability to generate websites from scratch.
1 paper · 0 benchmarks
WebLI (Web Language Image)
WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering,…
1 paper · 0 benchmarks
This dataset contains network traces collected in-lab and in a real-world setting.
1 paper · 0 benchmarks
https://proceedings.neurips.cc/paperfiles/paper/2022/hash/82ad13ec01f9fe44c01cb91814fd7b8c-Abstract-Conference.html
1 paper · 0 benchmarks
A dataset automatically generated using question generation neural models and alt-text video captions from the WebVid dataset, with 3M video-question-answer triplets.
1 paper · 0 benchmarks
http://websail-fe.cs.northwestern.edu/TabEL/#WebManual
1 paper · 0 benchmarks
Webis-ConcluGen-21 is a large-scale corpus of 136,996 samples of argumentative texts and their conclusions used for the task of generating informative conclusions.
1 paper · 0 benchmarks
We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications.
1 paper · 0 benchmarks
This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning.
1 paper · 0 benchmarks
Webly-Reference SR dataset is a test dataset for evaluating Ref-SR methods.
1 paper · 0 benchmarks
This dataset is used for user identity linkage across two online social networks in Chinese.
1 paper · 0 benchmarks
The dataset is a private dataset collected for automatic analysis of psychological distress.
1 paper · 1 benchmark
Werewolf Among Us is a dataset multimodal dataset for modeling persuasion behaviors.
1 paper · 0 benchmarks
Existing databases usually record posed or induced human behavior in individual or dyadic settings with biased annotations in which a basic emotional class label or a Valence–Arousal pair value represents the emotional states.
1 paper · 0 benchmarks
Manually labelled dataset of bird recordings from the species of interest inhabiting in the wetlands of the "Aiguamolls del Empord\{a}" natural park in Girona, Spain.
1 paper · 0 benchmarks
The Western Mediterranean Wetlands Bird Dataset is a collection of birds' vocalizations of different lengths that primarily consists of 5,795 labelled audio clips derived from 1,098 recordings, totalling 201.6 minutes or 12,096 seconds…
1 paper · 0 benchmarks
WhenAct (Temporal Human Action Localization in Lifestyle Vlogs)
We consider the task of temporal human action localization in lifestyle vlogs.
1 paper · 0 benchmarks
WhyAct is a dataset for identifying human action reasons in online videos, consisting of 1,077 visual actions manually annotated with their reasons.
1 paper · 0 benchmarks
WiFiCam dataset for through-wall imaging based on WiFi channel state information.
1 paper · 0 benchmarks
WiRLD (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
WiRLD_ (Wikidata Reference Logo Dataset)
The Wikidata Reference Logo Dataset (WiRLD), a comprehensive collection of reference logos specifically designed to address the challenges of large-scale logo identification.
1 paper · 0 benchmarks
WiTA (Writing in The Air)
WiTA (Writing in The Air) is a dataset for the challenging writing in the air (WiTA) task -- an elaborate task bridging vision and NLP.
1 paper · 0 benchmarks
This Wider-Test-200 dataset is introduced in the following paper: "Towards Unsupervised Blind Face Restoration using Diffusion Prior" Please visit our website and refer to our paper for more information on the dataset and our method:…
1 paper · 0 benchmarks
This public dataset contains 127,820 comments from Wikipedia Talk Pages labeled with whether or not they are toxic
1 paper · 0 benchmarks
The Wiki-Flick Event dataset for cross-modal event retrieval is a well-labelled but weakly-aligned dataset collected for cross-modality event retrieval.
1 paper · 0 benchmarks
We introduce a novel task for LVLMs, which involves reviewing the good and bad points of a given image.
1 paper · 0 benchmarks
Wiki-Reliability is the first dataset of English Wikipedia articles annotated with a wide set of content reliability issues.
1 paper · 0 benchmarks
Wiki-en is an annotated English dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Wiki-zh is an annotated Chinese dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Training data for Hebrew morphological word segmentation
1 paper · 0 benchmarks
A dataset comprising 8,551 ban evasion pairs on Wikipedia, where each pair comprises a parent account and the child account.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.