Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 229 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10945–10992 of 12,172
The WikiBio GPT-3 Hallucination Dataset is a benchmark dataset used for hallucination detection.
1 paper · 0 benchmarks
WikiBioCTE is a dataset for controllable text edition based on the existing dataset WikiBio (originally created for table-to-text generation).
1 paper · 0 benchmarks
WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings.
1 paper · 0 benchmarks
WikiContradiction is a novel wiki dataset for self-contradiction Wikipedia article detection.
1 paper · 0 benchmarks
Topical subsets of WikiData, assembled using the WikiDataSets python library.
1 paper · 0 benchmarks
WikiDes is a dataset for generating descriptions of Wikidata from Wikipedia paragraphs.
1 paper · 0 benchmarks
WikiEvalFacts This is an augmented version of the WikiEval dataset which additionally includes generated fact statements for each QA pair and human annotation of their truthfulness against each answer type.
1 paper · 0 benchmarks
The gold-standard and automatically-developed fine-grained Arabic named entity corpora are resources created by annotating Named Entities into 50 fine-grained classes.
1 paper · 0 benchmarks
WikiFactDiff is a dataset designed as a resource to perform atomic factual knowledge updates on language models, with the goal of aligning them with current knowledge.
1 paper · 0 benchmarks
WikiMulti (WikiMulti: a Corpus for Cross-Lingual Summarization)
wikimulti is a dataset for cross-lingual summarization based on Wikipedia articles in 15 languages.
1 paper · 0 benchmarks
a high-level explanation of the dataset characteristics We introduce WikiOFGraph, a novel large-scale, domain-diverse dataset synthesized by LLMs, ensuring superior graph-text consistency to advance general-domain graph-to-text generation.
1 paper · 1 benchmark
WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction.
1 paper · 0 benchmarks
WikiQAar (English-Arabic Wikipedia Question-Answering)
A publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
1 paper · 0 benchmarks
To collect WikiSuggest, Google Suggest API is used to harvest natural language questions and submit them to Google Search.
1 paper · 0 benchmarks
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
Semi-inductive link prediction (LP) in knowledge graphs (KG) is the task of predicting facts for new, previously unseen entities based on context information.
1 paper · 1 benchmark
Wikidated 1.0 is a dataset of Wikidata's full revision history, which encodes changes between Wikidata revisions as sets of deletions and additions of RDF triples.
1 paper · 0 benchmarks
The Wikimedia dataset refers to a collection of data related to Wikimedia projects, which include Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others.
1 paper · 0 benchmarks
Wikipedia is the largest and most read online free encyclopedia currently existing.
1 paper · 0 benchmarks
Wikipedia users activity for two language editions, Portuguese and Italian, for up to 8 January 2020.
1 paper · 0 benchmarks
WildAvatar is web-scale in-the-wild human avatar creation dataset extracted from YouTube, with 10,000+ different human subjects and scenes.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
WildQA is a video understanding dataset of videos recorded in outside settings.
1 paper · 1 benchmark
WildestFaces is tailored to study cross-domain recognition under a variety of adverse conditions.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks
WinSyn (WinSyn: A High Resolution Testbed for Synthetic Data)
75k photos of windows + 21k synthetic renders of building windows.
1 paper · 0 benchmarks
Air pollution management through wind speed forecasting: the time series exhibits a daily cyclical behavior and a long-term seasonality.
1 paper · 0 benchmarks
Summary: This dataset contains an estimation of the average yearly wind speed and of the wind power potential for Switzerland, at a spatial resolution of 250 x 250 meters and over the period from 2008 to 2017.
1 paper · 0 benchmarks
Test set of sentences in Hindi with complex coreference involving two entities inspired by WinoBias format of sentences in English.
1 paper · 0 benchmarks
This dataset consists of Winograd schemas that test coreference resolution systems' ability to differentiate singular vs plural they/them pronouns.
1 paper · 0 benchmarks
WinoPron is a novel dataset of Winogender-like template pairs in English, which fixes inconsistencies in Winogender Schemas and contains balanced template pairs for pronoun forms in 3 grammatical cases, which we find impacts performance…
1 paper · 0 benchmarks
The Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning.
1 paper · 1 benchmark
https://ieee-dataport.org/documents/wi-fi-signal-strength-measurements-smartphone-various-hand-gestures
1 paper · 0 benchmarks
WordNet-feelings, is an affective dataset that identifies 3664 word senses as feelings, and associates each of these with one of the 9 categories of feeling.
1 paper · 0 benchmarks
The Workflow Trace Archive (WTA) is an open-access archive of workflow traces from diverse computing infrastructures.
1 paper · 0 benchmarks
The goal of this dataset is to understand how people experience sexism and sexual harassment in the workplace by discovering themes in 2,362 experiences posted on the Everyday Sexism Project's website Source:…
1 paper · 0 benchmarks
The works-magnet aims at getting visible the AI-processed metadata for scholarly outputs and help curators improve those metadata.
1 paper · 0 benchmarks
Workshop Tools Dataset Motivated by the need for a dataset that also includes inertial information about the objects, we contribute the following dataset.
1 paper · 0 benchmarks
The World Across Time (WAT) dataset used in paper "CLNeRF: Continual Learning Meets NeRF".
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We present the World Wide Dishes dataset which seeks to assess disparities in representations of food through a decentralised data collection effort to gather perspectives directly from people with a wide variety of backgrounds from around…
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
WorldFloods: a newly compiled dataset of 119 globally verified flooding events from disaster response organizations
1 paper · 1 benchmark
The X-MARS dataset proposes new splits for the MARS dataset, to allow for cross-evaluation with the Market-1501 dataset without training and test overlap between the two datasets.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
https://arxiv.org/abs/2505.15372
1 paper · 0 benchmarks
Collections of images of the same rotating plastic object made in X-ray and visible spectra.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.