Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 28 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1297–1344 of 3,998
🤖 Robo3D - The WOD-C Benchmark WOD-C is an evaluation benchmark heading toward robust and reliable 3D perception in autonomous driving.
5 papers · 1 benchmark
WiGesture (Wireless Sensing Dataset for Gesture Recognition and People ID Identification with ESP32)
WiGesture dataset contains data related to gesture recognition and people id identification in a meeting room scenario.
5 papers · 2 benchmarks
WikiHowQA is a Community-based Question Answering dataset, which can be used for both answer selection and abstractive summarization tasks.
5 papers · 0 benchmarks
This dataset gathers 428,748 person and 12,236 animal infobox with descriptions based on Wikipedia dump (2018/04/01) and Wikidata (2018/04/12).
5 papers · 3 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
5 papers · 1 benchmark
mTVR is a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips.
5 papers · 0 benchmarks
7,672 human written natural language navigation instructions for routes in OpenStreetMap with a focus on visual landmarks.
5 papers · 2 benchmarks
The dataset consists in many runs of the same quantum circuit on different IBM quantum machines.
5 papers · 0 benchmarks
The $O2$Perm dataset is created from the Membrane Society of Australasia portal.
4 papers · 0 benchmarks
We established a 3D evaluation benchmark, 3D MM-Vet, to assess the 4-level capacity in embodied interaction scenarios, varying from basic perception to control statements generation.
4 papers · 1 benchmark
A Game Of Sorts is a collaborative image ranking task.
4 papers · 0 benchmarks
This dataset is described in the ALTA 2021 Shared Task website and associated CodaLab competition.
4 papers · 0 benchmarks
ARAUS (Affective Responses to Augmented Urban Soundscapes)
Choosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop…
4 papers · 0 benchmarks
ARC-DA (ARC Direct Answer Questions)
ARC Direct Answer Questions (ARC-DA) dataset consists of 2,985 grade-school level, direct-answer ("open response", "free form") science questions derived from the ARC multiple-choice question set released as part of the AI2 Reasoning…
4 papers · 0 benchmarks
Advising Corpus is a dataset based on an entirely new collection of dialogues in which university students are being advised which classes to take.
4 papers · 1 benchmark
Data set covering a set of debatable topics, where for each topic and stance, a set of triplets of the form is provided.
4 papers · 0 benchmarks
AutoChart is a dataset for chart-to-text generation, a task that consists on generating analytical descriptions of visual plots.
4 papers · 0 benchmarks
A whole-body FDG-PET/CT dataset with manually annotated tumor lesions (FDG-PET-CT-Lesions) 1,014 studies (900 patients)
4 papers · 0 benchmarks
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Bentham manuscripts refers to a large set of documents that were written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832).
4 papers · 1 benchmark
This brain tumor dataset contains 3064 T1-weighted contrast-enhanced images with three kinds of brain tumor.
4 papers · 0 benchmarks
Unsupervised Domain Adaptation demonstrates great potential to mitigate domain shifts by transferring models from labeled source domains to unlabeled target domains.
4 papers · 3 benchmarks
CAsT-snippets is a high-quality dataset for conversational information seeking containing snippet-level annotations for all queries in the TREC CAsT 2020 and 2022 datasets.
4 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CEDAR Signature is a database of off-line signatures for signature verification.
4 papers · 1 benchmark
CHOCOLATE (Captions Have Often ChOsen Lies About The Evidence)
CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions.
4 papers · 4 benchmarks
CI-MNIST (Correlated and Imbalanced MNIST)
CI-MNIST (Correlated and Imbalanced MNIST) is a variant of MNIST dataset with introduced different types of correlations between attributes, dataset features, and an artificial eligibility criterion.
4 papers · 0 benchmarks
CODA-19 is a human-annotated dataset that denotes the Background, Purpose, Method, Finding/Contribution, and Other for 10,966 English abstracts in the COVID-19 Open Research Dataset.
4 papers · 0 benchmarks
CSPubSum is a dataset for summarisation of computer science publications, created by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.
4 papers · 0 benchmarks
We introduce the Cambridge Law Corpus (CLC), a corpus for legal AI research.
4 papers · 0 benchmarks
The nine (moving camera) videos in this benchmark exhibit camouflaged animals that are difficult to see in a single frame, but can be detected based upon their motion across frames.
4 papers · 1 benchmark
Children's Song Dataset is open source dataset for singing voice research.
4 papers · 0 benchmarks
The ChineseLP dataset contains 411 vehicle images (mostly of passenger cars) with Chinese license plates (LPs).
4 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
ClimART (Climate Atmospheric Radiative Transfer)
Numerical simulations of Earth's weather and climate require substantial amounts of computation.
4 papers · 0 benchmarks
The topic of Climate Change (CC) has received limited attention in NLP despite its real world urgency.
4 papers · 1 benchmark
CoDesc is a large dataset of 4.2m Java source code and parallel data of their description from code search, and code summarization studies.
4 papers · 2 benchmarks
The CoNaLa Extended With Question Text is an extension to the original CoNaLa Dataset (Papers With Code Link) proposed in the NLP4Prog workshop paper "Reading StackOverflow Encourages Cheating: Adding Question Text Improves Extractive Code…
4 papers · 1 benchmark
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
CompMix is a crowdsourced QA benchmark which naturally demands the integration of a mixture of input sources.
4 papers · 0 benchmarks
Concadia is a publicly available Wikipedia-based corpus, which consists of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.
4 papers · 0 benchmarks
ConvRef is a conversational QA benchmark with reformulations.
4 papers · 0 benchmarks
We present the CrackVision12k dataset, a collection of 12,000 crack images derived from 13 publicly available crack datasets.
4 papers · 1 benchmark
DAMON (Dense Annotation of 3D Human Object contact in Natural Images)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
DDRel is a dataset for interpersonal relation classification in dyadic dialogues.
4 papers · 1 benchmark
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
DRAW-1K (Diverse Algebra Word Problem Set)
DRAW-1K is a dataset consisting of 1000 algebra word problems, semiautomatically annotated for the evaluation of automatic solvers.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.