Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 28 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1297–1344 of 3,998

🤖 Robo3D - The WOD-C Benchmark WOD-C is an evaluation benchmark heading toward robust and reliable 3D perception in autonomous driving.
5 papers · 1 benchmark
WiGesture (Wireless Sensing Dataset for Gesture Recognition and People ID Identification with ESP32)
WiGesture dataset contains data related to gesture recognition and people id identification in a meeting room scenario.
5 papers · 2 benchmarks
WikiHowQA is a Community-based Question Answering dataset, which can be used for both answer selection and abstractive summarization tasks.
5 papers · 0 benchmarks
This dataset gathers 428,748 person and 12,236 animal infobox with descriptions based on Wikipedia dump (2018/04/01) and Wikidata (2018/04/12).
5 papers · 3 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
Zero-shot Video Question Answering on LongVideoBench (A Benchmark for Long-context Interleaved Video-Language Understanding)
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
5 papers · 1 benchmark
mTVR is a large-scale multilingual video moment retrieval dataset, containing 218K English and Chinese queries from 21.8K TV show video clips.
5 papers · 0 benchmarks
7,672 human written natural language navigation instructions for routes in OpenStreetMap with a focus on visual landmarks.
5 papers · 2 benchmarks
The dataset consists in many runs of the same quantum circuit on different IBM quantum machines.
5 papers · 0 benchmarks
$O_2$Perm (Oxygen Permeability)
The $O2$Perm dataset is created from the Membrane Society of Australasia portal.
4 papers · 0 benchmarks
We established a 3D evaluation benchmark, 3D MM-Vet, to assess the 4-level capacity in embodied interaction scenarios, varying from basic perception to control statements generation.
4 papers · 1 benchmark
ALTA 2021 Shared Task (Automatic Grading of Evidence, 10 years later)
This dataset is described in the ALTA 2021 Shared Task website and associated CodaLab competition.
4 papers · 0 benchmarks
ARAUS (Affective Responses to Augmented Urban Soundscapes)
Choosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop…
4 papers · 0 benchmarks
ARC-DA (ARC Direct Answer Questions)
ARC Direct Answer Questions (ARC-DA) dataset consists of 2,985 grade-school level, direct-answer ("open response", "free form") science questions derived from the ARC multiple-choice question set released as part of the AI2 Reasoning…
4 papers · 0 benchmarks
Advising Corpus is a dataset based on an entirely new collection of dialogues in which university students are being advised which classes to take.
4 papers · 1 benchmark
Data set covering a set of debatable topics, where for each topic and stance, a set of triplets of the form is provided.
4 papers · 0 benchmarks
AutoChart is a dataset for chart-to-text generation, a task that consists on generating analytical descriptions of visual plots.
4 papers · 0 benchmarks
A whole-body FDG-PET/CT dataset with manually annotated tumor lesions (FDG-PET-CT-Lesions) 1,014 studies (900 patients)
4 papers · 0 benchmarks
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Bentham (Bentham project)
Bentham manuscripts refers to a large set of documents that were written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832).
4 papers · 1 benchmark
This brain tumor dataset contains 3064 T1-weighted contrast-enhanced images with three kinds of brain tumor.
4 papers · 0 benchmarks
Unsupervised Domain Adaptation demonstrates great potential to mitigate domain shifts by transferring models from labeled source domains to unlabeled target domains.
4 papers · 3 benchmarks
CAsT-snippets is a high-quality dataset for conversational information seeking containing snippet-level annotations for all queries in the TREC CAsT 2020 and 2022 datasets.
4 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CEDAR Signature is a database of off-line signatures for signature verification.
4 papers · 1 benchmark
CHOCOLATE (Captions Have Often ChOsen Lies About The Evidence)
CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions.
4 papers · 4 benchmarks
CI-MNIST (Correlated and Imbalanced MNIST)
CI-MNIST (Correlated and Imbalanced MNIST) is a variant of MNIST dataset with introduced different types of correlations between attributes, dataset features, and an artificial eligibility criterion.
4 papers · 0 benchmarks
CODA-19 is a human-annotated dataset that denotes the Background, Purpose, Method, Finding/Contribution, and Other for 10,966 English abstracts in the COVID-19 Open Research Dataset.
4 papers · 0 benchmarks
CSPubSum is a dataset for summarisation of computer science publications, created by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.
4 papers · 0 benchmarks
Cambridge Law Corpus (The Cambridge Law Corpus: A Dataset for Legal AI Research)
We introduce the Cambridge Law Corpus (CLC), a corpus for legal AI research.
4 papers · 0 benchmarks
The nine (moving camera) videos in this benchmark exhibit camouflaged animals that are difficult to see in a single frame, but can be detected based upon their motion across frames.
4 papers · 1 benchmark
Children's Song Dataset is open source dataset for singing voice research.
4 papers · 0 benchmarks
The ChineseLP dataset contains 411 vehicle images (mostly of passenger cars) with Chinese license plates (LPs).
4 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
ClimART (Climate Atmospheric Radiative Transfer)
Numerical simulations of Earth's weather and climate require substantial amounts of computation.
4 papers · 0 benchmarks
The topic of Climate Change (CC) has received limited attention in NLP despite its real world urgency.
4 papers · 1 benchmark
CoDesc is a large dataset of 4.2m Java source code and parallel data of their description from code search, and code summarization studies.
4 papers · 2 benchmarks
CoNaLa-Ext (CoNaLa Extended With Question Text)
The CoNaLa Extended With Question Text is an extension to the original CoNaLa Dataset (Papers With Code Link) proposed in the NLP4Prog workshop paper "Reading StackOverflow Encourages Cheating: Adding Question Text Improves Extractive Code…
4 papers · 1 benchmark
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
CompMix is a crowdsourced QA benchmark which naturally demands the integration of a mixture of input sources.
4 papers · 0 benchmarks
Concadia is a publicly available Wikipedia-based corpus, which consists of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.
4 papers · 0 benchmarks
ConvRef is a conversational QA benchmark with reformulations.
4 papers · 0 benchmarks
We present the CrackVision12k dataset, a collection of 12,000 crack images derived from 13 publicly available crack datasets.
4 papers · 1 benchmark
DAMON (Dense Annotation of 3D Human Object contact in Natural Images)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
4 papers · 0 benchmarks
DDRel is a dataset for interpersonal relation classification in dyadic dialogues.
4 papers · 1 benchmark
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
DRAW-1K (Diverse Algebra Word Problem Set)
DRAW-1K is a dataset consisting of 1000 algebra word problems, semiautomatically annotated for the evaluation of automatic solvers.
4 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.