Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 72 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3409–3456 of 3,998

In this paper, we introduce a victim dataset for the RoboCup Rescue competitions.
1 paper · 0 benchmarks
The SWC is a corpus of aligned Spoken Wikipedia articles from the English, German, and Dutch Wikipedia.
1 paper · 1 benchmark
We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I).
1 paper · 0 benchmarks
This is not a Dataset (This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models)
We introduce a large semi-automatically generated dataset of ~400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms that we use to…
1 paper · 1 benchmark
10000 instances of three-view numerical data set with 4 clusters and 2 feature components are considered.
1 paper · 0 benchmarks
Thunder-NUBench (Negation Understanding Benchmark) is a benchmark specifically designed to evaluate large language models’ (LLMs) sentence-level understanding of negation.
1 paper · 0 benchmarks
TiROD (Tiny Robotics Object Detection)
Dataset to benchmark Continual Learning for Object Detection in a Tiny Robotics settings.
1 paper · 1 benchmark
The TimberVision dataset consists of more than 2k annotated RGB images and contains a total of 51k trunk components including cut and lateral surfaces, thereby surpassing any existing dataset in this domain in terms of both quantity and…
1 paper · 0 benchmarks
TimeGraph (TimeGraph: Synthetic Benchmark Datasets for Robust Time-Series Causal Discovery)
TimeGraph is a comprehensive suite of synthetic datasets designed to benchmark causal discovery algorithms on time-series data.
1 paper · 0 benchmarks
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and difficulties in generating custom QA pairs.
1 paper · 0 benchmarks
Tinto (Tinto: Multisensor Benchmark for 3D Hyperspectral Point Cloud Segmentation in the Geosciences)
The increasing use of deep learning techniques has reduced interpretation time and, ideally, reduced interpreter bias by automatically deriving geological maps from digital outcrop models.
1 paper · 0 benchmarks
Tiny ImageNet-A is a subset of the Tiny ImageNet test set consisting of 3,374 images comprising real-world, unmodified, and naturally occurring examples that are misclassified by ResNet-18.
1 paper · 0 benchmarks
TinyChirp dataset for model training, validation and testing
1 paper · 0 benchmarks
The data consists of a set of 3 task types and 4 question types, creating 12 total scenarios.
1 paper · 0 benchmarks
A prevalent use case of topic models is that of topic discovery.
1 paper · 1 benchmark
Toulouse Vanishing Points Dataset is a public photographs database of Manhattan scenes taken with an iPad Air 1.
1 paper · 0 benchmarks
TraVLR is a synthetic dataset comprising four visio-linguistic reasoning tasks.
1 paper · 0 benchmarks
This data set is being released to support the spam and context-specific spam detection tasks on Twitter data.
1 paper · 3 benchmarks
This repository contains data for the NeurIPS conference paper titled "Harnessing Machine Learning for Single-Shot Measurement of Free Electron Laser Pulse Power".
1 paper · 0 benchmarks
The dataset contains procedurally generated images of transparent vessels containing liquid and objects .
1 paper · 1 benchmark
Trilemma Dataset (The Trilemma of Truth in Large Language Models)
The Trilemma of Truth is a multiclass probing dataset for evaluating the veracity-tracking mechanism of large language models.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable.
1 paper · 0 benchmarks
Tsinghua Dogs is a fine-grained classification dataset for dogs, over 65% of whose images are collected from people's real life.
1 paper · 0 benchmarks
Tweet Sentiment Extraction (Sentiment Analysis: Emotion in Text tweets with existing sentiment labels)
"My ridiculous dog is amazing." [sentiment: positive] With all of the tweets circulating every second it is hard to tell whether the sentiment behind a specific tweet will impact a company, or a person's, brand for being viral (positive),…
1 paper · 0 benchmarks
This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.
1 paper · 0 benchmarks
This dataset contains two subsets of flood images from Twitter: The Harz17 dataset comprises images from tweets containing flood-related keywords during the occurrence of a flood in the Harz region in Germany in July 2017.
1 paper · 0 benchmarks
Twitter MediaEval (MediaEval Benchmarking Initiative for Multimedia Evaluation)
The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video).
1 paper · 0 benchmarks
Twitter PoS VCB (Twitter part-of-speech vote-constrained-bootstrapping)
The data is about 1.5 million English tweets annotated for part-of-speech using Ritter's extension of the PTB tagset.
1 paper · 0 benchmarks
We introduce a dataset consisting of 1314 samples, including users’ tweets and bios.
1 paper · 0 benchmarks
Twitter-HyDrug is a real-world hypergraph data that describes the drug trafficking communities on Twitter.
1 paper · 1 benchmark
Twitter-HyDrug-UR (Twitter Hypergraph Drug for User Roles)
This benchmark hypergraph dataset, Twitter-HyDrug-UR, is derived from Twitter-HyDrug by HyGCL-DC.
1 paper · 1 benchmark
This task aims to extract named entities and entity types while further predicting segmentation masks of visual objects.
1 paper · 1 benchmark
The two Coiling Spiral is a 2d classification dataset composed of two classes; each spiral corresponds to one class.
1 paper · 0 benchmarks
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
UAVBillboards (UAV Billboards)
Mapping urban large-area advertising structures using drone imagery and deep learning-based spatial data analysis.
1 paper · 1 benchmark
UAVDB (Trajectory-Guided Adaptable Bounding Boxes for UAV Detection)
UAVDB is a high-resolution RGB video dataset meticulously designed for UAV detection tasks across diverse scales and complex backgrounds.
1 paper · 1 benchmark
The UAVVaste dataset consists to date of 772 images and 3716 annotations.
1 paper · 1 benchmark
https://github.com/zzr-idam/Under-Display-Camera-UAV
1 paper · 0 benchmarks
The UFPR-ADMR-v2 dataset contains 5,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4M consuming units in the Brazilian state of Paraná.
1 paper · 0 benchmarks
The UFPR-VCR dataset contains 10,039 images of 9,502 distinct vehicles across various categories, including cars, vans, buses, and trucks.
1 paper · 0 benchmarks
UICaption is a dataset of 114k UI images paired with descriptions of their functionality.
1 paper · 0 benchmarks
UIUC Scooping Dataset (Granular Materials Manipulation Dataset with Scooping/Digging/Excavation Action)
Overview: This dataset encompasses a compilation of 6,700 executed scoops (excavations), mapped across a vast spectrum of materials, terrain topography, and compositions.
1 paper · 0 benchmarks
Definitions of jargon/terms in computer science, mathematics, and physics
1 paper · 0 benchmarks
This package contains an anonymized packets of 802.11 probe requests captured throughout March of 2023 at Universitat Jaume I.
1 paper · 0 benchmarks
UK Key Stage Readability (UK Key Stage Readability for English Texts)
Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting.
1 paper · 1 benchmark
Bangladesh's legal system struggles with major challenges like delays, complexity, high costs, and millions of unresolved cases, which deter many from pursuing legal action due to lack of knowledge or financial constraints.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.