Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 44 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2065–2112 of 3,998
Schema2QA is the first large question answering dataset over real-world Schema.org data.
2 papers · 0 benchmarks
A multimodal empathetic dialogue dataset.
2 papers · 0 benchmarks
StoryBench (StoryBench: A Multifaceted Benchmark for Continuous Story Visualization)
StoryBench is a multi-task benchmark to reliably evaluate the ability of text-to-video models to generate stories from a sequence of captions and their duration.
2 papers · 1 benchmark
Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study File Descriptions File | Description --- | --- commitcategorizations.csv | Categorizations for the commits in our dataset.
2 papers · 0 benchmarks
SuMe (A Dataset Towards Summarizing Biomedical Mechanisms)
Can language models read biomedical texts and explain the biomedical mechanisms discussed?
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
This is a discourse dataset with multiple and subjective interpretations of English conversation in the form of perceived conversation acts and intents.
2 papers · 0 benchmarks
SuperCaustics is a simulation tool made in Unreal Engine for generating massive computer vision datasets that include transparent objects.
2 papers · 0 benchmarks
SynMirror consists of samples rendered from 3D assets of two widely used 3D object datasets - Objaverse and Amazon Berkeley Objects (ABO) placed in front of a mirror in a virtual blender environment.
2 papers · 0 benchmarks
Event cameras are sensors that are inspired by biological systems and specialize in capturing changes in brightness.
2 papers · 1 benchmark
TAP (Traffic Accident Prediction data repository)
The Traffic Accident Prediction (TAP) data repository offers extensive coverage for 1,000 US cities (TAP-city) and 49 states (TAP-state), providing real-world road structure data that can be easily used for graph-based machine learning…
2 papers · 0 benchmarks
TBBR (Thermal Bridges on Building Rooftops)
The dataset of Thermal Bridges on Building Rooftops (TBBR dataset) consists of annotated combined RGB and thermal drone images with a height map.
2 papers · 2 benchmarks
AI for science has generated a great deal of enthusiasm from both academia and industry.
2 papers · 0 benchmarks
This collection includes datasets from 20 subjects with primary newly diagnosed glioblastoma who were treated with surgery and standard concomitant chemo-radiation therapy (CRT) followed by adjuvant chemotherapy.
2 papers · 0 benchmarks
A new text effects dataset with 141,081 text effect/glyph pairs in total.
2 papers · 0 benchmarks
THFOOD-50 (Thai Food 50 Image Classification)
Fine-Grained Thai Food Image Classification Datasets THFOOD-50 containing 15,770 images of 50 famous Thai dishes.
2 papers · 0 benchmarks
TI1K Dataset (Thumb Index 1000 Hand & Fingertip Detection Dataset)
Thumb Index 1000 (TI1K) is a dataset of 1000 hand images with the hand bounding box, and thumb and index fingertip positions.
2 papers · 0 benchmarks
TLDR9+ is a large-scale summarization dataset containing over 9 million training instances extracted from Reddit discussion forum.
2 papers · 1 benchmark
TNCR Dataset (Table Net Detection and Classification Dataset)
We present TNCR, a new table dataset with varying image quality collected from free open source websites.
2 papers · 0 benchmarks
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TUSC (Tweets from US and Canada)
Tweets from US and Canada (TUSC) is a large dataset of more than 45 million geo-located tweets posted between 2015 and 2021 from US and Canada (TUSC), especially curated for natural language analysis
2 papers · 0 benchmarks
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
This mouse cerebellar atlas can be used for mouse cerebellar morphometry.
2 papers · 0 benchmarks
Text2KGBench is a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology.
2 papers · 0 benchmarks
Tweets and items from psychological scales for sexism detection with counterfactual examples.
2 papers · 0 benchmarks
ThermoHands is the first benchmark dataset specifically designed for egocentric 3D hand pose estimation from thermal images.
2 papers · 0 benchmarks
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
Tiny ImageNet-R is a subset of the ImageNet-R dataset by Hendrycks et al.
2 papers · 0 benchmarks
Titanic (Titanic - Machine Learning from Disaster)
Titanic Dataset Description Overview The data is divided into two groups: - Training set (train.csv): Used to build machine learning models.
2 papers · 1 benchmark
TorWIC (The Toronto Warehouse Incremental Change Dataset)
TorWIC is the dataset discussed in POCD: Probabilistic Object-Level Change Detection and Volumetric Mapping in Semi-Static Scenes.
2 papers · 0 benchmarks
A benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.
2 papers · 0 benchmarks
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic.
2 papers · 0 benchmarks
The data set contains 2500 manually-stance-labeled tweets, 1250 for each candidate (Joe Biden and Donald Trump).
2 papers · 2 benchmarks
Two-Path Computational Graph (CG) family introduced in "GENNAPE: Towards Generalized Neural Architecture Performance Estimators", accepted to AAAI-23.
2 papers · 0 benchmarks
A database of several hundred high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
2 papers · 0 benchmarks
This dataset contains 2,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
This dataset contains 2,000 images taken from inside a warehouse of the Energy Company of Paraná (Copel), which directly serves more than 4 million consuming units in the Brazilian state of Paraná.
2 papers · 1 benchmark
UI5k (Mobile App User Interface Dataset)
This dataset contains 54,987 UI screenshots and the metadata from 7,748 Android applications belonging to 25 application categories Download link: https://www.dropbox.com/sh/kfkhevxykzwputb/AAAhL6ipmOg4zZn4jULmyF0a?dl=0
2 papers · 0 benchmarks
UIT-ViSFD (Vietnamese Aspect-Based Sentiment Analysis Dataset)
UIT-ViSFD is a Vietnamese Smartphone Feedback Dataset as a new benchmark corpus built based on strict annotation schemes for evaluating aspect-based sentiment analysis, consisting of 11,122 human-annotated comments for mobile e-commerce,…
2 papers · 0 benchmarks
UK Biobank participants have generously provided a very wide range of information about their health and well-being since recruitment began in 2006.
2 papers · 1 benchmark
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
UPenn-GBM (The University of Pennsylvania glioblastoma (UPenn-GBM) cohort)
This collection comprises multi-parametric magnetic resonance imaging (mpMRI) scans for de novo Glioblastoma (GBM) patients from the University of Pennsylvania Health System, coupled with patient demographics, clinical outcome (e.g.,…
2 papers · 0 benchmarks
This dataset contains vibration data recorded on a rotating drive train.
2 papers · 0 benchmarks
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
V2VBench is a comprehensive benchmark designed to evaluate video editing methods.
2 papers · 0 benchmarks
325 word images intended for font recognition, whose fonts are included in [VFR-447] (and [VFR-2420]).
2 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.