Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 71 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3361–3408 of 3,998
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
TACO-BAAI (Topics in Algorithmic Code generation dataset)
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field.
1 paper · 1 benchmark
TADAC (Text Annotated Distortion, Appearance and Content Dataset)
We have developed a systematic method for constructing large text annotated image databases designed for exploiting vision-language modeling for image quality assessment and present the Text Annotated Distortion, Appearance and Content…
1 paper · 0 benchmarks
TAMPAR is a real-world dataset of parcel photos for tampering detection with annotations in COCO format.
1 paper · 0 benchmarks
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
1 paper · 0 benchmarks
TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
TBBR Raw (Hyperspectral (RGB + Thermal) drone images of Karlsruhe, Germany)
This dataset contains the raw images for the dataset of Thermal Bridges on Building Rooftops (TBBR) dataset.
1 paper · 0 benchmarks
TCB-DS (Toxigenic Cyanobacteria Dataset)
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera.
1 paper · 0 benchmarks
This dataset is a collection of paired wireless signal data and corresponding image ground truth specifically designed for underground object sensing and image reconstruction.
1 paper · 0 benchmarks
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
TF1-EN-3M: Three Million Synthetic Moral Fables for Open Language Models TF1-EN-3M is a large-scale synthetic dataset of 3,000,000 English-language moral fables, generated by instruction-tuned language models with no more than 8 billion…
1 paper · 0 benchmarks
TGB (Temporal Graph Benchmark)
TGB is a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust machine learning evaluation on temporal graphs.
1 paper · 0 benchmarks
THEOStereo is a dataset providing synthetic stereo image pairs and their corresponding scene depth and will be published along with [1].
1 paper · 0 benchmarks
THGP (Temporal Hands Guns and Phones Dataset)
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing).
1 paper · 0 benchmarks
THRED (Two-Hop Relation Extraction Dataset)
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
1 paper · 0 benchmarks
The TIC Dataset consists of 2056 images (512x640) of transmission line network footage in Greece (Northeast Attica) and annotations of three object classes, i.e.
1 paper · 0 benchmarks
TILT corpus (GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit)
A corpus of GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit (TILT).
1 paper · 0 benchmarks
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset.
1 paper · 1 benchmark
TML1M (Table-MovieLens1M)
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset.
1 paper · 1 benchmark
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
1 paper · 1 benchmark
TREx-2p is a dataset to probe whether a pretrained LM possesses “indirect” 2-hop knowledge.
1 paper · 0 benchmarks
TRR360D is based on the ICDAR2019MTD modern table detection dataset, it refers to the annotation format of the DOTA dataset.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
100 samples each of synthetic speech generated by 9 moderns TTS systems.
1 paper · 0 benchmarks
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks
TVPReid (Text-to-Video Person Re-identification)
The TVPReid dataset contains 6559 pedestrian videos, each of which is annotated with two text descriptions, for a total of 13118 descriptions.
1 paper · 0 benchmarks
TVRecap a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved.
1 paper · 4 benchmarks
The TWT16 dataset contains ~30k conversations in Twitter, collected from January to June 2016.
1 paper · 0 benchmarks
TXL-PBC dataset (a freely accessible labeled peripheral blood cell dataset)
The TXL-PBC Dataset is a comprehensive collection of re-annotated and integrated cell images from multiple cell datasets.
1 paper · 0 benchmarks
This data set contains real-world table tennis ball trajectories recorded with our custom developed table tennis ball launcher AIMY.
1 paper · 0 benchmarks
PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.
1 paper · 0 benchmarks
This dataset consists of RGB-D images captured using 12 Intel RealSense cameras.
1 paper · 0 benchmarks
TeleSim (TeleSim: A Network-Aware Testbed and Benchmark Dataset for Telerobotic Applications)
TeleSim is a network-aware hardware-in-the-loop dataset designed to evaluate the performance of telerobotic systems under varying network conditions.
1 paper · 0 benchmarks
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
1 paper · 0 benchmarks
Green family of datasets for emergent communications on relations.
1 paper · 0 benchmarks
We introduce TextAtlas5M, a dataset specifically designed for training and evaluating multimodal generation models on dense-text image generation.
1 paper · 0 benchmarks
Text present in images are not merely strings, they provide useful cues about the image.
1 paper · 0 benchmarks
TextWorld KG is a dynamic Knowledge Graph (KG) extraction dataset.
1 paper · 0 benchmarks
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
Using the Experience-Sampling Method (ESM), participants are asked to report TV consumption multiple times each day for a five week period.
1 paper · 0 benchmarks
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
Hyperspectral Imaging, employed in satellites for space remote sensing, like HYPSO-1, faces constraints due to few labeled data sets, affecting the training of AI models demanding these ground-truth annotations.
1 paper · 0 benchmarks
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction).
1 paper · 0 benchmarks
The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network.
1 paper · 0 benchmarks
This includes all data from the ACM IMC 2018 paper "The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem".
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.