Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 71 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3361–3408 of 3,998

TACM12K (Table-ACM12K)
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset.
1 paper · 1 benchmark
TACO-BAAI (Topics in Algorithmic Code generation dataset)
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field.
1 paper · 1 benchmark
TADAC (Text Annotated Distortion, Appearance and Content Dataset)
We have developed a systematic method for constructing large text annotated image databases designed for exploiting vision-language modeling for image quality assessment and present the Text Annotated Distortion, Appearance and Content…
1 paper · 0 benchmarks
TAMPAR is a real-world dataset of parcel photos for tampering detection with annotations in COCO format.
1 paper · 0 benchmarks
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
1 paper · 0 benchmarks
TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
TBBR Raw (Hyperspectral (RGB + Thermal) drone images of Karlsruhe, Germany)
This dataset contains the raw images for the dataset of Thermal Bridges on Building Rooftops (TBBR) dataset.
1 paper · 0 benchmarks
TCB-DS (Toxigenic Cyanobacteria Dataset)
The TCB-DS dataset is a specialized collection of microscopic images focusing on the automatic recognition of cyanobacteria genera.
1 paper · 0 benchmarks
This dataset is a collection of paired wireless signal data and corresponding image ground truth specifically designed for underground object sensing and image reconstruction.
1 paper · 0 benchmarks
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
TF1-EN-3M (klusai/ds-tf1-en-3m)
TF1-EN-3M: Three Million Synthetic Moral Fables for Open Language Models TF1-EN-3M is a large-scale synthetic dataset of 3,000,000 English-language moral fables, generated by instruction-tuned language models with no more than 8 billion…
1 paper · 0 benchmarks
TGB (Temporal Graph Benchmark)
TGB is a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust machine learning evaluation on temporal graphs.
1 paper · 0 benchmarks
THEOStereo is a dataset providing synthetic stereo image pairs and their corresponding scene depth and will be published along with [1].
1 paper · 0 benchmarks
THGP (Temporal Hands Guns and Phones Dataset)
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing).
1 paper · 0 benchmarks
THRED (Two-Hop Relation Extraction Dataset)
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
1 paper · 0 benchmarks
TIC Powerline Dataset (Tower Insulator Conductors Dataset for high voltage powerline inspection)
The TIC Dataset consists of 2056 images (512x640) of transmission line network footage in Greece (Northeast Attica) and annotations of three object classes, i.e.
1 paper · 0 benchmarks
TILT corpus (GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit)
A corpus of GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit (TILT).
1 paper · 0 benchmarks
TLF2K (Table-LastFm2K)
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset.
1 paper · 1 benchmark
TML1M (Table-MovieLens1M)
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset.
1 paper · 1 benchmark
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
1 paper · 1 benchmark
TREx-2p is a dataset to probe whether a pretrained LM possesses “indirect” 2-hop knowledge.
1 paper · 0 benchmarks
TRR360D is based on the ICDAR2019MTD modern table detection dataset, it refers to the annotation format of the DOTA dataset.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
100 samples each of synthetic speech generated by 9 moderns TTS systems.
1 paper · 0 benchmarks
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks
TVPReid (Text-to-Video Person Re-identification)
The TVPReid dataset contains 6559 pedestrian videos, each of which is annotated with two text descriptions, for a total of 13118 descriptions.
1 paper · 0 benchmarks
TVRecap a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved.
1 paper · 4 benchmarks
The TWT16 dataset contains ~30k conversations in Twitter, collected from January to June 2016.
1 paper · 0 benchmarks
TXL-PBC dataset (a freely accessible labeled peripheral blood cell dataset)
The TXL-PBC Dataset is a comprehensive collection of re-annotated and integrated cell images from multiple cell datasets.
1 paper · 0 benchmarks
This data set contains real-world table tennis ball trajectories recorded with our custom developed table tennis ball launcher AIMY.
1 paper · 0 benchmarks
Taskography (PDDLGym Taskography)
PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.
1 paper · 0 benchmarks
This dataset consists of RGB-D images captured using 12 Intel RealSense cameras.
1 paper · 0 benchmarks
TeleSim (TeleSim: A Network-Aware Testbed and Benchmark Dataset for Telerobotic Applications)
TeleSim is a network-aware hardware-in-the-loop dataset designed to evaluate the performance of telerobotic systems under varying network conditions.
1 paper · 0 benchmarks
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
1 paper · 0 benchmarks
We introduce TextAtlas5M, a dataset specifically designed for training and evaluating multimodal generation models on dense-text image generation.
1 paper · 0 benchmarks
Text present in images are not merely strings, they provide useful cues about the image.
1 paper · 0 benchmarks
TextWorld KG is a dynamic Knowledge Graph (KG) extraction dataset.
1 paper · 0 benchmarks
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
Using the Experience-Sampling Method (ESM), participants are asked to report TV consumption multiple times each day for a five week period.
1 paper · 0 benchmarks
The EMBO SourceData-NLP dataset (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
Hyperspectral Imaging, employed in satellites for space remote sensing, like HYPSO-1, faces constraints due to few labeled data sets, affecting the training of AI models demanding these ground-truth annotations.
1 paper · 0 benchmarks
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction).
1 paper · 0 benchmarks
The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network.
1 paper · 0 benchmarks
This includes all data from the ACM IMC 2018 paper "The Rise of Certificate Transparency and Its Implications on the Internet Ecosystem".
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.