Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 36 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 1681–1728 of 3,998

TTStroke-21 ME21 (TTStroke-21 for MediaEval 2021)
This task offers researchers an opportunity to test their fine-grained classification methods for detecting and recognizing strokes in table tennis videos.
3 papers · 2 benchmarks
TTStroke-21 ME22 (TTStroke-21 for MediaEval 2022)
TTStroke-21 for MediaEval 2022.
3 papers · 2 benchmarks
TaL Corpus (The Tongue and Lips Corpus)
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
A novel dataset of document-grounded task-based dialogues, where an Information Giver (IG) provides instructions (by consulting a document) to an Information Follower (IF), so that the latter can successfully complete the task.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
It is a freely available resource for research on handling negation and uncertainty in biomedical texts .
3 papers · 5 benchmarks
The Little Prince (The Little Prince Corpus)
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
Tiny ImageNetv2 is a subset of the ImageNetV2 (matched frequency) dataset by Recht et al.
3 papers · 0 benchmarks
The ToolLens dataset consists of 18,770 concise yet intentionally multifaceted queries, each associated with 1 to 3 verified tools out of a total of 464, designed to better mimic real-world user interactions.
3 papers · 1 benchmark
TriBERT dataset consists of 12,049 training, 2,527 validation and 2,560 test Human-Machine collaborative texts.
3 papers · 1 benchmark
TriageSQL is a cross-domain text-to-SQL question intention classification benchmark that requires models to distinguish four types of unanswerable questions from answerable questions.
3 papers · 0 benchmarks
Dataset of restaurant reviews from TripAdvisor that includes images and texts uploaded in reviews by users.
3 papers · 0 benchmarks
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
UAVA (UAV Assistant)
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
The UFPR-Periocular dataset has 16,830 images of both eyes (33,660 cropped images of each eye) from 1,122 subjects (2,244 classes).
3 papers · 0 benchmarks
UIIS10K (General Underwater Image Instance Segmentation dataset 10K)
We propose a large-scale underwater instance segmentation dataset, UIIS10K, which includes 10,048 images with pixel-level annotations for 10 categories.
3 papers · 0 benchmarks
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
VID Dataset (The Visual-Inertial-Dynamical Multirotor Dataset)
The Visual-Inertial-Dynamical (VID) dataset not only focuses on traditional six degrees of freedom (6-DOF) pose estimation, but also provides dynamical characteristics of the flight platform for external force perception or dynamics-aided…
3 papers · 0 benchmarks
Vehicle-Rear is a novel dataset for vehicle identification that contains more than three hours of high-resolution videos, with accurate information about the make, model, color and year of nearly 3,000 vehicles, in addition to the position…
3 papers · 0 benchmarks
a vessel dataset using 85 videos.
3 papers · 1 benchmark
Video Localized Narratives is a new form of multimodal video annotations connecting vision and language.
3 papers · 0 benchmarks
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
WFDD (Woven Fabric Defect Detection)
WFDD is a dataset for benchmarking anomaly detection methods with a focus on textile inspection.
3 papers · 1 benchmark
In our benchmark WHYSHIFT, we explore distribution shifts on 5 real-world tabular datasets from the economic and traffic sectors with natural spatiotemporal distribution shifts.We only pick 7 typical settings out of 22 settings and select…
3 papers · 0 benchmarks
Test-driven benchmark to challenge LLMs to write JavaScript React application GitHub Script
3 papers · 1 benchmark
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
Wiki-Convert is a 900,000+ sentences dataset of precise number annotations from English Wikipedia.
3 papers · 0 benchmarks
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
A multilingual dataset for the task of multilingual claim span identification.
3 papers · 0 benchmarks
We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection.
3 papers · 0 benchmarks
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
Youtbean is a dataset created from closed captions of YouTube product review videos.
3 papers · 0 benchmarks
A new English language dataset structured for task-oriented evaluation on unseen tasks.
3 papers · 0 benchmarks
This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mining methods.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
fluocells (Fluorescent Neuronal Cells)
By releasing this dataset, we aim at providing a new testbed for computer vision techniques using Deep Learning.
3 papers · 0 benchmarks
6981 SAT-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
3 papers · 0 benchmarks
A large, crowd-sourced dataset for the Native Language Identification (NLI) task.
3 papers · 1 benchmark
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
collected by one VLP-16 in a small vehicle (1m x 1m)
3 papers · 1 benchmark
2D site-percolation threshold (Daniel García Solla)
The dataset is a .h5 file comprised of entries with keys of the form (n,m), denoting the dimensions of the system matrix on which the simulations have been performed.
2 papers · 0 benchmarks
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
3D FRONT HUMAN is a dataset that extends the large-scale synthetic scene dataset 3D-FRONT.
2 papers · 0 benchmarks
3D-Point Cloud dataset of various geometrical terrains (3D-Point Cloud dataset of various geometrical terrains in urban environments recorded during human locomotion)
Depth vision has been recently used in many locomotion devices with the objective to ease the life of disabled people toward reaching more ecological lifestyle.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.