Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 36 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1681–1728 of 3,998
This task offers researchers an opportunity to test their fine-grained classification methods for detecting and recognizing strokes in table tennis videos.
3 papers · 2 benchmarks
TTStroke-21 for MediaEval 2022.
3 papers · 2 benchmarks
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
A novel dataset of document-grounded task-based dialogues, where an Information Giver (IG) provides instructions (by consulting a document) to an Information Follower (IF), so that the latter can successfully complete the task.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
It is a freely available resource for research on handling negation and uncertainty in biomedical texts .
3 papers · 5 benchmarks
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
Tiny ImageNetv2 is a subset of the ImageNetV2 (matched frequency) dataset by Recht et al.
3 papers · 0 benchmarks
The ToolLens dataset consists of 18,770 concise yet intentionally multifaceted queries, each associated with 1 to 3 verified tools out of a total of 464, designed to better mimic real-world user interactions.
3 papers · 1 benchmark
TriBERT dataset consists of 12,049 training, 2,527 validation and 2,560 test Human-Machine collaborative texts.
3 papers · 1 benchmark
TriageSQL is a cross-domain text-to-SQL question intention classification benchmark that requires models to distinguish four types of unanswerable questions from answerable questions.
3 papers · 0 benchmarks
Dataset of restaurant reviews from TripAdvisor that includes images and texts uploaded in reviews by users.
3 papers · 0 benchmarks
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
The UFPR-Periocular dataset has 16,830 images of both eyes (33,660 cropped images of each eye) from 1,122 subjects (2,244 classes).
3 papers · 0 benchmarks
UIIS10K (General Underwater Image Instance Segmentation dataset 10K)
We propose a large-scale underwater instance segmentation dataset, UIIS10K, which includes 10,048 images with pixel-level annotations for 10 categories.
3 papers · 0 benchmarks
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
VID Dataset (The Visual-Inertial-Dynamical Multirotor Dataset)
The Visual-Inertial-Dynamical (VID) dataset not only focuses on traditional six degrees of freedom (6-DOF) pose estimation, but also provides dynamical characteristics of the flight platform for external force perception or dynamics-aided…
3 papers · 0 benchmarks
Vehicle-Rear is a novel dataset for vehicle identification that contains more than three hours of high-resolution videos, with accurate information about the make, model, color and year of nearly 3,000 vehicles, in addition to the position…
3 papers · 0 benchmarks
a vessel dataset using 85 videos.
3 papers · 1 benchmark
Video Localized Narratives is a new form of multimodal video annotations connecting vision and language.
3 papers · 0 benchmarks
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
WFDD (Woven Fabric Defect Detection)
WFDD is a dataset for benchmarking anomaly detection methods with a focus on textile inspection.
3 papers · 1 benchmark
In our benchmark WHYSHIFT, we explore distribution shifts on 5 real-world tabular datasets from the economic and traffic sectors with natural spatiotemporal distribution shifts.We only pick 7 typical settings out of 22 settings and select…
3 papers · 0 benchmarks
Test-driven benchmark to challenge LLMs to write JavaScript React application GitHub Script
3 papers · 1 benchmark
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
Wiki-Convert is a 900,000+ sentences dataset of precise number annotations from English Wikipedia.
3 papers · 0 benchmarks
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
A multilingual dataset for the task of multilingual claim span identification.
3 papers · 0 benchmarks
We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection.
3 papers · 0 benchmarks
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
Youtbean is a dataset created from closed captions of YouTube product review videos.
3 papers · 0 benchmarks
A new English language dataset structured for task-oriented evaluation on unseen tasks.
3 papers · 0 benchmarks
This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mining methods.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
By releasing this dataset, we aim at providing a new testbed for computer vision techniques using Deep Learning.
3 papers · 0 benchmarks
6981 SAT-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.
3 papers · 0 benchmarks
A large, crowd-sourced dataset for the Native Language Identification (NLI) task.
3 papers · 1 benchmark
This package provides utilities for generation, filtering, solving, visualizing, and processing of mazes for training ML systems.
3 papers · 0 benchmarks
collected by one VLP-16 in a small vehicle (1m x 1m)
3 papers · 1 benchmark
The dataset is a .h5 file comprised of entries with keys of the form (n,m), denoting the dimensions of the system matrix on which the simulations have been performed.
2 papers · 0 benchmarks
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
3D FRONT HUMAN is a dataset that extends the large-scale synthetic scene dataset 3D-FRONT.
2 papers · 0 benchmarks
Depth vision has been recently used in many locomotion devices with the objective to ease the life of disabled people toward reaching more ecological lifestyle.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.