Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 70 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3313–3360 of 3,998

Songdo Vision (Songdo Vision: Vehicle Annotations from High-Altitude BeV Drone Imagery in a Smart City)
The Songdo Vision dataset provides high-resolution (4K, 3840×2160 pixels) RGB images annotated with categorized axis-aligned bounding boxes (BBs) for vehicle detection from a high-altitude bird’s-eye view (BeV) perspective.
1 paper · 1 benchmark
We collect a dataset of 805 clean videos that show the action of pouring water in a container.
1 paper · 1 benchmark
Ensemble Tagger Training and Testing Set This data includes two files: The training set used to create the SCANL Ensemble tagger [1] and the "unseen" testing set that includes words from systems that are not available in the training set.
1 paper · 0 benchmarks
Source code (Source code underlying the publication: Topology-Based Reconstruction Prevention for Decentralised Learning)
MATLAB code to reproduce results presented in the paper "Topology-Based Reconstruction Prevention for Decentralised Learning".
1 paper · 0 benchmarks
Scene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks.
1 paper · 0 benchmarks
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI).
1 paper · 0 benchmarks
Spatial Commonsense Graph Dataset (Spatial Commonsense Graph for Object Localisation in Partial Scenes)
Dataset built from partial reconstructions of real-world indoor scenes using RGB-D sequences from ScanNet, aimed at estimating the unknown position of an object (e.g.
1 paper · 0 benchmarks
Spectral Detection and Analysis Based Paper(SDAAP) dataset is the first open-source textual knowledge dataset for spectral analysis and detection and contains annotated literature data as well as corresponding knowledge instruction data,…
1 paper · 0 benchmarks
Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Spike-X4K (Spike-X4K Dataset)
Overview The Spike-X4K Dataset is a high-resolution image reconstruction resource tailored for the latest advancements in spike camera technology.
1 paper · 1 benchmark
Spot the Difference Corpus is a corpus of task-oriented spontaneous dialogues which contains 54 interactions between pairs of subjects interacting to find differences in two very similar scenes.
1 paper · 0 benchmarks
This dataset contains over 47,000 LEGO structures of over 28,000 unique 3D objects accompanied by detailed captions.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
The datasets used in our WACV paper High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer.
1 paper · 0 benchmarks
3D confocal stacks with corresponding 2D Light-field microscope images Confocal: -Single volume dimension: 1287x1287x64.
1 paper · 0 benchmarks
This data contains throughput, RTT, power consumption and speed data measured with Starlink during mobility setups.
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
When arriving at each state, each observation token gets a coin toss to see whether it will appear in the output observation string.
1 paper · 0 benchmarks
This dataset is a "part I" extension of the "Engineered cardiac microbundle time-lapse microscopy image dataset" and contains 732 experimental time-lapse image sequences of beating hiPSC-based cardiac microbundles using microbundle strain…
1 paper · 0 benchmarks
This dataset consists of EEG (Electroencephalogram) recordings collected from students at our college during an educational experiment.
1 paper · 0 benchmarks
The StudyAbroadGPT-Dataset is a collection of conversational data focused on university application requirements for various programs, including MBA, MS in Computer Science, Data Science, and Bachelor of Medicine.
1 paper · 0 benchmarks
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization.
1 paper · 0 benchmarks
Subjective Perception of Active Noise Reduction (SPANR) (Replication Data for: Anti-noise window: subjective perception of active noise reduction and effect of informational masking)
This repository contains replication data to the paper titled: "Anti-noise window: subjective perception of active noise reduction and effect of informational masking"
1 paper · 0 benchmarks
A large dataset of around 40000 Reddit posts was collected from r/suicidewatch and other non-suicidal subreddits.
1 paper · 0 benchmarks
The dataset contains 140 paragraphs from climate change reports with associated aspect-based (i.e.
1 paper · 0 benchmarks
Super-CLEVR-3D is a visual question answering (VQA) dataset where the questions are about the explicit 3D configuration of the objects from images (i.e.
1 paper · 0 benchmarks
This deposit is supplementary material to "Machine Learning Applications in Archaeological Practices: A Review".
1 paper · 0 benchmarks
Supplementary Material (Annotation Table of Review)
The file contains an annotated list of papers that are included in the literature survey.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Selected from the databricks/databricks-dolly-15k dataset - Generation Approach: Iterative evolution of instructions using a conversational…
1 paper · 0 benchmarks
Overview The LaMini Dataset is an instruction dataset generated using h2ogpt-gm-oasst1-en-2048-falcon-40b-v2.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Derived from the FLAN-v2 Collection.
1 paper · 0 benchmarks
There are 9,321 survey papers with high quality included in the SurvayBank in the domain of computer science.
1 paper · 0 benchmarks
The dataset contains cardiovascular medical records taken from 299 patients.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
To explore the nascent area of sustainable venture capital, a review of related research was conducted and social entrepreneurs & investors interviewed to construct a questionnaire assessing the interests and intentions of current & future…
1 paper · 0 benchmarks
Switchboard Dialog Act Corpus
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SynthBRSet (Synthetic Bike Rotation Dataset)
3D Computer Graphics is leveraged to generate a large and diverse dataset for training bike rotation estimators in bike parking assessment.
1 paper · 0 benchmarks
A dataset consisting of high-quality, synthetic chest X-rays from the CheXGenBench-benchmark leading model, Sana (0.6B).
1 paper · 0 benchmarks
A public open dataset of synthetic chest X-ray images of COVID-19.
1 paper · 0 benchmarks
This dataset is originally created for the Knowledge Graph Reasoning Challenge for Social Issues (KGRC4SI) Video data that simulates daily life actions in a virtual space from Scenario Data.
1 paper · 0 benchmarks
About Dataset The File contains 3D point cloud data of a Fabricate plant with 10 sequences.
1 paper · 0 benchmarks
Overview: This collection contains three synthetic datasets produced by gpt-4o-mini for sentiment analysis and PDT (Product Desirability Toolkit) testing.
1 paper · 0 benchmarks
Synthetic dataset comprising three different environments for multi-camera dynamic novel view synthesis for soccer.
1 paper · 0 benchmarks
Synthetic Speech Attribution Dataset.
1 paper · 0 benchmarks
Synthetic visual inspection data of structural elements in bridges.
1 paper · 0 benchmarks
We used the following procedure.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.