Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 52 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2449–2496 of 3,998
This dataset focuses on 50 articles about climate science, which were annotated completely by 49 students, 26 Upwork workers, 3 science and 3 journalism experts.
1 paper · 0 benchmarks
CropCOCO is a validation-only dataset of COCO val 2017 images cropped such that some keypoints annotations are outside of the image.
1 paper · 0 benchmarks
The standard evaluation protocol of Cross-View Time dataset allows for certain cameras to be shared between training and testing sets.
1 paper · 1 benchmark
CrowdSpeech is a publicly available large-scale dataset of crowdsourced audio transcriptions.
1 paper · 2 benchmarks
Official dataset of Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP.
1 paper · 0 benchmarks
The dataset contains 30 million cryptocurrency-related tweets from 10.10.2020 to 3.3.2021.
1 paper · 0 benchmarks
Abstract: Through digitization, maintaining and promoting cultural heritage is being strengthened.
1 paper · 0 benchmarks
D-OCC (Dynamic-OneCommon Corpus)
D-OCC is a large-scale dataset of 5,617 dialogues to enable fine-grained evaluation and analysis of various dialogue systems.
1 paper · 0 benchmarks
A large benchmark dataset containing 50K human judgments for 5K distinct sentence pairs in the English dative alternation.
1 paper · 0 benchmarks
DAPFAM (A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level)
Dataset DAPFAM See the accompanying paper: Ayaou et al., 2025 — “DAPFAM: A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level” (arXiv:2506.22141).
1 paper · 0 benchmarks
DARai (Daily Activity Recordings for AI and ML applications)
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings.
1 paper · 0 benchmarks
A comprehensive object-instance ReID dataset with multiple indoor object instances under varying lighting conditions.
1 paper · 0 benchmarks
A real world dataset for benchmarking global localization in complex indoor environments.
1 paper · 0 benchmarks
A ProcTHOR created synthetic dataset for benchmarking global localization in complex indoor environments.
1 paper · 0 benchmarks
DAVIS-Edit is a curated testing benchmark for video editing.
1 paper · 0 benchmarks
This dataset includes Direct Borohydride Fuel Cell (DBFC) impedance and polarization test in anode with Pd/C, Pt/C and Pd decorated Ni–Co/rGO catalysts.
1 paper · 0 benchmarks
[comment]:<> (Data for the paper "Deciphering Environmental Air Pollution with Large Scale City Data") Main Dataset citypollutiondata.csv Relevant Columns: Date: Date of the sample City: City of the sample Xmedian: Median value of the…
1 paper · 0 benchmarks
This dataset contains synthetic text data generated to train models for text generation.
1 paper · 0 benchmarks
Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV.
1 paper · 0 benchmarks
DIGITal (Digitally Generated Numerals)
Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9.
1 paper · 0 benchmarks
DMAD (Deepfake Massively Annotated Databases)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Medical VQA dataset built from the IDRiD and eOphta datasets.
1 paper · 0 benchmarks
The dataset was collected from DOTA 2 using OpenDota API via Python.
1 paper · 0 benchmarks
dataset in WWW 2019 "DPLink: User Identity Linkage via Deep Neural Network From Heterogeneous Mobility Data".
1 paper · 0 benchmarks
This dataset is a new variant of the voice cloning toolkit (VCTK) dataset: device-recorded VCTK (DR-VCTK), where the high-quality speech signals recorded in a semi-anechoic chamber using professional audio devices are played back and…
1 paper · 0 benchmarks
DRIFT (Domain-Adaptive Regression for Forest Monitoring)
The DRIFT dataset includes 25k image patches collected in five European countries sourced from aerial and nanosatellite image archives.
1 paper · 0 benchmarks
DUC 2006 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
1 paper · 0 benchmarks
The dataset comprises motion sensor data of 19 daily and sports activities each performed by 8 subjects in their own style for 5 minutes.
1 paper · 0 benchmarks
Daily load patterns (Daily load patterns of six global exchanges as recorded by vwd/Infront Financial Technology)
This data set provides fine-granular statistics on trading traffic generated by six global exchanges over the course of two days in February 2019 for a set of representative feeds and recorded by the systems of vwd Vereinigte…
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
A dataset of images obtained from DALL-E 3 for 67 countries and 10 concept classes, similar to DollarStreet images.
1 paper · 0 benchmarks
This archive contains raw data, intermediate results, statistics, and figures for the manuscript "Naïve individuals promote collective exploration in homing pigeons" Once unzipped, the folder structure will look as follow: - data/ [raw…
1 paper · 0 benchmarks
About the study This study was exploring the landscape of interpersonal conflicts during code review in following areas: - how these conflicts look like - what role do they play in software development - what are their consequences - what…
1 paper · 0 benchmarks
Data for "Image-based Backbone Reconstruction for Non-Slender Soft Robots" This dataset provides the data for the forthcoming paper "Image-based Backbone Reconstruction for Non-Slender Soft Robots".
1 paper · 0 benchmarks
Data from the "Resistance Against Manipulative AI: key factors and possible actions" article
1 paper · 0 benchmarks
Simulation data and pre-trained Graph Neural Network (GNN) models produced in [1].
1 paper · 0 benchmarks
Dataset information (e.g., google drive link) is attached in the GitHub repo: https://github.com/YY-GX/Annotated-Hands-Dataset Please find the description of the dataset in our paper: http://arxiv.org/abs/2401.15075
1 paper · 0 benchmarks
The dataset evaluates the number of vesicles observed in a Tcell that is close to a tumor cell.
1 paper · 0 benchmarks
This dataset is used for MPLP considering time windows constraints of customers and parking space.
1 paper · 0 benchmarks
Dataset for ZAugNet (Self-Supervised Z-Slice Augmentation for 3D Bio-Imaging via Knowledge Distillation)
Dataset used to train ZAugNet, a neural network for Z-slice augmentation, that encompasses a variety of shapes, textures, and microscopy techniques, as described below: Ascidian Embryos: This dataset consists of 3D confocal images of P.
1 paper · 0 benchmarks
Overview This dataset was collected during a pilot study that evaluated the virtual Cross Array Task (CAT) platform as an assessment tool for algorithmic thinking (AT) skills among K-12 students in Swiss compulsory education.
1 paper · 0 benchmarks
This dataset provides neutron and gamma-ray pulse signals for pulse shape discrimination experiments.
1 paper · 0 benchmarks
We release both the processed data and evaluation results from our own experiments, and the underlying raw data that can be used for future experiments and schemes in the domain of Zero-Interaction Security.
1 paper · 0 benchmarks
Overview of the scoping review paper corpus, sorted by their diferent intent types, categories, and subcategories.
1 paper · 0 benchmarks
The dataset is generated from the study of computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
This repository contains the dataset for the study of the computational reproducibility of Jupyter notebooks from biomedical publications.
1 paper · 0 benchmarks
Dataset outline This repository contains a novel time-series dataset for impact detection and localization on a plastic thin-plate, towards Structural Health Monitoring applications, using ceramic piezoelectric transducers (PZTs) connected…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.