Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 46 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2161–2208 of 3,998
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
This dataset is meant to be used to develop models for next-day fire hazard forecasting in Greece.
1 paper · 0 benchmarks
A curated evaluation dataset for end-to-end Relation Extraction of relationships between organisms and natural-products.
1 paper · 0 benchmarks
We present a dataset of dialogs in which journalists of The Guardian replied to reader comments and identify the reasons why.
1 paper · 0 benchmarks
This dataset contains 9 different seafood types collected from a supermarket in Izmir, Turkey for a university-industry collaboration project at Izmir University of Economics, and this work was published in ASYU 2020.
1 paper · 0 benchmarks
This dataset contains data of 125 1-hour simulations of ship motion during various sea states performing random maneuvers in 4 degrees of freedom (surge-sway-yaw-roll).
1 paper · 0 benchmarks
This dataset contains a collection of papers retrieved by using a PRISMA systematic review of Open Data and Public Domain data in Agriculture.
1 paper · 0 benchmarks
A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces.
1 paper · 0 benchmarks
This dataset contains a collection of 131 X-ray CT scans of pieces of modeling clay (Play-Doh) with various numbers of stones inserted, retrieved in the FleX-ray lab at CWI.
1 paper · 0 benchmarks
This dataset contains a collection of 235800 X-ray projections of 131 pieces of modeling clay (Play-Doh) with various numbers of stones inserted.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Neonatal seizures are a common emergency in the neonatal intensive care unit (NICU).
1 paper · 0 benchmarks
Dataset for A probabilistic forecast methodology for volatile electricity prices in the Australian National Electricity Market
1 paper · 0 benchmarks
This dataset is based on FB15k237 and a pre-trained language-model-based KGE.
1 paper · 0 benchmarks
This dataset is based on WN18RR and a pre-trained language-model-based KGE.
1 paper · 0 benchmarks
A2Dre (Subset of A2D Sentences which are not trivial)
We obtain A2Dre by selecting only instances that were labeled as non-trivial, which are 433 REs from 190 videos.
1 paper · 1 benchmark
A2Dre+ (Extension of A2D sentences where trivial cases where filtered)
A2Dre is a subset from the A2D test set including $433$~\textit{non-trivial} REs.
1 paper · 0 benchmarks
AAAC (Artificial Argument Analysis Corpus)
DeepA2 is a modular framework for deep argument analysis.
1 paper · 0 benchmarks
ABCD Study (Adolescent Brain Cognitive Development)
The ABCD Study is a prospective longitudinal study starting at the ages of 9-10 and following participants for 10 years.
1 paper · 0 benchmarks
AC-Bench: A Benchmark for Actual Causality Reasoning Dataset Description AC-Bench is designed to evaluate the actual causality (AC) reasoning capabilities of large language models (LLMs).
1 paper · 0 benchmarks
A benchmark environment based on the datasets "Adult" and "Names", which allows researchers to test how well their language model can abide by pre-defined access rights rules.
1 paper · 0 benchmarks
ACCORD CSQA is an extension of the popular CommonsenseQA (CSQA) dataset using ACCORD, a scalable framework for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop…
1 paper · 0 benchmarks
ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)
This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed.
1 paper · 0 benchmarks
This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs.
1 paper · 0 benchmarks
Dataset Card for the ACR Appropriateness Criteria Corpus This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help…
1 paper · 0 benchmarks
This is a comprehensive dataset of human arm motion during Activities of Daily Living (ADL).
1 paper · 0 benchmarks
Replication Material This document contains the necessary materials and instructions to replicate the findings presented in our paper.
1 paper · 0 benchmarks
The AI-GA (Artificial Intelligence Generated Abstracts) dataset is a collection of abstracts and titles, with half of the abstracts being AI-generated and the other half being original.
1 paper · 0 benchmarks
This project contains instructions and codes to reconstruct a dataset for the development and evaluation of forensic tools for detecting machine-generated text in social media.
1 paper · 0 benchmarks
We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients.
1 paper · 0 benchmarks
AIDERV2 (Aerial Image Dataset for Emergency Response Applications (version 2))
The dataset contains aerial images containing three commonly occurring natural disasters earthquake/collapsed buildings, flood, wildfire/fire, and a normal class; do not reflect any disaster.
1 paper · 1 benchmark
AIROGS (Rotterdam EyePACS AIROGS)
The Rotterdam EyePACS AIROGS dataset (in full, so including train and test) contains 113,893 color fundus images from 60,357 subjects and approximately 500 different sites with a heterogeneous ethnicity.
1 paper · 0 benchmarks
AISECKG (AISecKG: Knowledge Graph Dataset for Cybersecurity Education)
Cybersecurity education is exceptionally challenging as it involves learning the complex attacks; tools and developing critical problem-solving skills to defend the systems.
1 paper · 0 benchmarks
Video samples recorded in the field using the Azure Kinect DK.
1 paper · 0 benchmarks
ALLO (Anomaly Localization in Lunar Orbit)
ALLO is an anomaly detection and localization dataset for space stations in lunar orbit.
1 paper · 0 benchmarks
ANCHOLIK-NER (ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition)
We developed ANCHOLIK-NER, a Bangla Regional Named Entity Recognition dataset focusing on the Sylhet, Chittagong, and Barishal dialects.
1 paper · 0 benchmarks
AODRaw (Adverse condition Object Detection with RAW images)
We introduce the AODRaw dataset, which offers 7,785 high-resolution real RAW images with 135,601 annotated instances spanning 62 categories, capturing a broad range of indoor and outdoor scenes under 9 distinct light and weather conditions.
1 paper · 1 benchmark
AP (Adversarial Paraphrase)
This is a paraphrasing dataset created using the adversarial paradigm.
1 paper · 1 benchmark
A database of 56 high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits.
1 paper · 0 benchmarks
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
ARCT (Argument Reasoning Comprehension Task)
Freely licensed dataset with warrants for 2k authentic arguments from news comments.
1 paper · 0 benchmarks
This page contains ARINC 429 message data recorded from the hardware-in-a-loop simulator.
1 paper · 0 benchmarks
ASCAD (ANSSI SCA Database) is a set of databases that aims at providing a benchmarking reference for the SCA community: the purpose is to have something similar to the MNIST database that the Machine Learning community has been using for…
1 paper · 0 benchmarks
ASCAD database version 2.
1 paper · 0 benchmarks
ASLG-PC12 (English-ASL Gloss Parallel Corpus 2012)
An artificial corpus built using grammatical dependencies rules due to the lack of resources for Sign Language.
1 paper · 1 benchmark
ASOS Data (Automated Surface/Weather Observing Systems (ASOS/AWOS) Data)
The Automated Surface Observing Systems (ASOS) program is a joint effort of the National Weather Service (NWS), the Federal Aviation Administration (FAA), and the Department of Defense (DOD).
1 paper · 1 benchmark
ASRD (Anime Style Recognition Dataset)
A well-labeled challenging dataset, to facilitate the research on style recognition on anime images by collecting images from 190 anime and cartoon works covering 93 years from 13 countries and regions, 2D and 3D work into consideration…
1 paper · 0 benchmarks
ASyMOB (ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark)
ASyMOB (pronounced Asimov, in tribute to the renowned author), is a novel assessment framework focused exclusively on symbolic manipulation, featuring 17,092 unique math challenges, organized by similarity and complexity.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.