Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 50 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2353–2400 of 3,998

COCO-OOC goes beyond standard object detection to ask the question: Which objects are out-of-context (OOC)?
1 paper · 1 benchmark
CODD (Cooperative Driving Dataset)
The Cooperative Driving dataset is a synthetic dataset generated using CARLA that contains lidar data from multiple vehicles navigating simultaneously through a diverse set of driving scenarios.
1 paper · 0 benchmarks
COFFE (COFFE: A Code Efficiency Benchmark for Code Generation)
COFFE COFFE is a Python benchmark for evaluating the time efficiency of LLM-generated code.
1 paper · 0 benchmarks
The causal reasoning dataset is generated using the Causal Reasoning in Closed Daily Activities (COLD) framework that helps evaluate large language models (LLMs) on their causal reasoning abilities within real-world, everyday activities.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CORBEL (Conveyor belt pressure signal dataset))
Dataset included measuring static tension under 2 kg load in different points of the CB and measurements in dynamic conditions.
1 paper · 1 benchmark
CORRONA CERTAIN (Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions)
CERTAIN, or the Comparative Effectiveness Registry to Study Therapies for Arthritis and Inflammatory Conditions, is designed as a prospective nested substudy under our larger RA registry.
1 paper · 0 benchmarks
These datasets were used in the paper 'Evaluation of Thematic Coherence in Microblogs' (ACL, 2021).
1 paper · 0 benchmarks
This case surveillance public use dataset has 12 elements for all COVID-19 cases shared with CDC and includes demographics, any exposure history, disease severity indicators and outcomes, presence of any underlying medical conditions and…
1 paper · 0 benchmarks
The dataset contains Tweet IDs along with the location and tweet timestamp.
1 paper · 0 benchmarks
COVID-19 Vaccine Stance Dataset (COVID-19 Vaccination Stance with (De)Motivation Classification)
The data contains CSV files with anonymized user names, tweet texts, vaccine stance, cumulative score for the vaccine stance, location, and topic information.
1 paper · 0 benchmarks
COVMis-Stance is a stance detection dataset for COVID-19 misinformation.
1 paper · 0 benchmarks
CPNet (CorresPondenceNet)
CPNet dataset has a collection of 25 categories, 2,334 models based on ShapeNetCore, which includes 1,000+ correspondence sets with 104,861 points.
1 paper · 0 benchmarks
CPSC2019 (The 2nd China Physiological Signal Challenge (CPSC 2019))
Introduction The China Physiological Signal Challenge 2019 (CPSC 2019) aims to encourage the development of algorithms for challenging QRS detection and heart rate (HR) estimation from short-term single-lead ECG recordings usually with low…
1 paper · 0 benchmarks
CPSC2020 (The 3rd China Physiological Signal Challenge 2020)
Introduction Abnormality of cardiac conduction system can induce arrhythmia.
1 paper · 0 benchmarks
CPSC2021 (The 4th China Physiological Signal Challenge 2021)
Introduction The 4th China Physiological Signal Challenge 2021 (CPSC 2021) aims to encourage the development of algorithms for searching the paroxysmal atrial fibrillation (PAF) events from dynamic ECG recordings.
1 paper · 0 benchmarks
Intrusion alert dataset captured through the Collegiate Penetration Testing Competition (CPTC) 2018.
1 paper · 0 benchmarks
This is the price data that supports the cost estimation of data center providing flexibility, for the following paper "AI-focused HPC Data Centers Can Provide More Power Grid Flexibility and at Lower Cost".
1 paper · 0 benchmarks
CRCDX (TCGA-CRC-DX)
Histological images of colorectal cancer, derived from the TCGA database
1 paper · 0 benchmarks
CRED (Crowd Reaction Estimation Dataset)
In the realm of social media, understanding and predicting post reach is a significant challenge.
1 paper · 0 benchmarks
CRSB (Context Retrieval Supervision Benchmark)
The Official dataset proposed int the paper Context Awareness Gate For Retrieval Augmented Generation
1 paper · 0 benchmarks
CSAbstruct is a new dataset of annotated computer science abstracts with sentence labels according to their rhetorical roles.
1 paper · 0 benchmarks
CSI is a criminal conversational dataset for speaker identification built from the CSI television show.
1 paper · 0 benchmarks
The dataset contains gold-standard summary labels for 39 "CSI: Crime Scene Investigation" episodes from seasons 1-5.
1 paper · 0 benchmarks
CSPRD (Chinese Stock Policy Retrieval Dataset)
The Chinese Stock Policy Retrieval Dataset (CSPRD) contains a Chinese policy corpus of 10,002 articles and 709 prospectus examples from 545 companies listed on China’s Science and Technology Innovation Board (STAR Market).
1 paper · 0 benchmarks
Over 20,000 annotated synthetic images and web-scraped images of bicyclists with bounding box annotations in Pascal VOC format.
1 paper · 1 benchmark
CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents).
1 paper · 0 benchmarks
This dataset contains samples of CTI (Cyber Threat Intelligence) data in natural language, labeled with the corresponding adversarial techniques from the MITRE ATT&CK framework.
1 paper · 0 benchmarks
This dataset includes 720 directional B-format RIRs, i.e.
1 paper · 0 benchmarks
CUHK-QA is a dataset for natural language-based person search using iterative questioning.
1 paper · 0 benchmarks
CURE (A dataset for Clinical Understanding & Retrieval Evaluation)
CURE is a retrieval dataset with a monolingual and two cross-lingual conditions, with splits spanning ten medical domains.
1 paper · 0 benchmarks
CVE (Common Vulnerabilities and Exposures)
CVE stands for Common Vulnerabilities and Exposures.
1 paper · 0 benchmarks
CVR (Congressional Voting Records Data Set)
This data set includes votes for each of the U.S.
1 paper · 1 benchmark
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
This dataset comprises video files (converted into tif format) that depict glomerular activation in mice.
1 paper · 0 benchmarks
A collaborative effort between researchers at the Vascular Imaging Lab located at the University of Calgary and the Medical Image Computing Lab located at the University of Campinas (UNICAMP) originated the Calgary Campinas public brain…
1 paper · 0 benchmarks
https://zenodo.org/records/15301636
1 paper · 0 benchmarks
The CapMIT1003 database contains captions and clicks collected for images from the MIT1003 database, for which reference eye scanpath are available.
1 paper · 1 benchmark
Capriccio (Sentiment Analysis + Data Drift)
Capriccio is a sentiment classification dataset on tweets that simulates data drift.
1 paper · 0 benchmarks
Car_Price_Prediction (Second_Hand-Car_Price_Prediction)
In this dataset we added [Company Name, Car Model, Car Type, Fuel Type, Transmission, Engine (cc), Mileage, Kmsdriven, Buyers, Horsepower (kw), Year Price (Lakhs)]
1 paper · 1 benchmark
A dataset of games played in the card game "Cards Against Humanity" (CAH), by human players, derived from the online CAH labs.
1 paper · 0 benchmarks
A large-scale benchmark dataset involving well-labelled datasets to employ the state-of-the-art machine intelligence technologies for map text annotation recognition, map scene classification, map super-resolution reconstruction, and map…
1 paper · 0 benchmarks
Caselaw4 is a dataset of 350k common law judicial decisions from the U.S.
1 paper · 0 benchmarks
Casino Reviews (Online reviews of North American Casinos from Google Reviews)
This dataset contain online reviews gathered from google reviews written by north american casino users.
1 paper · 0 benchmarks
Cattle data set, which was introduced in a paper.
1 paper · 0 benchmarks
SyntaxGym, adapted for interventional interpretability.
1 paper · 1 benchmark
Classifying all cells in an organ is a relevant and difficult problem from plant developmental biology.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.