Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 78 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3697–3744 of 3,998
An axial turbine is a simplest hydrulic machine which is suitable for low-head conditions.
0 papers · 0 benchmarks
Dataset description We acquired the EEG from three Laplacian derivations, 3.5 cm (center-to- center) around the electrode positions (according to International 10-20 System of Electrode Placement) C3 (FC3, C5, CP3 and C1), Cz (FCz, C1, CPz…
0 papers · 0 benchmarks
Reflectance measurements of Bidirectional Texture Functions (BTFs) Database contains both flat samples: as well as 3D geometry with texture mapped BTFs: furthermore, there are some multispectral BTFs:
0 papers · 0 benchmarks
The dataset consists of 3265 text samples corresponding to the concatenation of lines spoken by fictional characters.
0 papers · 0 benchmarks
This dataset consists of odometer or speedometer images of bike and car vehicles.
0 papers · 0 benchmarks
Annotated and original images of billboards in Japanese street scapes
0 papers · 0 benchmarks
CBCT Walnut (Cone-Beam X-Ray CT Data Collection Designed for Machine Learning)
The scans are performed using a custom-built, highly flexible X-ray CT scanner, the FleX-ray scanner, developed by XRE nvand located in the FleX-ray Lab at the Centrum Wiskunde & Informatica (CWI) in Amsterdam, Netherlands.
0 papers · 0 benchmarks
CBLPRD-330k (China-Balanced-License-Plate-Recognition-Dataset-330k)
A high-quality, balanced dataset of 330,000 images featuring various types of Chinese license plates.
0 papers · 0 benchmarks
We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory.
0 papers · 0 benchmarks
CIDII Dataset (Correct Information and Disinformation about Islamic Issues)
The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues.
0 papers · 0 benchmarks
COCO-Facet is a benchmark for attribute-focused text-to-image retrieval, comprising 9,112 queries with 100 candidate images for each.
0 papers · 0 benchmarks
Applications of unmanned aerial vehicle (UAV) in logistics, agricultural automation, urban management, and emergency response are highly dependent on oriented object detection (OOD) to enhance visual perception.
0 papers · 0 benchmarks
With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic.
0 papers · 0 benchmarks
We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource.
0 papers · 0 benchmarks
This dataset is a collection of 4,000 images of cars in multiple scenes that are ready to use for optimizing the accuracy of computer vision models.
0 papers · 0 benchmarks
ChaBuD (Change detection for Burned area Delineation)
The dataset comprises patches of size 512x512 pixels collected from Sentinel-2 L2A satellite mission.
0 papers · 0 benchmarks
This is a dataset of paraphrases created by ChatGPT.
0 papers · 0 benchmarks
Dataset Description Our dataset contains questions from a well-known software testing book Introduction to Software Testing 2nd Edition by Ammann and Offutt.
0 papers · 0 benchmarks
Key Points - Purpose: Captures crossroad navigation under cloudy weather conditions.
0 papers · 0 benchmarks
CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations.
0 papers · 0 benchmarks
Colorectal-Liver-Metastases (Colorectal-Liver-Metastases | Preoperative CT and Survival Data for Patients Undergoing Resection of Colorectal Liver Metastases)
This collection consists of DICOM images and DICOM Segmentation Objects (DSOs) for 197 patients with Colorectal Liver Metastases (CRLM).
0 papers · 0 benchmarks
A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects.
0 papers · 0 benchmarks
ContraCAT (Contrastive Coreference Analytical Templates (for Machine Translation))
Current approaches to context-aware MT rely on a set of surface heuristics to translate pronouns, which break down when translations require real reasoning.
0 papers · 0 benchmarks
This dataset is the images of corn seeds considering the top and bottom view independently (two images for one corn seed: top and bottom).
0 papers · 0 benchmarks
This dataset includes CSV files that contain IDs and sentiment scores of the tweets related to the COVID-19 pandemic.
0 papers · 0 benchmarks
The study of material corrosion is an important research area, with corrosion degradation of metallic structures causing expenses up to 4% of the global domestic product annually along with major safety risks worldwide.
0 papers · 0 benchmarks
Crowd 11 (A Dataset for Fine Grained Crowd Behaviour Analysis)
This dataset defines a total of 11 crowd motion patterns and it is composed of over 6000 video sequences with an average length of 100 frames per sequence.
0 papers · 0 benchmarks
This dataset collects transparency disclosures about the sexual exploitation of children by social media and their reports about such activity and material to the national clearinghouse, the National Center for Missing and Exploited…
0 papers · 0 benchmarks
Description: CytoImageNet is an extensive collection of microscopy images, carefully curated to aid in the development of fast and automated methods for analyzing biological data.
0 papers · 0 benchmarks
| Name | Purpose | |------|---------| | FM100P | Evaluation of the single palette sorting | | KHTP | Evaluation of the palette pair sorting | | LHSP | Evaluation of the palette similarity measurement | | Perceptual Study | Perceptual Study…
0 papers · 0 benchmarks
In bioinformatics, the issue of mutation discovery and type determination remains a significant concern.
0 papers · 0 benchmarks
Dialog System Technology Challenges 8 (DSTC) Track 2 builds on the success of DSTC 7 Track 1 (NOESIS: Noetic End-to-End Response Selection Challenge).
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.