Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 58 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2737–2784 of 3,998

ILIAS (ILIAS: Instance-Level Image retrieval At Scale)
ILIAS is a large-scale test dataset for evaluation on Instance-Level Image retrieval At Scale.
1 paper · 0 benchmarks
A large dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
A small dataset from the Inductive Link Prediction Challenge 2022.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
IMCPT-SparseGM dataset is a new visual graph matching benchmark addressing partial matching and graphs with larger sizes, based on the novel stereo benchmark Image Matching Challenge PhotoTourism (IMC-PT) 2020.
1 paper · 1 benchmark
IMPACT Patent (A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents)
It is a large-scale multimodal patent dataset with detailed captions for design patent figures.
1 paper · 1 benchmark
INSANE Cross-Domain UAV Data Set (Cross-Domain UAV Data Sets with Increased Number of Sensors for developing Advanced and Novel Estimators)
This data set contains over 600GB of multimodal data from a Mars analog mission, including accurate 6DoF outdoor ground truth, indoor-outdoor transitions with continuous cross-domain ground truth, and indoor data with Optitrack…
1 paper · 0 benchmarks
INSTANCE (the Italian seismic dataset for machine learning)
INSTANCE is a data collection of more than 1.3 million seismic waveforms originating from a selection of about 54,000 earthquakes occurred since 2005 in Italy and surrounding regions and seismic noise recordings randomly extracted from…
1 paper · 0 benchmarks
IRV2V (IRregular V2V Dataset)
To facilitate research on asynchrony for collaborative perception, we simulate the first collaborative perception dataset with different temporal asynchronies based on CARLA, named IRregular V2V(IRV2V).
1 paper · 1 benchmark
ISAdetect dataset (ISAdetect binary file and object code dataset)
This repository holds two datasets: one with both the original binaries and the code sections extracted from them (“full dataset”), and one with only the code sections (“only code sections”).
1 paper · 0 benchmarks
ISBNet is a dataset of images of recyclables.
1 paper · 1 benchmark
ISOD (Indoor Small Object Dataset)
ISOD contains 2,000 manually labelled RGB-D images from 20 diverse sites, each featuring over 30 types of small objects randomly placed amidst the items already present in the scenes.
1 paper · 0 benchmarks
ITCPR dataset (Image-Text Composed Person Retrieval dataset)
The ITCPR dataset is a comprehensive collection specifically designed for the Zero-Shot Composed Person Retrieval (ZS-CPR) task.
1 paper · 1 benchmark
ITDD (Industrial Textile Defect Detection)
The Industrial Textile Defect Detection (ITDD) dataset includes 1885 industrial textile images categorized into 4 categories: cotton fabric, dyed fabric, hemp fabric, and plaid fabric.
1 paper · 2 benchmarks
IVM-Mix-1M provide over 1M image-instruction pairs with corresponding instruction-relevant mask labels.
1 paper · 0 benchmarks
After defining a taxonomy of the main stone deterioration patterns and anomalies, we selected 354 highly representative images of stone-built heritage, offering them a careful selection of labels to choose from.
1 paper · 1 benchmark
IgboNLP is a standard machine translation benchmark dataset for Igbo.
1 paper · 0 benchmarks
IllusionAnimalstest Dataset Characteristics IllusionAnimalstest is a generated dataset based on a synthetic collection of animal images, including 10 animal classes: cat, dog, pigeon, butterfly, elephant, horse, deer, snake, fish, and…
1 paper · 0 benchmarks
IllusionChartest Dataset Characteristics IllusionChartest is a generated dataset containing 3,300 samples of images that feature sequences of 3 to 5 random characters.
1 paper · 0 benchmarks
IllusionFashionMNISTtest Dataset Characteristics IllusionFashionMNISTtest is a generated dataset derived from the FashionMNIST dataset.
1 paper · 0 benchmarks
IllusionMNISTtest Dataset Characteristics IllusionMNISTtest is a generated dataset derived from the MNIST dataset.
1 paper · 0 benchmarks
This publicly available dataset contains 1613 RGB-D images of field-grown broccoli plants.
1 paper · 0 benchmarks
ImageNet-Atr (ImageNet with Adversarial Text Regions)
We build a new evaluation set by adding spotting words to the images of ImageNet 2012 evaluation sets.
1 paper · 0 benchmarks
This dataset contains 6,387 ChatGPT prompts collected from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to May 2023.
1 paper · 0 benchmarks
About A realistic visual-inertial dataset with 58 sequences spanning 5km of trajectories and 1.5 hours of recordings, designed for evaluating SLAM systems in indoor pedestrian-rich environments.
1 paper · 0 benchmarks
InFashAI (Inclusive Fashion AI)
AI algorithms, and in particular Machine Learning (ML) algorithms, learn from data tasks that have been traditionally done by humans such as: image classification, facial recognition, linguistic translation etc.
1 paper · 0 benchmarks
InHARD (Industrial Human Action Recognition Dataset in the Context of Industrial Collaborative Robotics)
We introduce a RGB+S dataset named “Industrial Human Action Recognition Dataset” (InHARD) from a real-world setting for industrial human action recognition with over 2 million frames, collected from 16 distinct subjects.
1 paper · 0 benchmarks
There was no predefined dataset of party symbols to be usedas a benchmark.
1 paper · 0 benchmarks
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
The Deepfake face detection task involves a facial image of unknown authenticity for testing.
1 paper · 0 benchmarks
IndraEye (IndraEye: Infrared Electro-Optical Drone-based Aerial Object Detection Dataset)
Deep neural networks (DNNs) have demonstrated superior performance when trained on well-illuminated environments, given that the images are captured through an Electro-Optical (EO) camera, which offers rich texture content.
1 paper · 0 benchmarks
This is a real-world industrial benchmark dataset from a major medical device manufacturer for the prediction of customer escalations.
1 paper · 0 benchmarks
A dataset of books for very young children.
1 paper · 0 benchmarks
A collection of large languge model responses to tasks of propositional logic.
1 paper · 0 benchmarks
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
Inshorts News (Inshorts English News dataset)
Inshorts News dataset Inshorts provides a news summary in 60 words or less.
1 paper · 1 benchmark
InstaCities1M is a dataset of social media images with associated text.
1 paper · 0 benchmarks
This newly curated synthetic dataset specifies an additional reference region to guide image harmonization.
1 paper · 0 benchmarks
For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions.
1 paper · 0 benchmarks
Invisible Mobile Keyboard Dataset contains user initial, age, type of mobile devices, size of the screen, time taken for typing each phrase, and annotation of typed phrases with coordinate values of the typed position (x and y points).
1 paper · 0 benchmarks
IoT-23 (IoT-23: A labeled dataset with malicious and benign IoT network traffic)
IoT-23 is a dataset of network traffic from Internet of Things (IoT) devices.
1 paper · 0 benchmarks
The dataset includes source code vulnerabilities in some of the most commonly used IoT frameworks.
1 paper · 0 benchmarks
Istella LETOR (Istella Learning to Rank)
The Istella LETOR full dataset is composed of 33,018 queries and 220 features representing each query-document pair.
1 paper · 0 benchmarks
JAMBO (A Multi-Annotator Image Dataset for Benthic Habitat Classification)
The JAMBO dataset contains 3290 underwater images of the seabed captured by an ROV in temperate waters in the Jammer Bay area off the North West coast of Jutland, Denmark.
1 paper · 0 benchmarks
JAZZVAR Dataset (JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting)
Jazz pianists often uniquely interpret jazz standards.
1 paper · 0 benchmarks
We processed 241 pairs of CXR and DES soft tissue images from the JSRT dataset by performing operations like inversion and contrast adjustment to convert these images into negative formats more frequently used in clinical settings.
1 paper · 0 benchmarks
The information contained in JUSThink Dialogue and Actions Corpus dataset includes dialogue transcripts, event logs, and test responses of children aged 9 through 12, as they participate in a robot-mediated human-human collaborative…
1 paper · 0 benchmarks
Source: Text mining methodologies with R: An application to central bank texts
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.