Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 16 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 721–768 of 3,998
Node classification on Film with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
GlobalOpinionQA consists of questions and answers from cross-national surveys designed to capture diverse opinions on global issues across different countries.
14 papers · 0 benchmarks
Inter-X is a large-scale dataset containing ~11K interaction sequences, more than 8.1M frames and 34K fine-grained human textual descriptions.
14 papers · 1 benchmark
LAV-DF (Localized Audio Visual DeepFake Dataset)
Localized Audio Visual DeepFake Dataset (LAV-DF).
14 papers · 1 benchmark
MAP (Maybe Ambiguous Pronoun)
Maybe Ambiguous Pronoun is a dataset similar to GAP dataset, but without binary gender constraints.
14 papers · 0 benchmarks
MAVE (MAVE: : A Product Dataset for Multi-source Attribute Value Extraction)
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
14 papers · 2 benchmarks
MIntRec is a novel dataset for multimodal intent recognition.
14 papers · 1 benchmark
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark
Open-Platypus is a family of fine-tuned and merged Large Language Models (LLMs) that achieves the strongest performance and currently stands at first place in HuggingFace's Open LLM Leaderboard.
14 papers · 0 benchmarks
A new large scale plane geometry problem solving dataset called PGPS9K, labeled both fine-grained diagram annotation and interpretable solution program.
14 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
SinD (A Drone Dataset at Signalized Intersection in China)
The SIND dataset is based on 4K video captured by drones, providing information including traffic participant trajectories, traffic light status, and high-definition maps
14 papers · 0 benchmarks
Purpose Medical imaging has become increasingly important in diagnosing and treating oncological patients, particularly in radiotherapy.
14 papers · 0 benchmarks
The T2Dv2 dataset consists of 779 tables originating from the English-language subset of the WebTables corpus.
14 papers · 4 benchmarks
Node classification on Texas with the fixed 48%/32%/20% splits provided by Geom-GCN.
14 papers · 2 benchmarks
speechocean762 is an open-source speech corpus designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are children.
14 papers · 3 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
13 papers · 2 benchmarks
This dataset contains benchmark scores for EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs).
13 papers · 1 benchmark
Flare7K, the first nighttime flare removal dataset, which is generated based on the observation and statistic of real-world nighttime lens flares.
13 papers · 1 benchmark
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
The Ghostbusters dataset leverages the GPT-3.5-turbo model for generating texts in the domains of creative writing, news, and student essays, providing 2,000 texts in the first two domains and 1,994 in the latter.
13 papers · 1 benchmark
GooAQ is a large-scale dataset with a variety of answer types.
13 papers · 0 benchmarks
HINT3 is a dataset for intent detection.
13 papers · 0 benchmarks
HopeEDI (HopeEDI: A Multilingual Hope Speech Detection Dataset for Equality, Diversity, and Inclusion)
Over the past few years, systems have been developed to control online content and eliminate abusive, offensive or hate speech content.
13 papers · 4 benchmarks
L3DAS22: MACHINE LEARNING FOR 3D AUDIO SIGNAL PROCESSING This dataset supports the L3DAS22 IEEE ICASSP Gand Challenge.
13 papers · 0 benchmarks
MedConceptsQA - Open Source Medical Concepts QA Benchmark The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA
13 papers · 2 benchmarks
The Multilingual Reuters Collection dataset comprises over 11,000 articles from six classes in five languages, i.e., English (E), French (F), German (G), Italian (I), and Spanish (S).
13 papers · 0 benchmarks
NTU4DRadLM is a novel 4D radar dataset specifically proposed for research on robust SLAM, based on 4D radar, thermal camera, and IMU.
13 papers · 0 benchmarks
It contains 15K triplets of essay problem statements, student-written, and LLM-generated essays.
13 papers · 0 benchmarks
PARANMT-50M is a dataset for training paraphrastic sentence embeddings.
13 papers · 0 benchmarks
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.
13 papers · 1 benchmark
Real HSI (End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention)
End-to-End Low Cost Compressive Spectral Imaging with Spatial-Spectral Self-Attention
13 papers · 1 benchmark
The goal of the Robust track is to improve the consistency of retrieval technology by focusing on poorly performing topics.
13 papers · 1 benchmark
The SARDet-100K dataset encompasses a total of 116,598 images, and 245,653 instances distributed across six categories: Aircraft, Ship, Car, Bridge, Tank, and Harbor.
13 papers · 1 benchmark
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.
13 papers · 2 benchmarks
This dataset encompasses a diverse range of tactile features that are instrumental in bifurcating various material properties.
13 papers · 0 benchmarks
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them.
13 papers · 1 benchmark
ViQuAE is a dataset for KVQAE (Knowledge-based Visual Question Answering about named Entities), a task which consists in answering questions about named entities grounded in a visual context using a Knowledge Base.
13 papers · 0 benchmarks
A Multi-Task 4D Radar-Camera Fusion Dataset for Autonomous Driving on Water Surfaces description of the dataset WaterScenes, the first multi-task 4D radar-camera fusion dataset on water surfaces, which offers data from multiple sensors,…
13 papers · 2 benchmarks
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
Yelp-Fraud (Multi-relational Graph Dataset for Yelp Spam Review Detection)
Yelp-Fraud is a multi-relational graph dataset built upon the Yelp spam review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
13 papers · 3 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
The dataset contains product information from AliExpress Sports & Entertainment category.
12 papers · 2 benchmarks
BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)
Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles.
12 papers · 1 benchmark
Bongard-HOI testifies to which extent your few-shot visual learner can quickly induce the true HOI concept from a handful of images and perform reasoning with it.
12 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.