Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 173 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 8257–8304 of 12,172
ERD (Educational Resource Discovery)
ERD (Educational Resource Discovery) is a corpus of 39,728 manually labeled web resources and 659 queries from NLP, Computer Vision (CV), and Statistics (STATS) for educational resource discovery.
1 paper · 0 benchmarks
ESA-AD (European Space Agency Dataset for Anomaly Detection in Satellite Telemetry)
ESA Anomaly Dataset is the first large-scale, real-life satellite telemetry dataset with curated anomaly annotations originated from three ESA missions.
1 paper · 0 benchmarks
Dataset Card for ESG/DLT Named Entity Recognition Dataset This dataset contains named entities related to Distributed Ledger Technology (DLT) and Environmental, Social, and Governance (ESG) topics created to support research in these areas…
1 paper · 0 benchmarks
We present ESG-FTSE, the first corpus comprised of news articles with Environmental, Social and Governance (ESG) relevance annotations.
1 paper · 0 benchmarks
ESP (Evaluation for Styled Prompt)
ESP dataset (Evaluation for Styled Prompt dataset) is a benchmark for zero-shot domain-conditional caption generation.
1 paper · 0 benchmarks
Provide: A multimodal ESSVP dataset is built with 224×224 size RGB images and 59-channel EEG data.
1 paper · 0 benchmarks
ETDII Dataset (Electric Transmission and Distribution Infrastructure Imagery Dataset)
Paper: GridTracer: Automatic Mapping of Power Grids using Deep Learning and Overhead Imagery Authors: Bohao Huang, Jichen Yang, Artem Streltsov, Kyle Bradbury, Leslie M.
1 paper · 1 benchmark
This dataset contains 27 ROS bags of point clouds produced by a Kinect based the ground truth obtained from a Vicon pose capture system.
1 paper · 0 benchmarks
This group of datasets was recorded with the aim to test point cloud registration algorithms in specific environments and conditions.
1 paper · 0 benchmarks
The ETHZ Shape dataset contains images of five diverse shape-based classes, collected from Flickr and Google Images.
1 paper · 0 benchmarks
EU Long-term Dataset with Multiple Sensors for Autonomous Driving was collected with a robocar, equipped with eleven heterogeneous sensors, in the downtown and suburban areas of Montbéliard in France.
1 paper · 0 benchmarks
EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework Authors: Weina Jin, Jianyu Fan, Diane Gromala, Philippe Pasquier, Ghassan Hamarneh Introduction: EUCA dataset is for modelling personalized or…
1 paper · 0 benchmarks
EUEN17037 Daylight and View Standard Test Dataset.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
This is an assembly dataset built on top of ShellcodeIA32, a dataset for automatically generating assembly from natural language descriptions that consists of 3,200 assembly instructions, commented in the English language, which were…
1 paper · 0 benchmarks
This dataset contains samples to generate Python code for security exploits.
1 paper · 0 benchmarks
Real and simulated lidar data of indoor and outdoor scenes, before and after geometric scene changes have occurred.
1 paper · 0 benchmarks
The EXPO-HD Dataset is a dataset of Expo whiteboard markers for the purpose of instance segmentation.
1 paper · 0 benchmarks
EarlyNSD (Early Nutrient Stress Detection of Plants)
Early detection of plant nutritional deficiencies, followed by corrective actions, is essential for sustaining crop yield.
1 paper · 1 benchmark
A Zero-Shot Sketch-based Inter-Modal Object Retrieval Scheme for Remote Sensing Images WITH the advancement in sensor technology, huge amounts of data are being collected from various satellites.
1 paper · 0 benchmarks
The dataset, generated from a scientific simulation, consists of a time series (251 steps) of 3D scalar fields on a spherical 180x201x360 grid covering 500 Myr of geological time.
1 paper · 0 benchmarks
A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner: 1.
1 paper · 0 benchmarks
Echo Corpus (Arviv et al, 2021) infused with information from KnowledJe (Halevy, 2023).
1 paper · 0 benchmarks
Echocardiography, or cardiac ultrasound, is the most widely used and readily available imaging modality to assess cardiac function and structure.
1 paper · 0 benchmarks
EchoNet-Dynamic is a dataset of over 10k echocardiogram, or cardiac ultrasound, videos from unique patients at Stanford University Medical Center.
1 paper · 0 benchmarks
Edge-Map-345C is a large-scale edge-map dataset including 290,281 edge-maps corresponding to 345 object categories of QuickDraw dataset.
1 paper · 0 benchmarks
Edina-DR is a novel corpus of discourse relation pairs; the first of its kind to attempt to identify the discourse relations connecting the dialogic turns in open-domain discourse.
1 paper · 0 benchmarks
Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based…
1 paper · 0 benchmarks
EgoISM-HOI is a new multimodal dataset composed of synthetic and real images of egocentric human-objects interactions in an industrial environment with rich annotations of hands and objects.
1 paper · 0 benchmarks
EgoMon (Egomon Gaze & Video dataset)
EgoMon Gaze & Video Dataset is an Egocentric (first person) Dataset that consists of 7 videos of 30 minutes, more or less, each one of them.
1 paper · 0 benchmarks
EgoPW training dataset with scene annotations Since we want to generalize to data captured with a real head-mounted camera, we also extended the EgoPW training dataset.
1 paper · 0 benchmarks
Contains 350 tweets with more than 8,000 words including 3,000 unique words written in Egyptian dialect.
1 paper · 0 benchmarks
A large comparable corpus for Basque-Spanish was prepared, on the basis of independently-produced news by the Basque public broadcaster EiTB.
1 paper · 0 benchmarks
EleThermal (EleThermal: Infrared Elephant Images Dataset)
This is the Infrared Elephant Images Dataset (named 'EleThermal dataset') collected from here and annotated by our project, released under GPLv3.
1 paper · 0 benchmarks
An open data corpus of 123.610 labeled samples, Source: Electro-Magnetic Side-Channel Attack Through Learned Denoising and Classification
1 paper · 0 benchmarks
Each HDF5 file has the following structure: energy Dataset {100000, 1} layer0 Dataset {100000, 3, 96} layer1 Dataset {100000, 12, 12} layer2 Dataset {100000, 12, 6} overflow Dataset {100000, 3} In practice, each file is a collection of…
1 paper · 0 benchmarks
A configurable synthetic dataset of simple shapes with ground truth concepts and known causal relationships between concepts and classes.
1 paper · 0 benchmarks
This is a detailed description of the dataset, a data sheet for the dataset as proposed by Gebru et al.
1 paper · 0 benchmarks
EmoFilm (Emotional speech from Films)
EmoFilm is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages.
1 paper · 0 benchmarks
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
1 paper · 0 benchmarks
The dataset provides News articles obtained from emol.cl including their content, title and all the comments it received in JSON format
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Predictions of energy consumption are crucial for energy retailers to minimize deviations from energy acquired in the day-ahead market and the actual consumption of their customers.
1 paper · 0 benchmarks
The "Microbundle Time-lapse Dataset" contains 24 experimental time-lapse images of cardiac microbundles using three distinct types of experimental testbed of beating lab grown hiPSC-based cardiac microbundles.
1 paper · 0 benchmarks
The dataset consists of a set of tasks in email and the associated people.
1 paper · 0 benchmarks
This is a dataset of open source software developed mainly by enterprises rather than volunteers.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.