Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 253 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 12097–12144 of 12,172
Datasets at https://zenodo.org/record/8105485 for Motion Robust CMR Reconstruction Code in https://github.com/syedmurtazaarshad/motion-robust-CMR
0 papers · 0 benchmarks
Vript (🎬 Vript: A Video Is Worth Thousands of Words)
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips).
0 papers · 0 benchmarks
The VXC TSG is based on samples taken from the ceramic tile industry and is comprised of 14 ceramic tile models, 42 surface grades and 960 pieces.
0 papers · 0 benchmarks
The WASABI Song Corpus is a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.
0 papers · 0 benchmarks
reference paper YOUNIS, H., RATTROUT, A., & YOUNIS, M.
0 papers · 0 benchmarks
WMT21 (Workshop on Machine Translation 2021) Translation Task focuses on news text translation.
0 papers · 0 benchmarks
WMT 2021 Ge'ez-Amharic is a Ge'ez-Amharic dataset prepared for NMT tasks of the 6th Workshop on NLP at Debre Berhan University, Ethiopia.
0 papers · 0 benchmarks
WRV (Wire-removal Dataset)
G2LP Wire-removal Dataset in G2LP-Net: Global to Local Progressive Video Inpainting Network Wire-removal Dataset The WRV dataset has been specifically curated for the challenges of video inpainting in irregularly slender regions.
0 papers · 0 benchmarks
WTA/TLA (WTA/TLA: A UAV-captured Dataset for Semantic Segmentation of Energy Infrastructure)
WTA (Wind Turbine Aerial) and TLA (Transmission Line Aerial) are public datasets which contain a set of RGB images from wind turbine farms and transmission towers and power lines, along with semantic ground truth for relevant classes.
0 papers · 0 benchmarks
WalnutData (A UAV Remote Sensing Dataset of Green Walnuts and Model Evaluation)
With the gradual maturity of UAV technology, it can provide extremely powerful support for smart agriculture and precise monitoring.
0 papers · 0 benchmarks
The RGB-D Scenes Dataset contains 8 scenes annotated with objects that belong to the Washington RGB-D Object Dataset.
0 papers · 0 benchmarks
The RGB-D Scenes Dataset v2 consists of 14 scenes containing furniture (chair, coffee table, sofa, table) and a subset of the objects in the RGB-D Object Dataset (bowls, caps, cereal boxes, coffee mugs, and soda cans).
0 papers · 0 benchmarks
WiSARD (Wilderness Search and Rescue Dataset)
WiSARD stands for Wilderness Search and Rescue Dataset (pronounced "wizard").
0 papers · 0 benchmarks
A method for automatically gathering massive amounts of naturally-occurring cross-document reference data is used to create the Wikilinks dataset comprising of 40 million mentions over 3 million entities.
0 papers · 0 benchmarks
Wikidata is a free and open knowledge base that can be read and edited by both humans and machines.
0 papers · 0 benchmarks
Wireless-Intelligence is a database website provided for AI-based wireless communication research, in which each dataset consists of hundreds and thousands of channel samples in different forms.
0 papers · 0 benchmarks
Social media message with sentiment label (positive, neutral, negative, question).
0 papers · 0 benchmarks
We provide a Mikolov-style word-analogy evaluation set specifically for Bangla, with a sample size of 16678, as well as a translated and curated version of the Mikolov dataset, which contains 10594 samples for cross-lingual research.
0 papers · 0 benchmarks
X-Wines (A Wine Dataset for Recommender Systems and Machine Learning)
X-Wines is a consistent wine dataset containing 100,646 instances and 21 million real evaluations carried out by users.
0 papers · 0 benchmarks
Yesno is an audio dataset consisting of 60 recordings of one individual saying yes or no in Hebrew; each recording is eight words long.
0 papers · 0 benchmarks
YTsubtitles is a remarkable tool designed for building a dataset from YouTube subtitles.
0 papers · 0 benchmarks
Z-Bench is a fascinating Chinese language model prompt dataset developed by an enthusiastic AI-focused team at Zhenfund.
0 papers · 0 benchmarks
Description: The ZakynthosTurtles dataset has been designed to support the development of numerical methods for the recognition and re-identification of individual sea turtles based on their unique scale patterns.
0 papers · 0 benchmarks
ZooScanNet (ZooScanNet: plankton images captured with the ZooScan)
Plankton was sampled with various nets, from bottom or 500m depth to the surface, in many oceans of the world.
0 papers · 0 benchmarks
academy 3 vs 1 with keeper on Google Research Football
0 papers · 0 benchmarks
A Chinese Mandarin speech corpus by Beijing DataTang Technology Co., Ltd, containing 200 hours of speech data from 600 speakers.
0 papers · 0 benchmarks
bioRxiv is a free online archive for unpublished preprints in the life sciences.
0 papers · 0 benchmarks
The cMedQA dataset is designed for Chinese community medical question answering.
0 papers · 0 benchmarks
Debate.org is a debate platform that is organized in rounds where each of two opponents submits posts arguing for their side.
0 papers · 0 benchmarks
contain the clinical trial dataset
0 papers · 0 benchmarks
eAppleScab (Apple Scab in the Early Stage of Development)
The study showed that the apple scab can be detected in the high-resolution RGB images in an early stage of its development.
0 papers · 0 benchmarks
This dataset was built based on a subset of foraminifer samples from the Yale Peabody Museum (YPM) Coretop Collection and the Natural History Museum, Lon- don (NHM) Henry A.
0 papers · 0 benchmarks
fNIRS2MW (The Tufts fNIRS to Mental Workload Dataset)
The Tufts fNIRS to Mental Workload (fNIRS2MW) open-access dataset is a new dataset for building machine learning classifiers that can consume a short window (30 seconds) of multivariate fNIRS recordings and predict the mental workload…
0 papers · 0 benchmarks
The data was captured from an overhead perspective, showcasing the swimming behavior of fish in a simulated flowing water channel.
0 papers · 0 benchmarks
A large set of images of flowers Homepage: https://www.tensorflow.org/tutorials/loaddata/images Dataset size: 221.83 MiB
0 papers · 0 benchmarks
Freefield1010 is a collection of 7,690 excerpts from field recordings around the world, gathered by the FreeSound project, and then standardised for research.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.