Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 251 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 12001–12048 of 12,172
Accompanying software package and data for the publication titled "Bayesian multi-exposure image fusion for robust high dynamic range ptychography".
0 papers · 0 benchmarks
Malaria, a mosquito-borne infectious disease affecting humans and other animals, is widespread in the tropical and subtropical regions.
0 papers · 0 benchmarks
The SweDN 1.0 dataset is a valuable resource for natural language processing (NLP) tasks, specifically text summarization.
0 papers · 0 benchmarks
The SweFAQ dataset is a collection of frequently asked questions from Swedish authorities' websites with shuffled answers.
0 papers · 0 benchmarks
SynD (A Synthetic Energy Dataset for Non-Intrusive Load Monitoring in Households)
SynD is a synthetic energy dataset with a focus on residential buildings.
0 papers · 0 benchmarks
TAL-SCQ5K-EN/TAL-SCQ5K-CN are high-quality mathematical competition datasets in English and Chinese language created by TAL Education Group, each consisting of 5K questions(3K training and 2K testing).
0 papers · 0 benchmarks
The TAU Spatial Sound Events 2019 - Ambisonic dataset contains recordings from a scene (along with the Microphone Array sister dataset).
0 papers · 0 benchmarks
The TAU Spatial Sound Events 2019 – Microphone Array dataset contains recordings from a scene (along with the Ambisonic sister dataset).
0 papers · 0 benchmarks
TCMP-300 (Traditional Chinese Medicinal Plant Dataset)
Traditional Chinese medicinal plants are often used to prevent and treat diseases for the human body.
0 papers · 1 benchmark
Dataset Introduction TFHAnnotatedDataset is an annotated patent dataset pertaining to thin film head technology in hard-disk.
0 papers · 0 benchmarks
Face detection and subsequent localization of facial landmarks are the primary steps in many face applications.
0 papers · 0 benchmarks
The increase in religiously motivated hate on social media is clear and ongoing.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
The TICO-19 dataset is a translation initiative focused on COVID-19 content, created by academic and industry partners along with Translators without Borders.
0 papers · 0 benchmarks
TLHDIBD2021 (Tai Le historical document image binarization dataset)
Hybrid-CBF: A hybrid classification and binarization framework for historical Tai Le document image binarization The binarization of historical documents is very important and more challenging than the binarization of ordinary documents.
0 papers · 0 benchmarks
The “Toyota Motor Europe (TME) Motorway Dataset” is composed by 28 clips for a total of approximately 27 minutes (30000+ frames) with vehicle annotation.
0 papers · 0 benchmarks
TRACT (Tweets Reporting Abuse Classification Task Corpus)
TRACT is a small scale manually annotated corpus for abuse classification problem.
0 papers · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
0 papers · 0 benchmarks
TS-TR (Turkish Scene Text Recognition Dataset)
The Turkish Scene Text Recognition (TS-TR) dataset was primarily developed to fill the gap in non-English text recognition resources, specifically addressing the unique challenges presented by the Turkish language, such as special…
0 papers · 0 benchmarks
Arabic multi-dialectal hate speech dataset.
0 papers · 0 benchmarks
The TURBID, is an open image dataset that has been generated to contribute with the underwater research area.
0 papers · 0 benchmarks
TURSpider (TURSpider: A Turkish Text-to-SQL Dataset)
TURSpider is a novel Turkish Text-to-SQL dataset that includes complex queries, akin to those in the original Spider dataset.
0 papers · 0 benchmarks
TURSpider is a novel Turkish Text-to-SQL dataset that includes complex queries, akin to those in the original Spider dataset.
0 papers · 0 benchmarks
TUT Rare Sound events 2017, development dataset consists of source files for creating mixtures of rare sound events (classes baby cry, gun shot, glass break) with background audio, as well a set of readily generated mixtures and recipes…
0 papers · 0 benchmarks
The TUT Sounds Event 2018 dataset consists of real-life first order Ambisonic (FOA) format recordings with stationary point sources each associated with a spatial coordinate.
0 papers · 0 benchmarks
TVPR (Top-View Person Re-Identification Dataset)
The TVPR (Top View Person Re-identification) dataset stores depth frames (640x480) collected using Asus Xtion Pro Live in top-view configuration.
0 papers · 0 benchmarks
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection.
0 papers · 0 benchmarks
Este conjunto de datos consiste en comentarios de publicaciones del MINSA (Perú) en Facebook sobre la vacuna contra el VPH entre los años 2019 y 2020.
0 papers · 0 benchmarks
Collection of stream of consciousness.
0 papers · 0 benchmarks
The Reddit COVID Dataset is a dataset of 4.51M Reddit posts and 17.8M comments - all mentions of COVID until 2021-10-25 across the entire Reddit social network.
0 papers · 0 benchmarks
Thermal Solar Plants Dataset (Market-Oriented Flow Allocation for Thermal Solar Plants: An Auction-Based Methodology with Artificial Intelligence)
This dataset includes 120 simulations of 10 loops of a 50 MW parabolic-trough solar plant with varying solar irradiances, optical efficiencies, thermal losses, ambient temperatures, input temperatures, and sector flow rates.
0 papers · 0 benchmarks
TiMoS (Tropes in Movie Synopses)
Tropes in Movie Synopses (TiMoS) is a dataset of movie tropes collected from a Wikipedia-style website, TVTropes3 with 5623 movie synopses associated with 95 most occurred tropes.
0 papers · 0 benchmarks
These sequences of hyperspectral radiance images have been taken from scenes undergoing natural illumination changes.
0 papers · 0 benchmarks
This dataset, commissioned by the Yandex Business Directory, contains 10,000 photos of organization information signs shot in the Russian Federation along with the INN (taxpayer ID) and OGRN (Primary State Registration Number) codes shown…
0 papers · 0 benchmarks
This datase, contains 1244 images of hot and cold water meters as well as their readings and coordinates of the displays showing those readings.
0 papers · 0 benchmarks
Top-notch Flutter App Development Company Our team of skilled Flutter app developers can assist you in creating platform-independent digital human experiences that transcend device boundaries.
0 papers · 0 benchmarks
The Tornado Network (TorNet) dataset is a large, high-resolution benchmark dataset developed to support machine learning research in tornado detection and prediction.
0 papers · 0 benchmarks
Toronto NeuroFace Dataset: A New Dataset for Facial Motion Analysis in Individuals with Neurological Disorders Toronto NeuroFace Dataset is a public dataset with videos of oro-facial gestures performed by individuals with oro-facial…
0 papers · 0 benchmarks
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.