Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 219 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10465–10512 of 12,172
The TIC Dataset consists of 2056 images (512x640) of transmission line network footage in Greece (Northeast Attica) and annotations of three object classes, i.e.
1 paper · 0 benchmarks
TILT corpus (GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit)
A corpus of GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit (TILT).
1 paper · 0 benchmarks
TITANIC-FGS is the first domain knowledge-enhanced, instruction-following dataset specifically designed for the Remote Sensing Fine-Grained Ship Classification (RS-FGSC) task.
1 paper · 0 benchmarks
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset.
1 paper · 1 benchmark
TLFM dataset (TLFM dataset for microscopy image sequence generation)
TLFM dataset structured in sequences of at least nine timesteps.
1 paper · 0 benchmarks
TLMSDD (Token Level Multi-target Stance Detection Dataset)
none
1 paper · 0 benchmarks
TMBuD is a dataset for building recognition and 3D reconstruction of human made structures in urban scenarios.
1 paper · 0 benchmarks
TML1M (Table-MovieLens1M)
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset.
1 paper · 1 benchmark
TPIC17 (Temporal Popularity Image Collection)
Image dataset with about 600K Flickr photos.
1 paper · 0 benchmarks
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
TRADES-LOB comprises simulated TRADES market data for Tesla and Intel, for 29/01 and 30/01.
1 paper · 0 benchmarks
TREC-05 (TREC 2005 Spam Public Corpora)
1 paper · 0 benchmarks
The dataset used for TREC 2017 Dynamic Domain Track consists of two domains: Ebola and New York Times.
1 paper · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
1 paper · 1 benchmark
The dataset is composed of 100 video sequences densely annotated with 60K bounding boxes, 17 sequence attributes, 13 action verb attributes and 29 target object attributes.
1 paper · 0 benchmarks
TREx-2p is a dataset to probe whether a pretrained LM possesses “indirect” 2-hop knowledge.
1 paper · 0 benchmarks
TRR360D is based on the ICDAR2019MTD modern table detection dataset, it refers to the annotation format of the DOTA dataset.
1 paper · 1 benchmark
Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers.
1 paper · 0 benchmarks
For slide editing, this benchmark dataset provide the pair of user instruction and corresponding slides (pptx).
1 paper · 0 benchmarks
TSFM-ScalingLaws-Dataset This is the dataset for the paper Towards Neural Scaling Laws for Time Series Foundation Models.
1 paper · 0 benchmarks
In this dataset, we provide detailed traffic stream data for the Spot robot, including both the Spot robot control traffic stream and the Spot video stream.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
100 samples each of synthetic speech generated by 9 moderns TTS systems.
1 paper · 0 benchmarks
TUAC (Temple University Artifact Corpus)
A new subset of the popular open source electroencephalogram (EEG) corpus – TUH EEG: - The Temple University Artifact Corpus (TUAR) consists of high yield artifact files annotated using a five-way classification system: 1.
1 paper · 0 benchmarks
TUD (Table Uniformity Dataset)
Dataset to reproduce results for the paper Detecting CSV file dialects by table uniformity measurement and data type inference, DOI ds240062.
1 paper · 1 benchmark
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks
The field of converting natural language into corresponding SQL queries using deep learning techniques has attracted significant attention in recent years.
1 paper · 0 benchmarks
TVPReid (Text-to-Video Person Re-identification)
The TVPReid dataset contains 6559 pedestrian videos, each of which is annotated with two text descriptions, for a total of 13118 descriptions.
1 paper · 0 benchmarks
TVRecap a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved.
1 paper · 4 benchmarks
The TWT16 dataset contains ~30k conversations in Twitter, collected from January to June 2016.
1 paper · 0 benchmarks
TXL-PBC dataset (a freely accessible labeled peripheral blood cell dataset)
The TXL-PBC Dataset is a comprehensive collection of re-annotated and integrated cell images from multiple cell datasets.
1 paper · 0 benchmarks
TYC Dataset (The TYC Dataset for Understanding Instance-Level Semantics and Motions of Cells in Microstructures)
We introduce the trapped yeast cell (TYC) dataset, a novel dataset for understanding instance-level semantics and motions of cells in microstructures.
1 paper · 0 benchmarks
TaRBench is a comprehensive benchmark that we developed to evaluate the effectiveness of TaRGet in automated test case repair.
1 paper · 0 benchmarks
In ICDAR-17, a Page-Object Detection (POD) competition was organized where the task was to identify page objects in documents which includes tables, figures and equations in document.
1 paper · 0 benchmarks
This data set contains real-world table tennis ball trajectories recorded with our custom developed table tennis ball launcher AIMY.
1 paper · 0 benchmarks
We present a benchmark dataset for tactile-based object tracking, featuring 12 distinct objects and 84 tracking trials—7 trials per object, each lasting an average of 10.2 seconds.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated version of the Alpaca dataset.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated versions of the Alpaca dataset and a subset of OpenOrca dataset.
1 paper · 0 benchmarks
Taobao dataset which is pre-processed in TGN Style.
1 paper · 0 benchmarks
Dataset for task complexity classification and the complexity score prediction
1 paper · 0 benchmarks
PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.
1 paper · 0 benchmarks
A truly multimodal dataset for benchmarking deep learning models on ecological tasks.
1 paper · 0 benchmarks
Please refer this paper @article{sarker2024tea, author = {Sarker, Swapnil Sharma and Islam, Ashiqul and Talukder Raktim, Raufun and Roshni, Sanjana and Joy, Sajib Kumar Saha and Shah, Faisal}, title = {Real-Time Tea Leaf Disease Detection…
1 paper · 0 benchmarks
This dataset consists of RGB-D images captured using 12 Intel RealSense cameras.
1 paper · 0 benchmarks
Tecnalia Hyperspectral Dataset contains different non-ferreous fractions of Waste from Electric and Electronic Equipment (WEEE) of Copper, Brass, Aluminum, Stainless Steel and White Copper.
1 paper · 0 benchmarks
The acquisition over the VIS and TIR data was performed by a commercial thermal camera Testo 882-3.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.