Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 121 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5761–5808 of 12,172
Contains 6016 image-pairs from the wild, shedding light upon a rich and diverse set of criteria employed by human beings.
3 papers · 0 benchmarks
We present the development of a Named Entity Recognition (NER) dataset for Tagalog.
3 papers · 0 benchmarks
TREC Submissions for all Ad Hoc Retrieval runs.
3 papers · 0 benchmarks
The TREC News Track features modern search tasks in the news domain.
3 papers · 1 benchmark
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
3 papers · 1 benchmark
TT100K (Tsinghua-Tencent 100K(official training and testing set))
Trainging and testing data: The original training set includes 6105 images, and the original testing set includes 3071 images.
3 papers · 1 benchmark
This task offers researchers an opportunity to test their fine-grained classification methods for detecting and recognizing strokes in table tennis videos.
3 papers · 2 benchmarks
TTStroke-21 for MediaEval 2022.
3 papers · 2 benchmarks
A novel dataset with a diverse set of sequences in different scenes for evaluating VI odometry.
3 papers · 0 benchmarks
TV-AD (Audio Description dataset for TV series)
TV-AD is a dataset that provides ground truth AD annotations that are aligned with TV series video, featuring episodes across multiple TV series including “The Big Bang Theory”, “Friends”, “Frasier”, “Seinfeld”, etc.
3 papers · 0 benchmarks
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
Taillar's permutation flow shop, the job shop, and the open shop scheduling problems instances: We restrict ourselves to basic problems: the processing times are fixed, there are neither set-up times nor due dates nor release dates, etc.
3 papers · 0 benchmarks
3D meshes of various garments of various sizes draped on people with various body poses and shapes.
3 papers · 0 benchmarks
A novel dataset of document-grounded task-based dialogues, where an Information Giver (IG) provides instructions (by consulting a document) to an Information Follower (IF), so that the latter can successfully complete the task.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
The TbV dataset is large-scale dataset created to allow the community to improve the state of the art in machine learning tasks related to mapping, that are vital for self-driving.
3 papers · 0 benchmarks
The dataset contains 2,000 houses, 13,478 rooms and 873 (some rooms have same textures so this number is smaller than the total number of rooms.) texture images with corresponding natural language descriptions.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
It is a freely available resource for research on handling negation and uncertainty in biomedical texts .
3 papers · 5 benchmarks
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
A vast amount of information in the biomedical domain is available as natural language free text.
3 papers · 0 benchmarks
Two datasets (synthetic and natural/real) containing simultaneously recorded egocentric and exocentric videos.
3 papers · 0 benchmarks
ThreeDWorld Transport Challenge is a visually-guided and physics-driven task-and-motion planning benchmark.
3 papers · 0 benchmarks
Tiny ImageNetv2 is a subset of the ImageNetV2 (matched frequency) dataset by Recht et al.
3 papers · 0 benchmarks
Title2Event is a large-scale sentence-level dataset for benchmarking Open Event Extraction without restricting event types.
3 papers · 0 benchmarks
Tobacco800 is a public subset of the complex document image processing (CDIP) test collection constructed by Illinois Institute of Technology, assembled from 42 million pages of documents (in 7 million multi-page TIFF images) released by…
3 papers · 0 benchmarks
The ToolLens dataset consists of 18,770 concise yet intentionally multifaceted queries, each associated with 1 to 3 verified tools out of a total of 464, designed to better mimic real-world user interactions.
3 papers · 1 benchmark
We ran 21 recommender systems on three datasets (BeerAdvocate, LibraryThing and MovieLens 1M).
3 papers · 0 benchmarks
Tragic Talkers is an audio-visual dataset consisting of excerpts from the "Romeo and Juliet" drama captured with microphone arrays and multiple co-located cameras for light-field video.
3 papers · 0 benchmarks
This dataset contains aircraft trajectories in an untowered terminal airspace collected over 8 months surrounding the Pittsburgh-Butler Regional Airport [ICAO:KBTP], a single runway GA airport, 10 miles North of the city of Pittsburgh,…
3 papers · 1 benchmark
Travel (Tour & Travels Customer Churn Prediction)
A Tour & Travels Company Wants To Predict Whether A Customer Will Churn Or Not Based On Indicators Given Below.
3 papers · 1 benchmark
TriBERT dataset consists of 12,049 training, 2,527 validation and 2,560 test Human-Machine collaborative texts.
3 papers · 1 benchmark
Three wild-type (C57BL/6J) male mice ran on a paper spool following odor trails (Mathis et al 2018).
3 papers · 1 benchmark
TriageSQL is a cross-domain text-to-SQL question intention classification benchmark that requires models to distinguish four types of unanswerable questions from answerable questions.
3 papers · 0 benchmarks
Dataset of restaurant reviews from TripAdvisor that includes images and texts uploaded in reviews by users.
3 papers · 0 benchmarks
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
Although promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images.
3 papers · 1 benchmark
TuSimple Lane is an extension of the TuSimple dataset with 14,336 lane boundaries annotations.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
TyDiP (A Dataset for Politeness Classification in Nine Typologically Diverse Languages)
A Dataset for Politeness Classification in Nine Typologically Diverse Languages (TyDiP) is a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
3 papers · 0 benchmarks
The archive contains original images from U2OS cells stained with Hoechst 33342 as PNG files.
3 papers · 0 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
UAV-GESTURE is a dataset for UAV control and gesture recognition.
3 papers · 0 benchmarks
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
UESTC-MMEA-CL (A multi-modal egocentric activity dataset for continual learning)
UESTC-MMEA-CL is a new multi-modal activity dataset for continual egocentric activity recognition, which is proposed to promote future studies on continual learning for first-person activity recognition in wearable applications.
3 papers · 0 benchmarks
The UFPR-Periocular dataset has 16,830 images of both eyes (33,660 cropped images of each eye) from 1,122 subjects (2,244 classes).
3 papers · 0 benchmarks
UIIS10K (General Underwater Image Instance Segmentation dataset 10K)
We propose a large-scale underwater instance segmentation dataset, UIIS10K, which includes 10,048 images with pixel-level annotations for 10 categories.
3 papers · 0 benchmarks
The dataset comprises 4,500 question-answer pairs collected from trusted medical sources, with at least one answer and at most four unique paraphrased answers per question
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.