Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 222 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 10609–10656 of 12,172
The lists of Tweet IDs for the experiments of the article: This Sample seems to be good enough!
1 paper · 0 benchmarks
"My ridiculous dog is amazing." [sentiment: positive] With all of the tweets circulating every second it is hard to tell whether the sentiment behind a specific tweet will impact a company, or a person's, brand for being viral (positive),…
1 paper · 0 benchmarks
The TwinSynths dataset is a novel benchmark designed to overcome common limitations found in earlier synthetic image datasets, such as low image quality, inadequate content preservation, and limited class diversity.
1 paper · 0 benchmarks
This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.
1 paper · 0 benchmarks
Twitter Cyberthreat Detection Dataset is a dataset that contains tweets from two sets of accounts related to cybersecurity.
1 paper · 0 benchmarks
This is a dataset for detection fake death hoaxes.
1 paper · 0 benchmarks
This dataset contains two subsets of flood images from Twitter: The Harz17 dataset comprises images from tweets containing flood-related keywords during the occurrence of a flood in the Harz region in Germany in July 2017.
1 paper · 0 benchmarks
The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video).
1 paper · 0 benchmarks
The data is about 1.5 million English tweets annotated for part-of-speech using Ritter's extension of the PTB tagset.
1 paper · 0 benchmarks
We introduce a dataset consisting of 1314 samples, including users’ tweets and bios.
1 paper · 0 benchmarks
Twitter-HyDrug is a real-world hypergraph data that describes the drug trafficking communities on Twitter.
1 paper · 1 benchmark
This benchmark hypergraph dataset, Twitter-HyDrug-UR, is derived from Twitter-HyDrug by HyGCL-DC.
1 paper · 1 benchmark
This task aims to extract named entities and entity types while further predicting segmentation masks of visual objects.
1 paper · 1 benchmark
The two Coiling Spiral is a 2d classification dataset composed of two classes; each spiral corresponds to one class.
1 paper · 0 benchmarks
Dataset accompanying paper Klein, N., Siegle, J.H., Teichert, T., Kass, R.E.
1 paper · 0 benchmarks
Two4Two (A Synthetic Dataset For Controlled Experiments)
Two4Two is a library to create synthetic image data crafted for human evaluations of interpretable ML approaches (esp.
1 paper · 0 benchmarks
Typography-MNIST is a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts.
1 paper · 0 benchmarks
This dataset supports the research detailed in the pre-print "Virtual Imaging Trials Improved the Transparency and Reliability of AI Systems in COVID-19 Imaging." The study employs both clinical and simulated CT data to evaluate AI models…
1 paper · 1 benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding.
1 paper · 0 benchmarks
UAGD (Uniform Age and Gender Dataset)
The source images of UAGD is manually selected from APPA-REAL, UTKFace and AgeDB datasets very carefully, which means only face images that are having large poses, containing noise pixels, bearing various expressions, and under different…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
UAV-PDD2023: A benchmark dataset for pavement distress detection based on UAV images
1 paper · 1 benchmark
Mapping urban large-area advertising structures using drone imagery and deep learning-based spatial data analysis.
1 paper · 1 benchmark
UAVDB (Trajectory-Guided Adaptable Bounding Boxes for UAV Detection)
UAVDB is a high-resolution RGB video dataset meticulously designed for UAV detection tasks across diverse scales and complex backgrounds.
1 paper · 1 benchmark
The UAVVaste dataset consists to date of 772 images and 3716 annotations.
1 paper · 1 benchmark
https://github.com/zzr-idam/Under-Display-Camera-UAV
1 paper · 0 benchmarks
UCCS (UnConstrained College Students Dataset)
Unconstrained Face Detection and Open-Set Face Recognition Challenge Paper: https://arxiv.org/abs/1708.02337 Official website: https://vast.uccs.edu/Opensetface UnOfficial website: https://exposing.ai/uccs Face detection and recognition…
1 paper · 0 benchmarks
The VIriors Action Recognition Challenge uses a subset of the UCF101 action recognition dataset: Train set: ~4.8K clips.
1 paper · 0 benchmarks
Existing benchmark datasets in real-world distribution shifts are generally synthetically generated via augmentations to simulate real-world shifts such as weather and camera rotation.
1 paper · 0 benchmarks
UCF50 is an action recognition data set with 50 action categories, consisting of realistic videos taken from youtube.
1 paper · 0 benchmarks
The SMS Spam Collection is a public set of SMS labeled messages that have been collected for mobile phone spam research.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
UCSF PDGM (The University of California San Francisco Preoperative Diffuse Glioma MRI Dataset)
MRI-based artificial intelligence (AI) research on patients with brain gliomas has been rapidly increasing in popularity in recent years in part due to a growing number of publicly available MRI datasets.
1 paper · 0 benchmarks
UDA-CH (Unsupervised Domain Adaptation on Cultural Heritage)
UDA-CH contains 16 objects that cover a variety of artworks which can be found in a museum like sculptures, paintings and books.
1 paper · 1 benchmark
Different from the setting of domain adaptation which uses all labeled source and unlabeled target domain examples for training, domain examples should be divided into two disjoint parts: training and test.
1 paper · 1 benchmark
Different from the setting of domain adaptation which uses all labeled source and unlabeled target domain examples for training, domain examples should be divided into two disjoint parts: training and test.
1 paper · 1 benchmark
We introduce a set of 425 panoramic X-rays with Human annotated Bounding Boxes and Polygons, the 425 images are a subset of UFBA-UESC Dental Dataset.
1 paper · 1 benchmark
The UFPR-ADMR-v2 dataset contains 5,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4M consuming units in the Brazilian state of Paraná.
1 paper · 0 benchmarks
The UFPR-VCR dataset contains 10,039 images of 9,502 distinct vehicles across various categories, including cars, vans, buses, and trucks.
1 paper · 0 benchmarks
Contains three difficult real-world scenarios: uncontrolled videos taken by UAVs and manned gliders, as well as controlled videos taken on the ground.
1 paper · 0 benchmarks
UHCSDB (Ultrahigh Carbon Steel micrograph DataBase)
DeCost, Hecht, Francis, Webler, Picard, and Holm.
1 paper · 0 benchmarks
UHGEvalDataset contains over 5000 news items.
1 paper · 0 benchmarks
UICaption is a dataset of 114k UI images paired with descriptions of their functionality.
1 paper · 0 benchmarks
UIT-ViCoQA (Conversational machine reading comprehension in the Vietnamese language)
UIT-ViCoQA is a new corpus for conversational machine reading comprehension in the Vietnamese language.
1 paper · 0 benchmarks
This dataset comprises over 26,000 full names annotated with genders.
1 paper · 0 benchmarks
Overview: This dataset encompasses a compilation of 6,700 executed scoops (excavations), mapped across a vast spectrum of materials, terrain topography, and compositions.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.