Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 125 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5953–6000 of 12,172
Amazon-PQA is a product question-answer dataset.
2 papers · 0 benchmarks
Ambiguous-HOI is a challenging dataset containing ambiguous human-object interaction images for HOI detection based on HICO-DET.
2 papers · 0 benchmarks
About Dataset Step right up to our AI data collection company, where we’ve got something special just for you: a unique set of American Sign Language datasets!
2 papers · 0 benchmarks
Amharic - English Parallel Corpus for Machine Translation contains 33,955 sentence pairs extracted text from such news platforms as Ethiopian Press Agency1, Fana Broadcasting Corporate2, and Walta Information Center3.
2 papers · 0 benchmarks
https://github.com/salesforce/xnliextension
2 papers · 0 benchmarks
Machine learning is transforming the video editing industry.
2 papers · 0 benchmarks
We present a novel Animation CelebHeads dataset (AnimeCeleb) to address an animation head reenactment.
2 papers · 0 benchmarks
ApisTox contains molecules in SMILES format for predicting pesticides toxicity to honey bees.
2 papers · 0 benchmarks
Our trajectory dataset consists of camera-based images, LiDAR scanned point clouds, and manually annotated trajectories.
2 papers · 1 benchmark
A new underwater dataset that has been recorded in an harbor and provides several sequences with synchronized measurements from a monocular camera, a MEMS-IMU and a pressure sensor.
2 papers · 0 benchmarks
Natural Language Inference processes pairs of sentences to extract their semantic relations.
2 papers · 0 benchmarks
Sentiment detection remains a pivotal task in natural language processing, yet its development in Arabic lags due to a scarcity of training materials compared to English.
2 papers · 0 benchmarks
AraCOVID19-MFH (AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset)
AraCOVID19-MFH is a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
2 papers · 0 benchmarks
ARCENE was obtained by merging three mass-spectrometry datasets to obtain enough training and test data for a benchmark.
2 papers · 0 benchmarks
The Argoverse 2 Map Change Dataset is a collection of 1,000 scenarios with ring camera imagery, lidar, and HD maps.
2 papers · 0 benchmarks
A real-world dataset, with hyper-accurate digital counterpart & comprehensive ground-truth annotation.
2 papers · 1 benchmark
The ArxivPapers dataset is an unlabelled collection of over 104K papers related to machine learning and published on arXiv.org between 2007–2020.
2 papers · 0 benchmarks
The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems.
2 papers · 0 benchmarks
Audio-alpaca: A preference dataset for aligning text-to-audio models Audio-alpaca is a pairwise preference dataset containing about 15k (prompt,chosen, rejected) triplets where given a textual prompt, chosen is the preferred generated…
2 papers · 0 benchmarks
The Aurora-2 data are based on a version of the original TIDigits (as available from LDC) downsampled at 8 kHz.
2 papers · 0 benchmarks
AutoLand (An Autonomous UAV Navigation and Landing System for Urban Search and Rescue Missions)
To faciliate training of neural networks and evaluation of alternate approaches for landing, we provide a synthetic dataset comprised of collapsed buildings.
2 papers · 0 benchmarks
This is dataset for A TIME SERIES IS WORTH 64 WORDS: LONG-TERM FORECASTING WITH TRANSFORMERS We evaluate the performance of our proposed PatchTST on 8 popular datasets, including Weather, Traffic, Electricity, ILI and 4 ETT datasets…
2 papers · 0 benchmarks
Global Symmetry Ground-truth for AVA dataset.
2 papers · 0 benchmarks
Avocado Research Email Collection consists of emails and attachments taken from 279 accounts of a defunct information technology company referred to as "Avocado".
2 papers · 0 benchmarks
Contains stock market closing prices of ten financial institutions.
2 papers · 0 benchmarks
BASEPROD (The Bardenas Semi-Desert Planetary Rover Dataset)
BASEPROD provides comprehensive rover sensor data collected over a 1.7 km traverse, accompanied by high-resolution 2D and 3D drone maps of the terrain.
2 papers · 0 benchmarks
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
Blender Cycles Ray-tracing (BCR) dataset contains 2449 high-quality images rendered from 1463 models.
2 papers · 0 benchmarks
BD-4SK-ASR (Basic Dataset for Sorani Kurdish Automatic Speech Recognition)
The Basic Dataset for Sorani Kurdish Automatic Speech Recognition (BD-4SK-ASR) is a dataset for automatic speech recognition for Sorani Kurdish.
2 papers · 0 benchmarks
Subsets of BDD100K Dataset that are used in Object Detection Under Rainy Conditions for Autonomous Vehicles: A Review of State-of-the-Art and Emerging Techniques
2 papers · 1 benchmark
BEE23 (Multi-bee Tracking Benchmark)
We collected 32 videos that record bee colony activity from different periods on several sunny days.
2 papers · 0 benchmarks
BGVP (BG Vulnerable Pedestrian)
BG Vulnerable Pedestrian (BGVP) is a dataset to help train well-rounded models and thus induce research to increase the efficacy of vulnerable pedestrian detection.
2 papers · 0 benchmarks
BN-HTRd (BN-HTRd: A Benchmark Dataset for Document Level Offline Bangla Handwritten Text Recognition (HTR))
We introduce a new Dataset (BN-HTRd) for offline Handwritten Text Recognition (HTR) from images of Bangla scripts comprising words, lines, and document-level annotations.
2 papers · 2 benchmarks
BS-RSCD is a dataset for rolling shutter correction and deblurring (RSCD).
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
BVI-LOWLIGHT (BVI-LOWLIGHT: FULLY REGISTERED DATASETS FOR LOW-LIGHT IMAGE AND VIDEO ENHANCEMENT)
Low-light images and video footage often exhibit issues due to the interplay of various parameters such as aperture, shutter speed, and ISO settings.
2 papers · 0 benchmarks
BabySLM is a language-acquisition-friendly benchmark to probe speech-based LMs at the lexical and syntactic levels, both of which are compatible with the vocabulary typical of children's language experiences.
2 papers · 0 benchmarks
A Dataset to Identify Manipulated Social Media News in Bangla We construct a publicly available Bangla dataset of 800 news-related social media items that are annotated as manipulated or not relative to 500 reference news articles.
2 papers · 0 benchmarks
Named entities in Bavarian text Details: Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm, Verena Blaschke, Ekaterina Artemova, and Barbara Plank.
2 papers · 0 benchmarks
A set of basque documents annotated with EusTimeML - a mark-up language for temporal information in Basque.
2 papers · 1 benchmark
BdSLImset (Bangladeshi Sign Language Image Dataset)
Bangladeshi Sign Language Image Dataset (BdSLImset) is a dataset that contains images of different Bangladeshi sign letters.
2 papers · 0 benchmarks
Dataset of the Beacon3D benchmark: Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis.
2 papers · 0 benchmarks
Contains the graph benchmark sets (regular set and irregular set) and experimental results.
2 papers · 0 benchmarks
Description The Berkeley Multimodal Human Action Database (MHAD) contains 11 actions performed by 7 male and 5 female subjects in the range 23-30 years of age except for one elderly subject.
2 papers · 0 benchmarks
The Berlin V2X dataset offers high-resolution GPS-located wireless measurements across diverse urban environments in the city of Berlin for both cellular and sidelink radio access technologies, acquired with up to 4 cars over 3 days.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.