Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 196 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9361–9408 of 12,172
Dataset (part 1/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach.
1 paper · 0 benchmarks
Dataset (part 2/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach.
1 paper · 0 benchmarks
Dataset (part 3/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach.
1 paper · 0 benchmarks
Dataset for User Verification part of MotionID: Human Authentication Approach.
1 paper · 0 benchmarks
From dataset repository for "2020 International BCI Competition": https://osf.io/pq7vb/?viewonly=08e7108d89fd42bab2adbd6b98fb683d
1 paper · 0 benchmarks
This dataset was generated to characterize mouse grooming behavior.
1 paper · 0 benchmarks
A large, annotated video dataset of mice performing a sequence of actions.
1 paper · 0 benchmarks
Movie Reviews (Movie Review Polarity Dataset Enriched with "Annotator Rationales")
This dataset is based on the movie review polarity dataset (v2.0) collected and maintained by Bo Pang and Lillian Lee.
1 paper · 0 benchmarks
MovieCLIP is a movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets designed for visual scene recognition.
1 paper · 0 benchmarks
The dataset contains entities from IMDB, TheMovieDB and TheTVDB with goldstandard matches between the sources.
1 paper · 0 benchmarks
MovieNet-TeViS is a synopsis-storyboard pair dataset to facilitate the Text synopsis to Video Storyboard (TeViS).
1 paper · 0 benchmarks
A parameterized synthetic dataset called Moving Symbols to support the objective study of video prediction networks.
1 paper · 0 benchmarks
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, jelly, and rigid collisions.
1 paper · 0 benchmarks
MuCeD, a dataset that is carefully curated and validated by expert pathologists from the All India Institute of Medical Science (AIIMS), Delhi, India.
1 paper · 0 benchmarks
Given an ongoing dialogue between a user and a dialogue assistant, for the user query, the model is required to predict both coreference links between the query and the dialogue context, and the self-contained rewritten user query that is…
1 paper · 0 benchmarks
MuLMS (Multi-Layer Materials Science)
The Multi-Layer Materials Science corpus (MuLMS) consists of 50 documents (licensed CC BY) from the materials science domain, spanning across the following 7 subareas: "Electrolysis", "Graphene", "Polymer Electrolyte Fuel Cell (PEMFC)",…
1 paper · 0 benchmarks
MuLVE (Multi-Language Vocabulary Evaluation)
Multi-Language Vocabulary Evaluation Data Set (MuLVE) is a dataset consisting of vocabulary cards and real-life user answers, labeled indicating whether the user answer is correct or incorrect.
1 paper · 0 benchmarks
This is the large version of the MuMiN dataset.
1 paper · 1 benchmark
This is the medium version of the MuMiN dataset.
1 paper · 1 benchmark
This is the small version of the MuMiN dataset.
1 paper · 1 benchmark
Early detection of retinal diseases is one of the most important means of preventing partial or permanent blindness in patients.
1 paper · 1 benchmark
MuSoHu (Toward human-like social robot navigation: A large-scale, multi-modal, social human navigation dataset)
A large-scale, egocentric, multimodal, and context-aware dataset of human demonstrations of social navigation.
1 paper · 0 benchmarks
A dataset of music videos with continuous valence/arousal ratings as well as emotion tags.
1 paper · 0 benchmarks
Dataset Description The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag.
1 paper · 1 benchmark
This dataset was created as part of the Master's thesis titled "Multi-Class Depression Detection Through Tweets Using Artificial Intelligence." It contains tweets labeled for five types of depression (Bipolar, Major, Psychotic, Atypical,…
1 paper · 0 benchmarks
Multi-CrossRE is a broadest multi-lingual dataset for Relation Extraction (RE) including 26 languages in addition to English, and covering six text domains.
1 paper · 0 benchmarks
A new multi-view egocentric dataset, Multi-Ego.
1 paper · 0 benchmarks
The Multi-Eup is a new multilingual benchmark dataset, comprising 22K multilingual documents collected from the European Parliament, spanning 24 languages.
1 paper · 0 benchmarks
This dataset is a multi-labelled SMILES odor dataset with 138 odor descriptors.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Two separate datasets of calibration runs in front of a calibration board: - 4IMUs+3Cams -4IMUs+4Cams sensor hz topic resolution MicroStrain GX3-25 500hz /gx325/data MicroStrain GX3-25 100hz /gx335/imudata Xsens MTI-100 400hz /imu/data…
1 paper · 0 benchmarks
This dataset contains approximately 25,000 seismic events recorded by Observatorio Vulcanológico Andes Sur (OVDAS, SERNAGEOMIN) of four Chilean volcanoes: Nevados de Chillán Volcanic Complex (NVChVC), Villarrica (VCA), Laguna del Maule…
1 paper · 0 benchmarks
The Multi-domain Image Characteristic Dataset consists of thousands of images sourced from the internet.
1 paper · 0 benchmarks
The database was acquired using a thermographic camera TESTO 882-3 equipped with an uncooled detector and a spectral sensitivity range from 8 to 14 μm.
1 paper · 0 benchmarks
Latent DNA Diffusion Dataset - Author: Zehui Li - Size: 1K<n<10K - License: MIT - Files: - .gitattributes (2.36 kB) - README.md (171 Bytes) - combined.tar.gz (336 MB LFS) - metadata.json (2.42 kB) - sequence.csv (330 MB LFS) -…
1 paper · 0 benchmarks
This dataset consists of four sets of flower images, from three different species: apple, peach, and pear, and accompanying ground truth images.
1 paper · 0 benchmarks
Mouse Brain MRI atlas (both in-vivo and ex-vivo) (repository relocated from the original webpage) List of atlases - FVBNCrl: Brain MRI atlas of the wild-type FVBNCrl mouse strain (used as the background strain for the rTg4510 which is a…
1 paper · 0 benchmarks
BernBypass70 is a dataset consisting of 70 surgical videos of LRYGB at Inselspital, Bern University Hospital, Switzerland.
1 paper · 0 benchmarks
MultiCite is a dataset of 12,653 citation contexts from over 1,200 computational linguistics papers used for Citation context analysis (CCA).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MultiRefKGC is a dataset created from conversations from Reddit designed for Knowledge-Grounded Dialogue Generation tasks.
1 paper · 0 benchmarks
MultiSenseBadminton (MultiSenseBadminton: Wearable Sensor–Based Biomechanical Dataset for Evaluation of Badminton Performance)
The sports industry is witnessing an increasing trend of utilizing multiple synchronized sensors for player data collection, enabling personalized training systems with multi-perspective real-time feedback.
1 paper · 0 benchmarks
MultiSenti presents a labeled dataset called MultiSenti for sentiment classification of code-switched informal short text, (2) explore the feasibility of adapting resources from a resource-rich language for an informal one, and (3) propose…
1 paper · 0 benchmarks
MultiSum is a dataset for multimodal summarization (MSMO).
1 paper · 0 benchmarks
MultiUN (Multilingual Corpus from United Nation Documents)
The MultiUN parallel corpus is extracted from the United Nations Website , and then cleaned and converted to XML at Language Technology Lab in DFKI GmbH (LT-DFKI), Germany.
1 paper · 0 benchmarks
MultiWOZ-coref, (or MultiWOZ 2.3) is an extension of the MultiWOZ dataset that adds co-reference annotations in addition to corrections of dialogue acts and dialogue states.
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845311#.ZK-jty9BxhE
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.