Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 137 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6529–6576 of 12,172
MSRA10K (MSRA10K Salient Object Database)
MSRA10K is a dataset for salient object detection that contains 10,000 images with pixel-level saliency labeling for 10K images from the MSRA salient object detection dataset.
2 papers · 0 benchmarks
the MTHS dataset contains 30Hz PPG signals obtained from 62 patients, including 35 men and 27 women.
2 papers · 2 benchmarks
MTST (Mobile Turkish Scene Text)
The Mobile Turkish Scene Text (MTST 200) dataset consists of 200 indoor and outdoor Turkish scene text images.
2 papers · 0 benchmarks
The MUSE dataset contains bilingual dictionaries for 110 pairs of languages.
2 papers · 2 benchmarks
MUSIC (Multi-Spectral Imaging via Computed Tomography)
The Multi-Spectral Imaging via Computed Tomography (MUSIC) dataset is a two-part (2D- and 3D spectral) open access dataset for advanced image analysis of spectral radiographic (x-ray) scans, their tomographic reconstruction and the…
2 papers · 0 benchmarks
We introduce the first dataset, MUSIC-AVQA-R, to evaluate the robustness of AVQA models.
2 papers · 0 benchmarks
MUStARD (Multimodal Sarcasm Detection Dataset)
We release the MUStARD dataset which is a multimodal video corpus for research in automated sarcasm discovery.
2 papers · 0 benchmarks
This dataset includes time-synchronized multimodal data records of students (learning logs, videos, EEG brainwaves) as they work in various subjects from Squirrel AI Learning System (SAIL) to solve problems of varying difficulty levels.
2 papers · 0 benchmarks
MVSep is a synthetic dataset for the vocal separation task created by combining random vocal and instrumental samples, publicly available on the internet.
2 papers · 0 benchmarks
MVTec D2S (MVTec Densely Segmented Supermarket)
MVTec D2S is a benchmark for instance-aware semantic segmentation in an industrial domain.
2 papers · 0 benchmarks
MaNGA (Mapping Nearby Galaxies at APO)
MaNGA is a component of the Fourth-Generation Sloan Digital Sky Survey whose goal is to map the detailed composition and kinematic structure of nearby galaxies.
2 papers · 0 benchmarks
The E-MASAC Dataset is a collection of code-mixed conversations sourced from an Indian TV series, focusing on Hindi-English interactions.
2 papers · 1 benchmark
Consists of visual arithmetic problems automatically generated using a grammar model--And-Or Graph (AOG).
2 papers · 0 benchmarks
MagicBathyNet is a benchmark dataset made up of image patches of Sentinel-2, SPOT-6 and aerial imagery, bathymetry in raster format and seabed classes annotations.
2 papers · 0 benchmarks
Context Malicious URLs or malicious website is a very serious threat to cybersecurity.
2 papers · 0 benchmarks
This dataset contains 1203 individuals captured from two disjoint camera views.
2 papers · 0 benchmarks
All animal procedures are overseen by veterinary staff of the MIT and Broad Institute Department of Comparative Medicine, in compliance with the NIH guide for the care and use of laboratory animals and approved by the MIT and Broad…
2 papers · 1 benchmark
MathBridge (MathBridge: A Large Corpus Dataset for Translating Spoken Mathematical Expressions into LaTeX Formulas for Improved Readability)
Understanding sentences that contain mathematical expressions in text form poses significant challenges.
2 papers · 0 benchmarks
Our primary objective in creating this dataset is to support researchers in the advancement of algorithms for keypoints detection and the pretraining of large models on retinal images using a self-supervised approach.
2 papers · 0 benchmarks
Each dataset in the Mechanical MNIST collection contains the results of 70,000 (60,000 training examples + 10,000 test examples) finite element simulation of a heterogeneous material subject to large deformation.
2 papers · 0 benchmarks
The Mechanical MNIST Crack Path dataset contains Finite Element simulation results from phase-field models of quasi-static brittle fracture in heterogeneous material domains subjected to prescribed loading and boundary conditions.
2 papers · 0 benchmarks
MedMNIST-C is an open-source data set collection comprising algorithmically generated corruptions applied to the test sets of the MedMNIST collection following the concept of ImageNet-C.
2 papers · 0 benchmarks
The process by which sections in a document are demarcated and labeled is known as section identification.
2 papers · 2 benchmarks
MedleyVox is an evaluation dataset for multiple singing voices separation that corresponds to such categories.
2 papers · 0 benchmarks
The MegaNegRaising dataset, also known as MegaNeRd, is a collection of data that captures patterns of neg-raising inferences and acceptability judgments for 925 clause-embedding verbs of English in various syntactic structures.
2 papers · 0 benchmarks
Results of a high-throughput biological assay measuring the stability of proteins https://github.com/Rocklin-Lab/cdna-display-proteolysis-pipeline From the paper "Here we present cDNA display proteolysis, a method for measuring…
2 papers · 0 benchmarks
MentSum (Mental Health Summarization Dataset)
Mental health remains a significant challenge of public health worldwide.
2 papers · 1 benchmark
MAUD is an expert-annotated merger agreement reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points study, where lawyers and law students answered 92 questions about 152 merger agreements.
2 papers · 0 benchmarks
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
This dataset contains the extraction made in 2022 of all the 622 datasets that existed then at the UCI Machine Learning Repository.
2 papers · 0 benchmarks
The Metaphorical Connections dataset is a poetry dataset that contains annotations between metaphorical prompts and short poems.
2 papers · 0 benchmarks
Middlebury 2003 is a stereo dataset for indoor scenes.
2 papers · 0 benchmarks
Middlebury MVS is the earliest MVS dataset for multi-view stereo network evaluation.
2 papers · 0 benchmarks
Mila Simulated Floods Dataset is a 1.5 square km virtual world using the Unity3D game engine including urban, suburban and rural areas.
2 papers · 1 benchmark
Minecraft House is a crowd sourced dataset that collects examples of humans building houses in Minecraft.
2 papers · 0 benchmarks
MiniWob++ is a suite of web-browser based tasks introduced in Liu et al.
2 papers · 0 benchmarks
MoB (Malicious or Benign Cartoon Videos)
A dataset of cartoon video clips.
2 papers · 1 benchmark
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
A large scale OCSR dataset, proposed in paper “MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild“ MolParser-7M contains nearly 8 million paired image-SMILES data.
2 papers · 0 benchmarks
Dates with Moon phases extended days until next phase (1992/1/4 to 2027/12/20) Incorporate lunar data to your research.
2 papers · 0 benchmarks
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
This dataset consists of blurred, noisy and defocused images.
2 papers · 0 benchmarks
A version of the CMU Movie Summary Corpus (http://www.cs.cmu.edu/~ark/personas/), which was originally scraped from plot summaries from Wikipedia, with some cleaning and sentences turned into events & sorted into "genres" (via LDA).
2 papers · 0 benchmarks
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MuViHand is a dataset for 3D Hand Pose Estimation that consists of multi-view videos of the hand along with ground-truth 3D pose labels.
2 papers · 0 benchmarks
Multi Task Crowd is a new 100 image dataset fully annotated for crowd counting, violent behaviour detection and density level classification.
2 papers · 0 benchmarks
This is a multi-codec DASH dataset comprising AVC, HEVC, VP9, and AV1 in order to enable interoperability testing and streaming experiments for the efficient usage of these codecs under various conditions.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.