Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 136 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 6481–6528 of 12,172
Usually, the information related to the crop types available in a given territory is annual information, that is, we only know the type of main crop grown over a year and we do not know any crops that have followed one another during the…
2 papers · 1 benchmark
We randomly selected three videos from the Internet, that are longer than 1.5K frames and have their main objects continuously appearing.
2 papers · 1 benchmark
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
LuViRA (Lund University Vision, Radio, and Audio)
The Lund University Vision, Radio, and Audio (LuViRA) positioning dataset consists of 89 trajectories that are recorded in the Lund University Humanities Lab's Motion Capture (Mocap) Studio using a MIR200 robot as the targeted platform.
2 papers · 0 benchmarks
LymphoMNIST is a comprehensive dataset designed for the nuanced classification of lymphocyte images.
2 papers · 0 benchmarks
Lyra Dataset (A Dataset for Greek Traditional and Folk Music)
Lyra is a dataset of 1570 traditional and folk Greek music pieces that includes audio and video (timestamps and links to YouTube videos), along with annotations that describe aspects of particular interest for this dataset, including…
2 papers · 0 benchmarks
M-PCCD (MPEG Point Cloud Compression Dataset)
The emerging MPEG point cloud codecs (V-PCC and G-PCC variants) are assessed, and best practices for rate allocation are investigated [1].
2 papers · 1 benchmark
M2KR (Multi-task Multi-modal Knowledge Retrieval)
The M2KR is a collection of datasets designed for training and evaluating general-purpose vision-language retrievers.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
The Model for Attended Awareness in Driving (MAAD) is a dataset of third-person estimates of a driver’s attended awareness.
2 papers · 0 benchmarks
MAMe (Museum Art Medium dataset)
The MAMe dataset contains images of high-resolution and variable shape of artworks from 3 different museums: - The Metropolitan Museum of Art of New York - The Los Angeles County Museum of Art - The Cleveland Museum of Art Source:…
2 papers · 1 benchmark
MAPS-KB is a million-scale probabilistic simile knowledge base, covering 4.3 million triplets over 0.4 million terms from 70 GB corpora.
2 papers · 0 benchmarks
MASRI-HEADSET is a corpus that was developed by the MASRI project at the University of Malta.
2 papers · 0 benchmarks
MAVS (Multilingual Audio-Visual Smartphone dataset)
MAVS is an audio-visual smartphone dataset captured in five different recent smartphones.
2 papers · 0 benchmarks
Dataset page: https://github.com/mosamdabhi/MBW-Data MBW - Zoo is a challenging dataset consisting image frames of tail-end distribution categories (such as Fish, Colobus Monkeys, Chimpanzees, etc.) with their corresponding 2D, 3D, and…
2 papers · 0 benchmarks
A dataset of around 2 million examples for machine reading-comprehension.
2 papers · 0 benchmarks
MCAD (Multi-Camera Action Dataset)
Designed to evaluate the open view classification problem under the surveillance environment.
2 papers · 0 benchmarks
MCSCSet is a large-scale specialist-annotated dataset, designed for the task of Medical-domain Chinese Spelling Correction that contains about 200k samples.
2 papers · 0 benchmarks
MCSI (Mpox Close Skin Images)
The Mpox Close Skin Images dataset (MCSI) is a collection of skin images obtained from diverse public sources, that we accurately pre-processed (i.e., cropped and zoomed) in order to focus the skin lesion (if present), and to evaluate…
2 papers · 0 benchmarks
MD Gender (Multi-Dimensional Gender Bias Datasets)
Provides eight automatically annotated large scale datasets with gender information.
2 papers · 0 benchmarks
MDD (Movie Dialog dataset)
Movie Dialog dataset (MDD) is designed to measure how well models can perform at goal and non-goal orientated dialog centered around the topic of movies (question answering, recommendation and discussion).
2 papers · 0 benchmarks
MECD (Multi-Event Causal Discovery)
Provide: 1,105 lifestyle videos that span diverse scenarios.
2 papers · 1 benchmark
MFW+ is a benchmark dataset for masked face recognition and an extended version of MFW.
2 papers · 1 benchmark
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
This repository provides a cleaned dataset, which is intended to be used for text classification, language modeling, and AI-generated content detection tasks.
2 papers · 0 benchmarks
MHSMA (The Modified Human Sperm Morphology Analysis)
The MHSMA dataset is a collection of human sperm images from 235 patients with male factor infertility.
2 papers · 0 benchmarks
MIAD contains more than 100K high-resolution color images in various outdoor industrial scenarios, designed for unsupervised anomaly detection.
2 papers · 0 benchmarks
You need to request access to download and use the dataset.
2 papers · 1 benchmark
This database is provided and maintained by Dr.
2 papers · 1 benchmark
MIDGARD is an open-source simulator for autonomous robot navigation in outdoor unstructured environments.
2 papers · 0 benchmarks
MIMIC II (Multi-parameter Intelligent Monitoring for Critical Care II database)
The data used in this research is a subset of the Multi-parameter Intelligent Monitoring for Critical Care (MIMIC) II database.
2 papers · 0 benchmarks
The MIMIC PERform Testing dataset contains the following physiological signals recorded from 200 critically-ill patients during routine clinical care: - electrocardiogram (ECG) - photoplethysmogram (PPG) - impedance pneumography (imp),…
2 papers · 2 benchmarks
To support the machine learning (ML) community in developing a time-cost-effective diagnostic assistant, we collaborate with ED clinicians to curate a benchmark, called MIMIC-ED-Assist, that is derived from MIMIC-IV and MIMIC-ED.
2 papers · 0 benchmarks
A dataset of 21 WSIs of CMC completely annotated for MF.
2 papers · 0 benchmarks
Persian-English parallel corpus with more than one million sentence pairs collected from masterpieces of literature.
2 papers · 0 benchmarks
MLP (Multimodal Lecture Presentations)
Multimodal Lecture Presentations (MLP) is a large-scale benchmark dataset for testing the capabilities of machine learning models in multimodal understanding of educational content.
2 papers · 0 benchmarks
We provide a dataset called MMAC Captions for sensor-augmented egocentric-video captioning.
2 papers · 0 benchmarks
MMSD2.0 (Towards a Reliable Multi-modal Sarcasm Detection System)
Multi-modal sarcasm detection has attracted much recent attention.
2 papers · 0 benchmarks
MMVR (Millimeter-wave Multi-View Radar (MMVR) Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
MMVax-Stance includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
MOMAland is an open source Python library for developing and comparing multi-objective multi-agent reinforcement learning algorithms by providing a standard API to communicate between learning algorithms and environments, as well as a…
2 papers · 0 benchmarks
MOTFront provides photo-realistic RGB-D images with their corresponding instance segmentation masks, class labels, 2D & 3D bounding boxes, 3D geometry, 3D poses and camera parameters.
2 papers · 0 benchmarks
MPHOI-72 (Multi-person Human-object Interaction Dataset 72)
MPHOI-72 is a multi-person human-object interaction dataset that can be used for a wide variety of HOI/activity recognition and pose estimation/object tracking tasks.
2 papers · 0 benchmarks
A multi-domain question rewriting dataset is constructed from human contributed Stack Exchange question edit histories.
2 papers · 0 benchmarks
MS-BioGraphs (MS-BioGraphs: Sequence SImilarity Graph Datasets)
https://doi.org/10.21227/gmd9-1534
2 papers · 0 benchmarks
MSC is a dataset for Macro-Management in StarCraft 2 based on the platfrom SC2LE.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
Multi-modal situated reasoning in 3D scenes
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.