Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 116 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 5521–5568 of 12,172
A new resource to train and evaluate multitask systems on samples in multiple modalities and three languages.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 2 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
The main goal of the data collection is to acquire highly natural conversations that cover a wide variety of styles and scenarios.
3 papers · 2 benchmarks
MMDB (Multimodal Dyadic Behavior)
Multimodal Dyadic Behavior (MMDB) dataset is a unique collection of multimodal (video, audio, and physiological) recordings of the social and communicative behavior of toddlers.
3 papers · 0 benchmarks
Contains 25,165 textual news articles collected from hundreds of news media sites (e.g., Yahoo News, Google News, CNN News.) and 76,516 image posts shared on Flickr social media, which are annotated according to 412 real-world events.
3 papers · 0 benchmarks
MMID (Massively Multilingual Image Dataset)
A large-scale multilingual corpus of images, each labeled with the word it represents.
3 papers · 0 benchmarks
Includes challenging sequences and extensive data stratification in-terms of camera and object motion, velocity magnitudes, direction, and rotational speeds.
3 papers · 0 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
The pioneering eyeblink detection dataset is characterized by three key features: (1) Sample with multi-human instances.
3 papers · 2 benchmarks
MRPB 1.0 is a mobile robot local planning benchmark.
3 papers · 0 benchmarks
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
Open Dataset: Mobility Scenario FIMU An open, multidimensional (6 categorical attributes), and synthetic dataset of faked virtual humans generated by an optimization approach applied to a real-life call-detail-records-based anonymized…
3 papers · 0 benchmarks
MSR-VTT Adverbs is a subset from MSR-VTT with extracted verb-adverb annotations.
3 papers · 2 benchmarks
Frame-to-frame video alignment/synchronization
3 papers · 1 benchmark
MUC-4 (Fourth Message Uunderstanding Conference)
A dataset for evaluate system's understanding of given passages.
3 papers · 1 benchmark
MULTI-Benchmark is a cutting-edge benchmark for evaluating Multimodal Large Language Models (MLLMs).
3 papers · 0 benchmarks
MUTE (Multimodal Bengali Hateful Memes Dataset)
MUTE This is the first open-source Bengali Hateful Meme dataset, consisting of around 4200 memes annotated with two labels: hate and not hate.
3 papers · 0 benchmarks
MVHand is a new multi-view hand posture dataset to obtain complete 3D point clouds of the hand in the real world.
3 papers · 0 benchmarks
MaleX is a curated dataset of malware and benign Windows executable samples for malware researchers.
3 papers · 0 benchmarks
Marine Debris Turntable is a dataset for sonar perception.
3 papers · 0 benchmarks
MeLa BitChute is a near-complete dataset of over 3M videos from 61K channels over 2.5 years (June 2019 to December 2021) from the social video hosting platform BitChute, a commonly used alternative to YouTube.
3 papers · 0 benchmarks
Medical Question Pairs (MQP) Dataset This repository contains a dataset of 3048 similar and dissimilar medical question pairs hand-generated and labeled by Curai's doctors.
3 papers · 0 benchmarks
The “Medico automatic polyp segmentation challenge” aims to develop computer-aided diagnosis systems for automatic polyp segmentation to detect all types of polyps (for example, irregular polyp, smaller or flat polyps) with high efficiency…
3 papers · 1 benchmark
The MeltingTemp dataset is collected from Polyinfo.
3 papers · 0 benchmarks
A large-scale dataset of memes with captions and class labels.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
MineRLis an imitation learning dataset with over 60 million frames of recorded human player data.
3 papers · 0 benchmarks
There are other gridworld Gym environments out there, but this one is designed to be particularly simple, lightweight and fast.
3 papers · 0 benchmarks
Mirrored-Human is a dataset for 3D pose estimation from a single view.
3 papers · 0 benchmarks
Contains mitochondria instances.
3 papers · 1 benchmark
MobiBits (Multimodal Mobile Biometric Database)
A novel database comprising representations of five different biometric characteristics, collected in a mobile, unconstrained or semi-constrained setting with three different mobile devices, including characteristics previously unavailable…
3 papers · 0 benchmarks
We sample 2025 frames of images from the original KITTI for Mono3DRefer, containing 41,140 expressions in total and a vocabulary of 5,271 words.
3 papers · 0 benchmarks
The dataset, comprising 1204 meticulously curated images, serves as a comprehensive resource for advancing real-time mosquito detection models.
3 papers · 0 benchmarks
This dataset contains a large set (~3.2 Million) of high quality expert trajectories generated from a geometrically consist hybrid planner in a wide variety of environment (~575,000 environments).
3 papers · 0 benchmarks
In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Multi-CPR (Multi Domain Chinese Dataset for Passage Retrieval)
Multi-CPR is a multi-domain Chinese dataset for passage retrieval.
3 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
MultiQ is a multi-hop QA dataset for Russian, suitable for general open-domain question answering, information retrieval, and reading comprehension tasks.
3 papers · 1 benchmark
Multilingual TOP is a dataset for multilingual semantic parsing with human-written sentences as opposed to machine translated ones.
3 papers · 0 benchmarks
Contains squared blocks of 48×48 pixels including 13 Sentinel-2 bands.
3 papers · 1 benchmark
New refined labels for the MusicNet dataset obtained by the EM process as described in the paper: Ben Maman and Amit Bermano, "Unaligned Supervision for Automatic Music Transcription in The Wild"
3 papers · 0 benchmarks
Named Entity (NER) annotations of the Hebrew Treebank (Haaretz newspaper) corpus, including: morpheme and token level NER labels, nested mentions, and more.
3 papers · 3 benchmarks
X-rays, CT Images and Genomic Sequences representing cases of tuberculosis.
3 papers · 0 benchmarks
NILUT (NILUT 3D LUT Dataset and enhanced images from MIT5K)
Read all the details about the dataset in our paper "NILUT: Conditional Neural Implicit 3D Lookup Tables for Image Enhancement" We host the dataset in Kaggle: https://www.kaggle.com/datasets/photolab/nilut-3d-lut-dataset More information…
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.