Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 193 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9217–9264 of 12,172
MTA-KDD'19 (Malware Traffic Analysis Knowledge Dataset 2019)
Malware Traffic Analysis Knowledge Dataset 2019 (MTA-KDD'19) is an updated and refined dataset specifically tailored to train and evaluate machine learning based malware traffic analysis algorithms.
1 paper · 0 benchmarks
MTC is a financial-domain dataset of the multi-label topic classification task.
1 paper · 0 benchmarks
MTNeuro is a multi-task neuroimaging benchmark built on volumetric, micrometer-resolution X-ray microtomography images spanning a large thalamocortical section of mouse brain, encompassing multiple cortical and subcortical regions.
1 paper · 0 benchmarks
MTTN is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc.
1 paper · 0 benchmarks
Periodic Tic sounds (T0=1s) sampled at 16kHz with duration of nearly 10s.
1 paper · 0 benchmarks
MUNO21 is a large-scale and comprehensive dataset for the map update task.
1 paper · 0 benchmarks
MUSIED is a large-scale Chinese event detection dataset based on user reviews, text conversations, and phone conversations in a leading e-commerce platform for food service, designed for event detection tasks.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
MVB (Multi View Baggage) is a dataset for baggage ReID task which has some essential differences from person ReID.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MVME (Multi-View Medical Evaluation Benchmark)
The benchmark assesses the real-time interactive consultation capabilities of LLMs across three critical dimensions.
1 paper · 0 benchmarks
Contains about 1, 000 videos from 10 queries and their video tags, manual annotations, and associated web images.
1 paper · 0 benchmarks
MVTec-AC is a curated refinement of the widely-used MVTec-AD dataset, specifically designed for anomaly classification—distinguishing between different types of anomalies rather than merely detecting if an image is anomalous.
1 paper · 2 benchmarks
MVTec-FS (MVTec few-shot detection and classfication dataset)
The MVTec-FS dataset is a refined version of the MVTec AD dataset, designed for few-shot learning research.
1 paper · 0 benchmarks
MVX incorporates realistic physical world simulation with a differentiable accurate ray tracing wireless simulation that includes multi-agent and multimodal datasets for AI-driven digital twin applications in vehicular communication…
1 paper · 1 benchmark
Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature.
1 paper · 0 benchmarks
If you plan to test your method on our road network, you can find road network files in env\map.
1 paper · 0 benchmarks
MacaquePose is an animal pose estimation dataset containing pictures of macaque monkeys and manually labeled annotations on them.
1 paper · 1 benchmark
Dataset of 1,517,419 quantum reaction rate constant products kQM(T)QR(T) computed from the transmission coefficient for model single and double barrier minimum energy paths.
1 paper · 0 benchmarks
Dataset used in "Machine Learning for Analyzing Atomic Force Microscopy (AFM) Images Generated from Polymer Blends".
1 paper · 0 benchmarks
This dataset is used to train and evaluate models for the detection of machine-paraphrased text.
1 paper · 0 benchmarks
Dataset introduction There are four dimension in MBTI.
1 paper · 0 benchmarks
A collection of over 700 games of Mafia, in which players are randomly assigned either deceptive or non-deceptive roles and then interact via forum postings.
1 paper · 0 benchmarks
These are the games used for testing models in the paper Aligning Superhuman AI with Human Behavior Chess as a Model System.
1 paper · 0 benchmarks
Maintenance of Wakefulness Test (MWT) is a dataset of recordings with microsleep episodes and drowsiness.
1 paper · 0 benchmarks
ML-ready Global Dataset of elevation map.
1 paper · 0 benchmarks
Makeup216 contains a variety and representation of logo (captured from the real world) and is among the largest and most complex logo datasets in the field.
1 paper · 0 benchmarks
This is a dataset of Internet malicious activity (mal-activity in short).
1 paper · 0 benchmarks
MalVis (MalVis: A Large-Scale Android Malware Visualization Dataset and Framework for Improved Classification)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MalayalamMixSentiment is a Sentiment Analysis Dataset for Code-Mixed Malayalam-English.
1 paper · 0 benchmarks
The malnutrition data, from the United Nations Children's Fund data warehouse, include two variables, stunted growth and the prevalence of low birth weight, collected in 77 countries from 1985 to 2019.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The Norwegian Historical Data Centre, 2021, "Manually annotated 3-digit occupation code training set from the Norwegian 1950 census", https://doi.org/10.18710/7JWAZX, DataverseNO, V1
1 paper · 0 benchmarks
Manually annotated 3-digit occupation codes from the Norwegian full count 1950 population census.
1 paper · 0 benchmarks
MapEval contains 700 question-answer pairs.
1 paper · 0 benchmarks
MapEval-Textual contains 300 question-answer pairs.
1 paper · 1 benchmark
MapEval-Textual contains 300 context-question-answer triplets.
1 paper · 1 benchmark
MapEval-Visual contains 400 image-question-answer triplets.
1 paper · 1 benchmark
MapReader in GeoHumanities workshop (SIGSPATIAL 2022): Gold standards and outputs Refer to: https://github.com/Living-with-machines/MapReader/wiki/GeoHumanities-workshop-in-SIGSPATIAL-2022
1 paper · 0 benchmarks
This dataset was developed within an analysis of research data generated and managed within the University of Bologna, with respect to the differences and commonalities between disciplines and potential challenges for institutional data…
1 paper · 0 benchmarks
This dataset accompanies the ICWSM 2022 paper "Mapping Topics in 100,000 Real-Life Moral Dilemmas".
1 paper · 0 benchmarks
Marine Microalgae Detection in Microscopy Images dataset contains a total number of images in the dataset is 937 and all the objects in these images were annotated.
1 paper · 0 benchmarks
Describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains.
1 paper · 0 benchmarks
This dataset is useful for doing research in the field of mars surface monocular depth estimation.
1 paper · 1 benchmark
It contains grayscale mono and stereo images (NavCam and LocCam) from laboratory tests performed by a prototype rover on a martian-like testbed.
1 paper · 0 benchmarks
This dataset is a collection of marxist fragments mixed and cut randomly from the Marxist archive (marxists.org).
1 paper · 0 benchmarks
We present MatSci-NLP, a natural language benchmark for evaluating the performance of natural language processing (NLP) models on materials science text.
1 paper · 0 benchmarks
MatSeg (Dataset for Zero-Shot Material States Segmentation)
MatSeg Dataset for Zero-Shot Material States Segmentation: The dataset contains large-scale synthetic images for training data and highly diverse real-world image benchmarks for testing.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.