Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 195 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9313–9360 of 12,172
There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks.
1 paper · 0 benchmarks
There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks.
1 paper · 0 benchmarks
There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks.
1 paper · 0 benchmarks
The dataset utilized in this study consists of demographic, clinical, and laboratory data for 2,401 individuals.
1 paper · 0 benchmarks
Metaphorics is a newly introduced non-contextual skeleton action dataset.
1 paper · 0 benchmarks
Metric-Type of Numerical Tables is a dataset extracted from scientific papers (ACL anthology website) consisting of header tables, captions, and metric-types.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MiMIC (Multi-Modal Indian Earnings Calls Dataset)
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources.
1 paper · 0 benchmarks
MiSCS (Microscopic Shrub Cross Sections)
Microscopy images of shrub cross sections for instance segmentation of tree rings.
1 paper · 0 benchmarks
MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.
1 paper · 0 benchmarks
Pulmonary hypertension (PH) is a syndrome complex that accompanies a number of diseases of different etiologies, associated with basic mechanisms of structural and functional changes of the pulmonary circulation vessels and revealed…
1 paper · 0 benchmarks
Microscopy Images of the Drosophila Wing dataset are divided into two folders, Tumor/ No Tumor.
1 paper · 0 benchmarks
This dataset contains annotations for 5000 music files on the following music properties: Melodiousness Articulation Rhythmic stability Rhythmic complexity Dissonance Tonal stability Modality The annotations were given by musicians and…
1 paper · 0 benchmarks
Experiments on a metal milling machine for different speeds, feeds, and depth of cut.
1 paper · 0 benchmarks
MindReader is a novel dataset providing explicit user ratings over a knowledge graph within the movie domain.
1 paper · 0 benchmarks
MinNav is a synthetic dataset based on the sandbox game Minecraft.
1 paper · 0 benchmarks
Minecraft Segmentation is a segmentation dataset for the Minecraft House that adds semantic segmentation labels for sub-components of the house.
1 paper · 0 benchmarks
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
The MiniHAREM, a reiteration of the 2005 evaluation, used the same methodology and platform.
1 paper · 0 benchmarks
Minsk2019 ALS database is a dataset collected in Republican Research and Clinical Center of Neurology and Neurosurgery (Minsk, Belarus).
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
4 different synthetic datasets generated by Blender
1 paper · 0 benchmarks
MixedWM38 Dataset(WaferMap) has more than 38000 wafer maps, including 1 normal pattern, 8 single defect patterns, and 29 mixed defect patterns, a total of 38 defect patterns.
1 paper · 2 benchmarks
MmCows is a large-scale multimodal dataset for behavior monitoring, health management, and dietary management of dairy cattle.
1 paper · 0 benchmarks
Different types of cells play a vital role in the initiation, development, invasion, metastasis and therapeutic response of tumors of various organs.
1 paper · 1 benchmark
MoToMQA (Multi-Order Theory of Mind Question & Answer)
The MoToMQA (Multi-Order Theory of Mind Question & Answer) benchmark is a test suite introduced to examine the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about…
1 paper · 0 benchmarks
https://huggingface.co/datasets/OpenDFM/MobA-MobBench
1 paper · 0 benchmarks
MobiFace is the first dataset for single face tracking in mobile situations.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset consists of the Graphcast model checkpoints produced during the fine-tuning process of (Subich 2024).
1 paper · 0 benchmarks
Modern Hebrew Sentiment Dataset is a sentiment analysis benchmark for Hebrew, based on 12K social media comments, and provide two instances of these data: in token-based and morpheme-based settings.
1 paper · 0 benchmarks
The Modified Swiss Dwellings (MSD) dataset is an ML-ready dataset for floor plan generation and analysis at building-level scale.
1 paper · 0 benchmarks
The set is created using molecule SMILES retrieved from the database PubChem.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We present MoleculeCLA: a large-scale dataset consisting of approximately 140,000 small molecules derived from computational ligand-target binding analysis, providing nine properties that cover chemical, physical, and biological aspects.
1 paper · 0 benchmarks
The French national meteorological service published an open-access dataset of hourly weather observations in Brittany, France, for the month of January 2014.
1 paper · 0 benchmarks
Mon(IoT)r Testbed The Mon(IoT)r Testbed is the traffic capture software developed for the Mon(IoT)r Lab.
1 paper · 0 benchmarks
This dataset of medical misinformation was collected and is published by Kempelen Institute of Intelligent Technologies (KInIT).
1 paper · 0 benchmarks
This dataset is used for neural co-training.
1 paper · 0 benchmarks
Monovab (An Annotated Corpus for Bangla Multi-label Emotion Detection)
The most popular news portal's Facebook pages such as Prothom Alo, BBC Bangla, BD News 24, Bangla Tribune, Kaler Kantho, Daily Jugantor are picked to build the dataset.
1 paper · 0 benchmarks
In our work, we have designed and implemented a novel workflow with several heuristic methods to combine state-of-the-art methods related to CVE fix commits gathering.
1 paper · 0 benchmarks
Morpho-MNIST (Morpho-MNIST: Quantitative Assessment and Diagnostics for Representation Learning)
Revealing latent structure in data is an active field of research, having introduced exciting technologies such as variational autoencoders and adversarial networks, and is essential to push machine learning towards unsupervised knowledge…
1 paper · 0 benchmarks
Dataset can be used by anyone who is interested to perform morphological classification of galaxies.
1 paper · 0 benchmarks
This dataset is for evaluation of morphosyntactic analyzers.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Motion Capture Data for Hand Motion Embodiment contains demonstrations of different hand motion recorded with the Qualisys MOCAP system.
1 paper · 0 benchmarks
This is a dataset of robot motions based on physics simulations.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.