{"url":"/dataset/mhist","name":"MHIST","full_name":"Minimalist Histopathology image analysis dataset","description_markdown":"The **m**inimalist **hist**opathology image analysis dataset (**MHIST**) is a binary classification dataset of 3,152 fixed-size images of colorectal polyps, each with a gold-standard label determined by the majority vote of seven board-certified gastrointestinal pathologists. MHIST also includes each image’s annotator agreement level. As a minimalist dataset, MHIST occupies less than 400 MB of disk space, and a ResNet-18 baseline can be trained to convergence on MHIST in just 6 minutes using approximately 3.5 GB of memory on a NVIDIA RTX 3090. As example use cases, the authors use MHIST to study natural questions that arise in histopathology image classification such as how dataset size, network depth, transfer learning, and high-disagreement examples affect model performance.\r\n\r\nSource: [Wei et al.](https://arxiv.org/pdf/2101.12355.pdf)\r\n\r\nImage source: [Wei et al.](https://arxiv.org/pdf/2101.12355.pdf)","description_withheld":null,"homepage":"https://bmirds.github.io/MHIST/","introduced_date":"2021-01-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-petri-dish-for-histopathology-image","title":"A Petri Dish for Histopathology Image Analysis","first_author":"Jerry Wei","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Biology","url":"/datasets/modality/biology"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"}],"languages":[],"variants":["MHIST"],"data_loaders":[],"num_papers_in_archive":28,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/classification-on-mhist","task":"Classification","dataset_variant":"MHIST","rows":9,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"MoCo-v2 (ResNet-50)","paper":"/paper/improved-transferability-of-self-supervised","metrics":{"Accuracy":"88.03"},"code_links":[{"title":"vpulab/bn_finetuning","url":"https://github.com/vpulab/bn_finetuning"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/improved-transferability-of-self-supervised","title":"Improved transferability of self-supervised learning models through batch normalization finetuning","date":"2024-08-15","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/benchmarking-self-supervised-learning-on","title":"Benchmarking Self-Supervised Learning on Diverse Pathology Datasets","date":"2022-12-09","rows_on_this_dataset":6,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}