{"url":"/dataset/polymnist","name":"PolyMNIST","full_name":null,"description_markdown":"The dataset is based on the original MNIST dataset.\r\nCompared to the original dataset, the digits are scaled down by a factor of $0.75$ such that there is more space for the random translation.The PolyMNIST consists of 5 different modalities.\r\n\r\nThe background of every modality $\\mathbf{x}_m$ consists of random patches of size $28 \\times 28$ from a large image.\r\nAnd the digit is placed at a random position of the patch.\r\nUsing this setup, every modality has modality-specific information given by its background image and shared information given by the digit, which is shared between all modalities.\r\nAn additional difficulty compared to the original PolyMNIST is the random translation of the digits","description_withheld":null,"homepage":"https://github.com/thomassutter/MoPoE","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/generalized-multimodal-elbo-1","title":"Generalized Multimodal ELBO","first_author":"Thomas M. Sutter","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[],"languages":[],"variants":["PolyMNIST"],"data_loaders":[],"num_papers_in_archive":14,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}