{"url":"/dataset/oracle-mnist","name":"Oracle-MNIST","full_name":"Oracle-MNIST: a Realistic Image Dataset for Benchmarking Machine Learning Algorithms","description_markdown":"We introduce the Oracle-MNIST dataset, comprising of 2828 grayscale images of 30,222 ancient characters from 10 categories, for benchmarking pattern classification, with particular challenges on image noise and distortion. The training set totally consists of 27,222 images, and the test set contains 300 images per class. Oracle-MNIST shares the same data format with the original MNIST dataset, allowing for direct compatibility with all existing classifiers and systems, but it constitutes a more challenging classification task than MNIST. The images of ancient characters suffer from 1) extremely serious and unique noises caused by three-thousand years of burial and aging and 2) dramatically variant writing styles by ancient Chinese, which all make them realistic for machine learning research. The dataset is freely available at https://github.com/wm-bupt/oracle-mnist.","description_withheld":null,"homepage":"https://paperswithcode.com/paper/oracle-mnist-a-realistic-image-dataset-for","introduced_date":"2022-05-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/oracle-mnist-a-realistic-image-dataset-for","title":"Oracle-MNIST: a Realistic Image Dataset for Benchmarking Machine Learning Algorithms","first_author":"Mei Wang","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Oracle-MNIST"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/image-classification-on-oracle-mnist","task":"Image Classification","dataset_variant":"Oracle-MNIST","rows":4,"metrics":["Accuracy","Trainable Parameters"],"first_row_in_archive_order":{"model":"ResNet-18 + Vision Eagle Attention","paper":"/paper/vision-eagle-attention-a-new-lens-for","metrics":{"Accuracy":"97.20"},"code_links":[{"title":"MahmudulHasan11085/Vision-Eagle-Attention","url":"https://github.com/MahmudulHasan11085/Vision-Eagle-Attention"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/vision-eagle-attention-a-new-lens-for","title":"Vision Eagle Attention: a new lens for advancing image classification","date":"2024-11-15","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/learning-local-discrete-features-in","title":"Learning local discrete features in explainable-by-design convolutional neural networks","date":"2024-10-31","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/a-block-based-convolutional-neural-network","title":"LR-Net: A Block-based Convolutional Neural Network for Low-Resolution Image Classification","date":"2022-07-19","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}