{"url":"/dataset/semeion","name":"Semeion","full_name":"Semeion Handwritten Digit Data Set","description_markdown":"1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.\r\n\r\nThe dataset was created by Tactile Srl, Brescia, Italy (http://www.tattile.it) and donated in 1994 to Semeion Research Center of Sciences of Communication, Rome, Italy (http://www.semeion.it), for machine learning research.\r\n\r\nFor any questions, e-mail Massimo Buscema (m.buscema '@' semeion.it) or Stefano Terzi (s.terzi '@' semeion.it)\r\n\r\n##Data Set Information:\r\n\r\n1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values. Then each pixel of each image was scaled into a boolean (1/0) value using a fixed threshold.\r\n\r\nEach person wrote on a paper all the digits from 0 to 9, twice. The commitment was to write the digit the first time in the normal way (trying to write each digit accurately) and the second time in a fast way (with no accuracy).\r\n\r\nThe best validation protocol for this dataset seems to be a 5x2CV, 50% Tune (Train +Test), and completely blind 50% Validation\r\n\r\n\r\n##Attribute Information:\r\nThis dataset consists of 1593 records (rows) and 256 attributes (columns).\r\nEach record represents a handwritten digit, originally scanned with a resolution of 256 grays scale (28).\r\nEach pixel of each original scanned image was first stretched, and after scaled between 0 and 1 (setting to 0 for every pixel whose value was under the value 127 of the grey scale (127 included) and setting to 1 for each pixel whose original value in the grey scale was over 127).\r\n\r\nFinally, each binary image was scaled again into a 16x16 square box (the final 256 binary attributes).","description_withheld":null,"homepage":"https://archive.ics.uci.edu/ml/datasets/semeion+handwritten+digit","introduced_date":"2008-11-11","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"}],"languages":[],"variants":["Semeion"],"data_loaders":[{"repo":"https://github.com/pytorch/vision","url":"https://pytorch.org/vision/stable/generated/torchvision.datasets.SEMEION.html","frameworks":["pytorch"]}],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}