{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/in-search-of-a-dataset-for-handwritten","title":"In Search of a Dataset for Handwritten Optical Music Recognition: Introducing MUSCIMA++","arxiv_id":"1703.04824","date":"2017-03-14","proceeding":null,"authors":["Jan Hajič jr.","Pavel Pecina"],"abstract":"Optical Music Recognition (OMR) has long been without an adequate dataset and\nground truth for evaluating OMR systems, which has been a major problem for\nestablishing a state of the art in the field. Furthermore, machine learning\nmethods require training data. We analyze how the OMR processing pipeline can\nbe expressed in terms of gradually more complex ground truth, and based on this\nanalysis, we design the MUSCIMA++ dataset of handwritten music notation that\naddresses musical symbol recognition and notation reconstruction. The MUSCIMA++\ndataset version 0.9 consists of 140 pages of handwritten music, with 91255\nmanually annotated notation symbols and 82261 explicitly marked relationships\nbetween symbol pairs. The dataset allows training and evaluating models for\nsymbol classification, symbol localization, and notation graph assembly, both\nin isolation and jointly. Open-source tools are provided for manipulating the\ndataset, visualizing the data and further annotation, and the dataset itself is\nmade available under an open license.","url_abs":"http://arxiv.org/abs/1703.04824v1","url_pdf":"http://arxiv.org/pdf/1703.04824v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"in-search-of-a-dataset-for-handwritten","repo_url":"https://github.com/OMR-Research/muscima-pp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.04824","atlas_url":"https://app.syntology.ai/?focus=1703.04824","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}