{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-learning-of-harmonic-priors-for","title":"Efficient Learning of Harmonic Priors for Pitch Detection in Polyphonic Music","arxiv_id":"1705.07104","date":"2017-05-19","proceeding":null,"authors":["Pablo A. Alvarado","Dan Stowell"],"abstract":"Automatic music transcription (AMT) aims to infer a latent symbolic\nrepresentation of a piece of music (piano-roll), given a corresponding observed\naudio recording. Transcribing polyphonic music (when multiple notes are played\nsimultaneously) is a challenging problem, due to highly structured overlapping\nbetween harmonics. We study whether the introduction of physically inspired\nGaussian process (GP) priors into audio content analysis models improves the\nextraction of patterns required for AMT. Audio signals are described as a\nlinear combination of sources. Each source is decomposed into the product of an\namplitude-envelope, and a quasi-periodic component process. We introduce the\nMat\\'ern spectral mixture (MSM) kernel for describing frequency content of\nsingles notes. We consider two different regression approaches. In the sigmoid\nmodel every pitch-activation is independently non-linear transformed. In the\nsoftmax model several activation GPs are jointly non-linearly transformed. This\nintroduce cross-correlation between activations. We use variational Bayes for\napproximate inference. We empirically evaluate how these models work in\npractice transcribing polyphonic music. We demonstrate that rather than\nencourage dependency between activations, what is relevant for improving pitch\ndetection is to learnt priors that fit the frequency content of the sound\nevents to detect.","url_abs":"http://arxiv.org/abs/1705.07104v2","url_pdf":"http://arxiv.org/pdf/1705.07104v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-learning-of-harmonic-priors-for","repo_url":"https://github.com/PabloAlvarado/MSMK","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"music-transcription","task_name":"Music Transcription"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}