{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hf0-a-hybrid-pitch-extraction-method-for","title":"hf0: A hybrid pitch extraction method for multimodal voice","arxiv_id":"1904.09765","date":"2019-04-22","proceeding":null,"authors":["Pradeep Rengaswamy","Gurunath Reddy M","Krothapalli Sreenivasa Rao"],"abstract":"Pitch or fundamental frequency (f0) extraction is a fundamental problem\nstudied extensively for its potential applications in speech and clinical\napplications. In literature, explicit mode specific (modal speech or singing\nvoice or emotional/ expressive speech or noisy speech) signal processing and\ndeep learning f0 extraction methods that exploit the quasi periodic nature of\nthe signal in time, harmonic property in spectral or combined form to extract\nthe pitch is developed. Hence, there is no single unified method which can\nreliably extract the pitch from various modes of the acoustic signal. In this\nwork, we propose a hybrid f0 extraction method which seamlessly extracts the\npitch across modes of speech production with very high accuracy required for\nmany applications. The proposed hybrid model exploits the advantages of deep\nlearning and signal processing methods to minimize the pitch detection error\nand adopts to various modes of acoustic signal. Specifically, we propose an\nordinal regression convolutional neural networks to map the periodicity rich\ninput representation to obtain the nominal pitch classes which drastically\nreduces the number of classes required for pitch detection unlike other deep\nlearning approaches. Further, the accurate f0 is estimated from the nominal\npitch class labels by filtering and autocorrelation. We show that the proposed\nmethod generalizes to the unseen modes of voice production and various noises\nfor large scale datasets. Also, the proposed hybrid model significantly reduces\nthe learning parameters required to train the deep model compared to other\nmethods. Furthermore,the evaluation measures showed that the proposed method is\nsignificantly better than the state-of-the-art signal processing and deep\nlearning approaches.","url_abs":"http://arxiv.org/abs/1904.09765v1","url_pdf":"http://arxiv.org/pdf/1904.09765v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hf0-a-hybrid-pitch-extraction-method-for","repo_url":"https://github.com/Pradeepiit/hf0","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}