{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sum-product-networks-for-robust-automatic","title":"Sum-Product Networks for Robust Automatic Speaker Identification","arxiv_id":"1910.11969","date":"2020-08-13","proceeding":null,"authors":[],"abstract":"We introduce sum-product networks (SPNs) for robust speech processing through\na simple robust automatic speaker identification (ASI) task. SPNs are deep\nprobabilistic graphical models capable of answering multiple probabilistic\nqueries. We show that SPNs are able to remain robust by using the marginal\nprobability density function (PDF) of the spectral features that reliably\nrepresent speech. Though current SPN toolkits and learning algorithms are in\ntheir infancy, we aim to show that SPNs have the potential to become a useful\ntool for robust speech processing in the future. SPN speaker models are\nevaluated here on real-world non-stationary and coloured noise sources at\nmultiple signal-to-noise ratio (SNR) levels. In terms of ASI accuracy, we find\nthat SPN speaker models are more robust than two recent convolutional neural\nnetwork (CNN)-based ASI systems. Additionally, SPN speaker models consist of\nsignificantly fewer parameters than their CNN-based counterparts. The results\nindicate that SPN speaker models could be a robust, parameter-efficient\nalternative for ASI. Additionally, this work demonstrates that SPNs have\npotential in related tasks, such as robust automatic speech recognition (ASR)\nand automatic speaker verification (ASV). Availability: The SPN ASI system is\navailable at https://github.com/anicolson/SPN-ASI.","url_abs":"http://arxiv.org/abs/1910.11969v3","url_pdf":"http://arxiv.org/pdf/1910.11969v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sum-product-networks-for-robust-automatic","repo_url":"https://github.com/anicolson/SPN-ASI","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speaker-identification","task_name":"Speaker Identification"},{"task_slug":"speaker-verification","task_name":"Speaker Verification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}