{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speech-and-speaker-recognition-from-raw","title":"Speech and Speaker Recognition from Raw Waveform with SincNet","arxiv_id":"1812.05920","date":"2018-12-13","proceeding":null,"authors":["Mirco Ravanelli","Yoshua Bengio"],"abstract":"Deep neural networks can learn complex and abstract representations, that are\nprogressively obtained by combining simpler ones. A recent trend in speech and\nspeaker recognition consists in discovering these representations starting from\nraw audio samples directly. Differently from standard hand-crafted features\nsuch as MFCCs or FBANK, the raw waveform can potentially help neural networks\ndiscover better and more customized representations. The high-dimensional raw\ninputs, however, can make training significantly more challenging. This paper\nsummarizes our recent efforts to develop a neural architecture that efficiently\nprocesses speech from audio waveforms. In particular, we propose SincNet, a\nnovel Convolutional Neural Network (CNN) that encourages the first layer to\ndiscover meaningful filters by exploiting parametrized sinc functions. In\ncontrast to standard CNNs, which learn all the elements of each filter, only\nlow and high cutoff frequencies of band-pass filters are directly learned from\ndata. This inductive bias offers a very compact way to derive a customized\nfront-end, that only depends on some parameters with a clear physical meaning.\nOur experiments, conducted on both speaker and speech recognition, show that\nthe proposed architecture converges faster, performs better, and is more\ncomputationally efficient than standard CNNs.","url_abs":"http://arxiv.org/abs/1812.05920v1","url_pdf":"http://arxiv.org/pdf/1812.05920v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speech-and-speaker-recognition-from-raw","repo_url":"https://github.com/mravanelli/SincNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"speech-and-speaker-recognition-from-raw","repo_url":"https://github.com/meiyor/SincNet-for-Autism-EEG-based-Emotion-Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"inductive-bias","task_name":"Inductive Bias"},{"task_slug":"speaker-recognition","task_name":"Speaker Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}