{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-speech-recognition-from-the-raw","title":"End-to-End Speech Recognition From the Raw Waveform","arxiv_id":"1806.07098","date":"2018-06-19","proceeding":null,"authors":["Neil Zeghidour","Nicolas Usunier","Gabriel Synnaeve","Ronan Collobert","Emmanuel Dupoux"],"abstract":"State-of-the-art speech recognition systems rely on fixed, hand-crafted\nfeatures such as mel-filterbanks to preprocess the waveform before the training\npipeline. In this paper, we study end-to-end systems trained directly from the\nraw waveform, building on two alternatives for trainable replacements of\nmel-filterbanks that use a convolutional architecture. The first one is\ninspired by gammatone filterbanks (Hoshen et al., 2015; Sainath et al, 2015),\nand the second one by the scattering transform (Zeghidour et al., 2017). We\npropose two modifications to these architectures and systematically compare\nthem to mel-filterbanks, on the Wall Street Journal dataset. The first\nmodification is the addition of an instance normalization layer, which greatly\nimproves on the gammatone-based trainable filterbanks and speeds up the\ntraining of the scattering-based filterbanks. The second one relates to the\nlow-pass filter used in these approaches. These modifications consistently\nimprove performances for both approaches, and remove the need for a careful\ninitialization in scattering-based trainable filterbanks. In particular, we\nshow a consistent improvement in word error rate of the trainable filterbanks\nrelatively to comparable mel-filterbanks. It is the first time end-to-end\nmodels trained from the raw signal significantly outperform mel-filterbanks on\na large vocabulary task under clean recording conditions.","url_abs":"http://arxiv.org/abs/1806.07098v2","url_pdf":"http://arxiv.org/pdf/1806.07098v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-speech-recognition-from-the-raw","repo_url":"https://github.com/renyuanL/ry-Speech-commands","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"instance-normalization","method_name":"Instance Normalization"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.07098","atlas_url":"https://app.syntology.ai/?focus=1806.07098","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}