{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-spectro-temporal-features-in-end-to","title":"Exploring spectro-temporal features in end-to-end convolutional neural networks","arxiv_id":"1901.00072","date":"2019-01-01","proceeding":null,"authors":["Sean Robertson","Gerald Penn","Yingxue Wang"],"abstract":"Triangular, overlapping Mel-scaled filters (\"f-banks\") are the current\nstandard input for acoustic models that exploit their input's time-frequency\ngeometry, because they provide a psycho-acoustically motivated time-frequency\ngeometry for a speech signal. F-bank coefficients are provably robust to small\ndeformations in the scale. In this paper, we explore two ways in which filter\nbanks can be adjusted for the purposes of speech recognition. First, triangular\nfilters can be replaced with Gabor filters, a compactly supported filter that\nbetter localizes events in time, or Gammatone filters, a\npsychoacoustically-motivated filter. Second, by rearranging the order of\noperations in computing filter bank features, features can be integrated over\nsmaller time scales while simultaneously providing better frequency resolution.\nWe make all feature implementations available online through open-source\nrepositories. Initial experimentation with a modern end-to-end CNN phone\nrecognizer yielded no significant improvements to phone error rate due to\neither modification. The result, and its ramifications with respect to learned\nfilter banks, is discussed.","url_abs":"http://arxiv.org/abs/1901.00072v1","url_pdf":"http://arxiv.org/pdf/1901.00072v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-spectro-temporal-features-in-end-to","repo_url":"https://github.com/sdrobert/more-or-let","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}