{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-end-to-end-spoken-language","title":"Towards end-to-end spoken language understanding","arxiv_id":"1802.08395","date":"2018-02-23","proceeding":null,"authors":["Dmitriy Serdyuk","Yongqiang Wang","Christian Fuegen","Anuj Kumar","Baiyang Liu","Yoshua Bengio"],"abstract":"Spoken language understanding system is traditionally designed as a pipeline\nof a number of components. First, the audio signal is processed by an automatic\nspeech recognizer for transcription or n-best hypotheses. With the recognition\nresults, a natural language understanding system classifies the text to\nstructured data as domain, intent and slots for down-streaming consumers, such\nas dialog system, hands-free applications. These components are usually\ndeveloped and optimized independently. In this paper, we present our study on\nan end-to-end learning system for spoken language understanding. With this\nunified approach, we can infer the semantic meaning directly from audio\nfeatures without the intermediate text representation. This study showed that\nthe trained model can achieve reasonable good result and demonstrated that the\nmodel can capture the semantic attention directly from the audio features.","url_abs":"http://arxiv.org/abs/1802.08395v1","url_pdf":"http://arxiv.org/pdf/1802.08395v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-end-to-end-spoken-language","repo_url":"https://github.com/dmitriy-serdyuk/arxiv2kindle","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"spoken-language-understanding","task_name":"Spoken Language Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1802.08395","atlas_url":"https://app.syntology.ai/?focus=1802.08395","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}