{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quaternion-convolutional-neural-networks-for-1","title":"Quaternion Convolutional Neural Networks for End-to-End Automatic Speech Recognition","arxiv_id":"1806.07789","date":"2018-06-20","proceeding":null,"authors":["Titouan Parcollet","Ying Zhang","Mohamed Morchid","Chiheb Trabelsi","Georges Linarès","Renato De Mori","Yoshua Bengio"],"abstract":"Recently, the connectionist temporal classification (CTC) model coupled with\nrecurrent (RNN) or convolutional neural networks (CNN), made it easier to train\nspeech recognition systems in an end-to-end fashion. However in real-valued\nmodels, time frame components such as mel-filter-bank energies and the cepstral\ncoefficients obtained from them, together with their first and second order\nderivatives, are processed as individual elements, while a natural alternative\nis to process such components as composed entities. We propose to group such\nelements in the form of quaternions and to process these quaternions using the\nestablished quaternion algebra. Quaternion numbers and quaternion neural\nnetworks have shown their efficiency to process multidimensional inputs as\nentities, to encode internal dependencies, and to solve many tasks with less\nlearning parameters than real-valued models. This paper proposes to integrate\nmultiple feature views in quaternion-valued convolutional neural network\n(QCNN), to be used for sequence-to-sequence mapping with the CTC model.\nPromising results are reported using simple QCNNs in phoneme recognition\nexperiments with the TIMIT corpus. More precisely, QCNNs obtain a lower phoneme\nerror rate (PER) with less learning parameters than a competing model based on\nreal-valued CNNs.","url_abs":"http://arxiv.org/abs/1806.07789v1","url_pdf":"http://arxiv.org/pdf/1806.07789v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"quaternion-convolutional-neural-networks-for-1","repo_url":"https://github.com/Riccardo-Vecchi/Pytorch-Quaternion-Neural-Networks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"phoneme-recognition","task_name":"Phoneme Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-timit","task":"Speech Recognition","dataset":"TIMIT","model":"QCNN-10L-256FM","rank_in_archive_order":19,"of":22,"metrics":{"Percentage error":"19.64"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.07789","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}