{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-only-recognition-of-normal-whispered","title":"Visual-Only Recognition of Normal, Whispered and Silent Speech","arxiv_id":"1802.06399","date":"2018-02-18","proceeding":null,"authors":["Stavros Petridis","Jie Shen","Doruk Cetin","Maja Pantic"],"abstract":"Silent speech interfaces have been recently proposed as a way to enable\ncommunication when the acoustic signal is not available. This introduces the\nneed to build visual speech recognition systems for silent and whispered\nspeech. However, almost all the recently proposed systems have been trained on\nvocalised data only. This is in contrast with evidence in the literature which\nsuggests that lip movements change depending on the speech mode. In this work,\nwe introduce a new audiovisual database which is publicly available and\ncontains normal, whispered and silent speech. To the best of our knowledge,\nthis is the first study which investigates the differences between the three\nspeech modes using the visual modality only. We show that an absolute decrease\nin classification rate of up to 3.7% is observed when training and testing on\nnormal and whispered, respectively, and vice versa. An even higher decrease of\nup to 8.5% is reported when the models are tested on silent speech. This\nreveals that there are indeed visual differences between the 3 speech modes and\nthe common assumption that vocalized training data can be used directly to\ntrain a silent speech recognition system may not be true.","url_abs":"http://arxiv.org/abs/1802.06399v1","url_pdf":"http://arxiv.org/pdf/1802.06399v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"silent-speech-recognition","task_name":"Silent Speech Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"visual-speech-recognition","task_name":"Visual Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[{"slug":"av-digits-database","name":"AV Digits Database","full_name":"AV Digits Database"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}