{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-word-embeddings-for-visual-speech","title":"Deep word embeddings for visual speech recognition","arxiv_id":"1710.11201","date":"2017-10-30","proceeding":null,"authors":["Themos Stafylakis","Georgios Tzimiropoulos"],"abstract":"In this paper we present a deep learning architecture for extracting word\nembeddings for visual speech recognition. The embeddings summarize the\ninformation of the mouth region that is relevant to the problem of word\nrecognition, while suppressing other types of variability such as speaker, pose\nand illumination. The system is comprised of a spatiotemporal convolutional\nlayer, a Residual Network and bidirectional LSTMs and is trained on the\nLipreading in-the-wild database. We first show that the proposed architecture\ngoes beyond state-of-the-art on closed-set word identification, by attaining\n11.92% error rate on a vocabulary of 500 words. We then examine the capacity of\nthe embeddings in modelling words unseen during training. We deploy\nProbabilistic Linear Discriminant Analysis (PLDA) to model the embeddings and\nperform low-shot learning experiments on words unseen during training. The\nexperiments demonstrate that word-level visual speech recognition is feasible\neven in cases where the target words are not included in the training set.","url_abs":"http://arxiv.org/abs/1710.11201v1","url_pdf":"http://arxiv.org/pdf/1710.11201v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-word-embeddings-for-visual-speech","repo_url":"https://github.com/tstafylakis/Lipreading-ResNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"lipreading","task_name":"Lipreading"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"visual-speech-recognition","task_name":"Visual Speech Recognition"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}