{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/zero-shot-keyword-spotting-for-visual-speech","title":"Zero-shot keyword spotting for visual speech recognition in-the-wild","arxiv_id":"1807.08469","date":"2018-07-23","proceeding":"ECCV 2018 9","authors":["Themos Stafylakis","Georgios Tzimiropoulos"],"abstract":"Visual keyword spotting (KWS) is the problem of estimating whether a text\nquery occurs in a given recording using only video information. This paper\nfocuses on visual KWS for words unseen during training, a real-world, practical\nsetting which so far has received no attention by the community. To this end,\nwe devise an end-to-end architecture comprising (a) a state-of-the-art visual\nfeature extractor based on spatiotemporal Residual Networks, (b) a\ngrapheme-to-phoneme model based on sequence-to-sequence neural networks, and\n(c) a stack of recurrent neural networks which learn how to correlate visual\nfeatures with the keyword representation. Different to prior works on KWS,\nwhich try to learn word representations merely from sequences of graphemes\n(i.e. letters), we propose the use of a grapheme-to-phoneme encoder-decoder\nmodel which learns how to map words to their pronunciation. We demonstrate that\nour system obtains very promising visual-only KWS results on the challenging\nLRS2 database, for keywords unseen during training. We also show that our\nsystem outperforms a baseline which addresses KWS via automatic speech\nrecognition (ASR), while it drastically improves over other recently proposed\nASR-free KWS methods.","url_abs":"http://arxiv.org/abs/1807.08469v2","url_pdf":"http://arxiv.org/pdf/1807.08469v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"zero-shot-keyword-spotting-for-visual-speech","repo_url":"https://github.com/lilianemomeni/KWS-Net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"visual-keyword-spotting","task_name":"Visual Keyword Spotting"},{"task_slug":"visual-speech-recognition","task_name":"Visual Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.08469","atlas_url":"https://app.syntology.ai/?focus=1807.08469","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}