{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lipreading-using-temporal-convolutional","title":"Lipreading using Temporal Convolutional Networks","arxiv_id":"2001.08702","date":"2020-01-23","proceeding":null,"authors":["Brais Martinez","Pingchuan Ma","Stavros Petridis","Maja Pantic"],"abstract":"Lip-reading has attracted a lot of research attention lately thanks to advances in deep learning. The current state-of-the-art model for recognition of isolated words in-the-wild consists of a residual network and Bidirectional Gated Recurrent Unit (BGRU) layers. In this work, we address the limitations of this model and we propose changes which further improve its performance. Firstly, the BGRU layers are replaced with Temporal Convolutional Networks (TCN). Secondly, we greatly simplify the training procedure, which allows us to train the model in one single stage. Thirdly, we show that the current state-of-the-art methodology produces models that do not generalize well to variations on the sequence length, and we addresses this issue by proposing a variable-length augmentation. We present results on the largest publicly-available datasets for isolated word recognition in English and Mandarin, LRW and LRW1000, respectively. Our proposed model results in an absolute improvement of 1.2% and 3.2%, respectively, in these datasets which is the new state-of-the-art performance.","url_abs":"https://arxiv.org/abs/2001.08702v1","url_pdf":"https://arxiv.org/pdf/2001.08702v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lipreading-using-temporal-convolutional","repo_url":"https://github.com/Yondijr/FlowerPower","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"lipreading-using-temporal-convolutional","repo_url":"https://github.com/mpc001/Lipreading_using_Temporal_Convolutional_Networks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"lip-reading","task_name":"Lip Reading"},{"task_slug":"lipreading","task_name":"Lipreading"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/lipreading-on-lrw-1000-1","task":"Lipreading","dataset":"LRW-1000","model":"3D Conv + ResNet-18 + MS-TCN","rank_in_archive_order":1,"of":4,"metrics":{"Top-1 Accuracy":"41.4%"},"uses_additional_data":false},{"leaderboard":"/sota/lipreading-on-lip-reading-in-the-wild","task":"Lipreading","dataset":"Lip Reading in the Wild","model":"3D Conv + ResNet-18 + MS-TCN","rank_in_archive_order":12,"of":22,"metrics":{"Top-1 Accuracy":"85.30"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2001.08702","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2001.08702"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Yondijr/FlowerPower","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mpc001/Lipreading_using_Temporal_Convolutional_Networks","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran_fixture":1,"unverified":1},"by_repo_kind":{"listed":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"463db9779e4d1664","entry":"threeD_to_2D_tensor","repo":"Yondijr/FlowerPower","repo_kind":"listed","path":"lipreading/model.py","file_url":"https://github.com/Yondijr/FlowerPower/blob/HEAD/lipreading/model.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"463db9779e4d1664"}},{"code_sha256_prefix":"3308f7059a23a0f2","entry":"landmarks_interpolate","repo":"Yondijr/FlowerPower","repo_kind":"listed","path":"preprocessing/crop_mouth_from_video.py","file_url":"https://github.com/Yondijr/FlowerPower/blob/HEAD/preprocessing/crop_mouth_from_video.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"3308f7059a23a0f2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}