{"url":"/sota/lipreading-on-lrs2","task":{"name":"Lipreading","url":"/task/lipreading","note":null},"dataset":{"name":"LRS2","url":"/dataset/lrs2"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing. \r\n\r\nDeep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks.  It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc. \r\n\r\nThe primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Word Error Rate (WER)"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Word Error Rate (WER)":"lower"}},"counts":{"rows":25,"rows_with_code":13,"rows_with_paper_page":25,"rows_dated":25,"rows_using_additional_data":17},"rows":[{"rank_in_archive_order":1,"model":"Auto-AVSR","metrics":{"Word Error Rate (WER)":"14.6"},"uses_additional_data":true,"paper_date":"2023-03-25","paper":"/paper/auto-avsr-audio-visual-speech-recognition","paper_url":"https://arxiv.org/abs/2303.14307v3","paper_title":"Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels","code":"https://github.com/mpc001/auto_avsr","n_code_links":2,"syntology":{"n_ran":0,"n_unverified":6,"n_samples":6,"n_pointer_only_licence":0}},{"rank_in_archive_order":2,"model":"USR","metrics":{"Word Error Rate (WER)":"15.4"},"uses_additional_data":true,"paper_date":"2024-11-04","paper":"/paper/unified-speech-recognition-a-single-model-for","paper_url":"https://arxiv.org/abs/2411.02256v1","paper_title":"Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs","code":"https://github.com/ahaliassos/usr","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":9,"n_samples":9,"n_pointer_only_licence":9}},{"rank_in_archive_order":3,"model":"SyncVSR","metrics":{"Word Error Rate (WER)":"16.5"},"uses_additional_data":true,"paper_date":"2024-06-18","paper":"/paper/syncvsr-data-efficient-visual-speech","paper_url":"https://arxiv.org/abs/2406.12233v1","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","code":"https://github.com/KAIST-AILab/SyncVSR","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"RAVEn Large","metrics":{"Word Error Rate (WER)":"18.6"},"uses_additional_data":true,"paper_date":"2022-12-12","paper":"/paper/jointly-learning-visual-and-auditory-speech","paper_url":"https://arxiv.org/abs/2212.06246v2","paper_title":"Jointly Learning Visual and Auditory Speech Representations from Raw Data","code":"https://github.com/ahaliassos/raven","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"VTP (more data)","metrics":{"Word Error Rate (WER)":"22.6"},"uses_additional_data":true,"paper_date":"2021-10-14","paper":"/paper/sub-word-level-lip-reading-with-visual","paper_url":"https://arxiv.org/abs/2110.07603v2","paper_title":"Sub-word Level Lip Reading With Visual Attention","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":6,"model":"ES³ Large + extLM","metrics":{"Word Error Rate (WER)":"24.6"},"uses_additional_data":true,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":7,"model":"CTC/Attention (LRW+LRS2/3+AVSpeech)","metrics":{"Word Error Rate (WER)":"25.5"},"uses_additional_data":true,"paper_date":"2022-02-26","paper":"/paper/visual-speech-recognition-for-multiple","paper_url":"https://arxiv.org/abs/2202.13084v2","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","code":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages","n_code_links":2,"syntology":null},{"rank_in_archive_order":8,"model":"ES³ Large","metrics":{"Word Error Rate (WER)":"26.7"},"uses_additional_data":true,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":9,"model":"ES³ Base + extLM","metrics":{"Word Error Rate (WER)":"28.7"},"uses_additional_data":true,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":10,"model":"VTP","metrics":{"Word Error Rate (WER)":"28.9"},"uses_additional_data":true,"paper_date":"2021-10-14","paper":"/paper/sub-word-level-lip-reading-with-visual","paper_url":"https://arxiv.org/abs/2110.07603v2","paper_title":"Sub-word Level Lip Reading With Visual Attention","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":11,"model":"SyncVSR","metrics":{"Word Error Rate (WER)":"28.9"},"uses_additional_data":false,"paper_date":"2024-06-18","paper":"/paper/syncvsr-data-efficient-visual-speech","paper_url":"https://arxiv.org/abs/2406.12233v1","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","code":"https://github.com/KAIST-AILab/SyncVSR","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"ES³ Base* + extLM","metrics":{"Word Error Rate (WER)":"29.3"},"uses_additional_data":false,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":13,"model":"ES³ Base","metrics":{"Word Error Rate (WER)":"30.7"},"uses_additional_data":true,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":14,"model":"ES³ Base*","metrics":{"Word Error Rate (WER)":"31.4"},"uses_additional_data":false,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":15,"model":"CTC/Attention","metrics":{"Word Error Rate (WER)":"32.9"},"uses_additional_data":false,"paper_date":"2022-02-26","paper":"/paper/visual-speech-recognition-for-multiple","paper_url":"https://arxiv.org/abs/2202.13084v2","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","code":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages","n_code_links":2,"syntology":null},{"rank_in_archive_order":16,"model":"Hybrid CTC / Attention","metrics":{"Word Error Rate (WER)":"39.1"},"uses_additional_data":false,"paper_date":"2021-02-12","paper":"/paper/end-to-end-audio-visual-speech-recognition","paper_url":"https://arxiv.org/abs/2102.06657v1","paper_title":"End-to-end Audio-visual Speech Recognition with Conformers","code":"https://github.com/zziz/pwc","n_code_links":3,"syntology":null},{"rank_in_archive_order":17,"model":"MoCo + wav2vec (w/o extLM)","metrics":{"Word Error Rate (WER)":"43.2"},"uses_additional_data":false,"paper_date":"2022-02-24","paper":"/paper/leveraging-uni-modal-self-supervised-learning-1","paper_url":"https://arxiv.org/abs/2203.07996v2","paper_title":"Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition","code":"https://github.com/lumia-group/leveraging-self-supervised-learning-for-avsr","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":6,"n_samples":6,"n_pointer_only_licence":0}},{"rank_in_archive_order":18,"model":"Multi-head Visual-Audio Memory","metrics":{"Word Error Rate (WER)":"44.5"},"uses_additional_data":true,"paper_date":"2022-04-04","paper":"/paper/distinguishing-homophenes-using-multi-head-1","paper_url":"https://arxiv.org/abs/2204.01725v1","paper_title":"Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading","code":"https://github.com/ms-dot-k/Multi-head-Visual-Audio-Memory","n_code_links":1,"syntology":null},{"rank_in_archive_order":19,"model":"TM-seq2seq + extLM","metrics":{"Word Error Rate (WER)":"48.3"},"uses_additional_data":true,"paper_date":"2018-09-06","paper":"/paper/deep-audio-visual-speech-recognition","paper_url":"http://arxiv.org/abs/1809.02108v2","paper_title":"Deep Audio-Visual Speech Recognition","code":"https://github.com/lordmartian/deep_avsr","n_code_links":4,"syntology":null},{"rank_in_archive_order":20,"model":"LF-MMI TDNN","metrics":{"Word Error Rate (WER)":"48.86"},"uses_additional_data":true,"paper_date":"2020-01-06","paper":"/paper/audio-visual-recognition-of-overlapped-speech","paper_url":"https://arxiv.org/abs/2001.01656v1","paper_title":"Audio-visual Recognition of Overlapped speech for the LRS2 dataset","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":21,"model":"Hybrid CTC / Attention","metrics":{"Word Error Rate (WER)":"50"},"uses_additional_data":false,"paper_date":"2018-09-28","paper":"/paper/audio-visual-speech-recognition-with-a-hybrid","paper_url":"http://arxiv.org/abs/1810.00108v1","paper_title":"Audio-Visual Speech Recognition With A Hybrid CTC/Attention Architecture","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":22,"model":"Conv-seq2seq","metrics":{"Word Error Rate (WER)":"51.7"},"uses_additional_data":true,"paper_date":"2019-10-01","paper":"/paper/spatio-temporal-fusion-based-convolutional","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Zhang_Spatio-Temporal_Fusion_Based_Convolutional_Sequence_Learning_for_Lip_Reading_ICCV_2019_paper.html","paper_title":"Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":23,"model":"CTC + KD ASR","metrics":{"Word Error Rate (WER)":"53.2"},"uses_additional_data":true,"paper_date":"2019-11-28","paper":"/paper/asr-is-all-you-need-cross-modal-distillation","paper_url":"https://arxiv.org/abs/1911.12747v2","paper_title":"ASR is all you need: cross-modal distillation for lip reading","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":24,"model":"TM-CTC + extLM","metrics":{"Word Error Rate (WER)":"54.7"},"uses_additional_data":true,"paper_date":"2018-09-06","paper":"/paper/deep-audio-visual-speech-recognition","paper_url":"http://arxiv.org/abs/1809.02108v2","paper_title":"Deep Audio-Visual Speech Recognition","code":"https://github.com/lordmartian/deep_avsr","n_code_links":4,"syntology":null},{"rank_in_archive_order":25,"model":"LIBS","metrics":{"Word Error Rate (WER)":"65.29"},"uses_additional_data":false,"paper_date":"2019-11-26","paper":"/paper/hearing-lips-improving-lip-reading-by","paper_url":"https://arxiv.org/abs/1911.11502v1","paper_title":"Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers","code":"https://github.com/zju-vipa/KamalEngine","n_code_links":1,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,795 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9581,"papers_checked":6795,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":2785},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":3,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":3,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":21,"n_samples":21,"n_pointer_only_licence":9,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":21,"n_samples":21,"n_pointer_only_licence":9,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}