{"url":"/sota/lipreading-on-lrs3-ted","task":{"name":"Lipreading","url":"/task/lipreading","note":null},"dataset":{"name":"LRS3-TED","url":"/dataset/lrs3-ted"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing. \r\n\r\nDeep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks.  It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc. \r\n\r\nThe primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Word Error Rate (WER)"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Word Error Rate (WER)":"lower"}},"counts":{"rows":23,"rows_with_code":15,"rows_with_paper_page":23,"rows_dated":23,"rows_using_additional_data":19},"rows":[{"rank_in_archive_order":1,"model":"LP + Conformer","metrics":{"Word Error Rate (WER)":"12.8"},"uses_additional_data":true,"paper_date":"2023-02-17","paper":"/paper/conformers-are-all-you-need-for-visual-speech","paper_url":"https://arxiv.org/abs/2302.10915v2","paper_title":"Conformers are All You Need for Visual Speech Recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":2,"model":"Auto-AVSR","metrics":{"Word Error Rate (WER)":"19.1"},"uses_additional_data":true,"paper_date":"2023-03-25","paper":"/paper/auto-avsr-audio-visual-speech-recognition","paper_url":"https://arxiv.org/abs/2303.14307v3","paper_title":"Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels","code":"https://github.com/mpc001/auto_avsr","n_code_links":2,"syntology":{"n_ran":0,"n_unverified":6,"n_samples":6,"n_pointer_only_licence":0}},{"rank_in_archive_order":3,"model":"SyncVSR","metrics":{"Word Error Rate (WER)":"21.5"},"uses_additional_data":true,"paper_date":"2024-06-18","paper":"/paper/syncvsr-data-efficient-visual-speech","paper_url":"https://arxiv.org/abs/2406.12233v1","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","code":"https://github.com/KAIST-AILab/SyncVSR","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"USR (self + semi-supervised)","metrics":{"Word Error Rate (WER)":"21.5"},"uses_additional_data":true,"paper_date":"2024-11-04","paper":"/paper/unified-speech-recognition-a-single-model-for","paper_url":"https://arxiv.org/abs/2411.02256v1","paper_title":"Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs","code":"https://github.com/ahaliassos/usr","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":9,"n_samples":9,"n_pointer_only_licence":9}},{"rank_in_archive_order":5,"model":"USR (self-supervised)","metrics":{"Word Error Rate (WER)":"22.3"},"uses_additional_data":true,"paper_date":"2024-11-04","paper":"/paper/unified-speech-recognition-a-single-model-for","paper_url":"https://arxiv.org/abs/2411.02256v1","paper_title":"Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs","code":"https://github.com/ahaliassos/usr","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":9,"n_samples":9,"n_pointer_only_licence":9}},{"rank_in_archive_order":6,"model":"RAVEn Large","metrics":{"Word Error Rate (WER)":"23.4"},"uses_additional_data":true,"paper_date":"2022-12-12","paper":"/paper/jointly-learning-visual-and-auditory-speech","paper_url":"https://arxiv.org/abs/2212.06246v2","paper_title":"Jointly Learning Visual and Auditory Speech Representations from Raw Data","code":"https://github.com/ahaliassos/raven","n_code_links":1,"syntology":null},{"rank_in_archive_order":7,"model":"VSP-LLM","metrics":{"Word Error Rate (WER)":"25.4"},"uses_additional_data":true,"paper_date":"2024-02-23","paper":"/paper/where-visual-speech-meets-language-vsp-llm","paper_url":"https://arxiv.org/abs/2402.15151v2","paper_title":"Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing","code":"https://github.com/sally-sh/vsp-llm","n_code_links":1,"syntology":{"n_ran":8,"n_unverified":4,"n_samples":12,"n_pointer_only_licence":12}},{"rank_in_archive_order":8,"model":"AV-HuBERT Large + Relaxed Attention + LM","metrics":{"Word Error Rate (WER)":"25.51"},"uses_additional_data":true,"paper_date":"2022-09-20","paper":"/paper/relaxed-attention-for-transformer-models","paper_url":"https://arxiv.org/abs/2209.09735v1","paper_title":"Relaxed Attention for Transformer Models","code":"https://github.com/Oguzhanercan/Vision-Transformers","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"DistillAV","metrics":{"Word Error Rate (WER)":"26.2"},"uses_additional_data":true,"paper_date":"2025-02-09","paper":"/paper/audio-visual-representation-learning-via","paper_url":"https://arxiv.org/abs/2502.05766v1","paper_title":"Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models","code":"https://github.com/jxzhanggg/DistillAV","n_code_links":1,"syntology":null},{"rank_in_archive_order":10,"model":"AV-HuBERT Large","metrics":{"Word Error Rate (WER)":"26.9"},"uses_additional_data":true,"paper_date":"2022-01-05","paper":"/paper/learning-audio-visual-speech-representation-1","paper_url":"https://arxiv.org/abs/2201.02184v2","paper_title":"Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction","code":"https://github.com/facebookresearch/av_hubert","n_code_links":2,"syntology":null},{"rank_in_archive_order":11,"model":"VTP (more data)","metrics":{"Word Error Rate (WER)":"30.7"},"uses_additional_data":true,"paper_date":"2021-10-14","paper":"/paper/sub-word-level-lip-reading-with-visual","paper_url":"https://arxiv.org/abs/2110.07603v2","paper_title":"Sub-word Level Lip Reading With Visual Attention","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":12,"model":"SyncVSR","metrics":{"Word Error Rate (WER)":"31.2"},"uses_additional_data":false,"paper_date":"2024-06-18","paper":"/paper/syncvsr-data-efficient-visual-speech","paper_url":"https://arxiv.org/abs/2406.12233v1","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","code":"https://github.com/KAIST-AILab/SyncVSR","n_code_links":1,"syntology":null},{"rank_in_archive_order":13,"model":"CTC/Attention (LRW+LRS2/3+AVSpeech)","metrics":{"Word Error Rate (WER)":"31.5"},"uses_additional_data":true,"paper_date":"2022-02-26","paper":"/paper/visual-speech-recognition-for-multiple","paper_url":"https://arxiv.org/abs/2202.13084v2","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","code":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages","n_code_links":2,"syntology":null},{"rank_in_archive_order":14,"model":"RNN-T","metrics":{"Word Error Rate (WER)":"33.6"},"uses_additional_data":true,"paper_date":"2019-11-08","paper":"/paper/recurrent-neural-network-transducer-for-audio","paper_url":"https://arxiv.org/abs/1911.04890v1","paper_title":"Recurrent Neural Network Transducer for Audio-Visual Speech Recognition","code":"https://github.com/around-star/Speech-Recognition","n_code_links":1,"syntology":null},{"rank_in_archive_order":15,"model":"ES³ Large","metrics":{"Word Error Rate (WER)":"37.1"},"uses_additional_data":false,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":16,"model":"ES³ Base","metrics":{"Word Error Rate (WER)":"40.3"},"uses_additional_data":false,"paper_date":"2024-01-01","paper":"/paper/es3-evolving-self-supervised-learning-of","paper_url":"http://openaccess.thecvf.com//content/CVPR2024/html/Zhang_ES3_Evolving_Self-Supervised_Learning_of_Robust_Audio-Visual_Speech_Representations_CVPR_2024_paper.html","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":17,"model":"VTP","metrics":{"Word Error Rate (WER)":"40.6"},"uses_additional_data":true,"paper_date":"2021-10-14","paper":"/paper/sub-word-level-lip-reading-with-visual","paper_url":"https://arxiv.org/abs/2110.07603v2","paper_title":"Sub-word Level Lip Reading With Visual Attention","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":18,"model":"Hyb + Conformer","metrics":{"Word Error Rate (WER)":"43.3"},"uses_additional_data":true,"paper_date":"2021-02-12","paper":"/paper/end-to-end-audio-visual-speech-recognition","paper_url":"https://arxiv.org/abs/2102.06657v1","paper_title":"End-to-end Audio-visual Speech Recognition with Conformers","code":"https://github.com/zziz/pwc","n_code_links":3,"syntology":null},{"rank_in_archive_order":19,"model":"CTC-V2P","metrics":{"Word Error Rate (WER)":"55.1"},"uses_additional_data":true,"paper_date":"2018-07-13","paper":"/paper/large-scale-visual-speech-recognition","paper_url":"http://arxiv.org/abs/1807.05162v3","paper_title":"Large-Scale Visual Speech Recognition","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":20,"model":"EG-seq2seq","metrics":{"Word Error Rate (WER)":"57.8"},"uses_additional_data":false,"paper_date":"2020-05-12","paper":"/paper/discriminative-multi-modality-speech","paper_url":"https://arxiv.org/abs/2005.05592v2","paper_title":"Discriminative Multi-modality Speech Recognition","code":"https://github.com/JackSyu/Discriminative-Multi-modality-Speech-Recognition","n_code_links":2,"syntology":null},{"rank_in_archive_order":21,"model":"TM-seq2seq","metrics":{"Word Error Rate (WER)":"58.9"},"uses_additional_data":true,"paper_date":"2018-09-06","paper":"/paper/deep-audio-visual-speech-recognition","paper_url":"http://arxiv.org/abs/1809.02108v2","paper_title":"Deep Audio-Visual Speech Recognition","code":"https://github.com/lordmartian/deep_avsr","n_code_links":4,"syntology":null},{"rank_in_archive_order":22,"model":"CTC + KD","metrics":{"Word Error Rate (WER)":"59.8"},"uses_additional_data":true,"paper_date":"2019-11-28","paper":"/paper/asr-is-all-you-need-cross-modal-distillation","paper_url":"https://arxiv.org/abs/1911.12747v2","paper_title":"ASR is all you need: cross-modal distillation for lip reading","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":23,"model":"Conv-seq2seq","metrics":{"Word Error Rate (WER)":"60.1"},"uses_additional_data":true,"paper_date":"2019-10-01","paper":"/paper/spatio-temporal-fusion-based-convolutional","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Zhang_Spatio-Temporal_Fusion_Based_Convolutional_Sequence_Learning_for_Lip_Reading_ICCV_2019_paper.html","paper_title":"Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading","code":null,"n_code_links":0,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9581,"papers_checked":6264,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":3316},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":4,"rows_with_any_sample_ran":1,"distinct_papers_with_graph_line":3,"distinct_papers_with_any_sample_ran":1,"samples_over_distinct_papers":{"n_ran":8,"n_unverified":19,"n_samples":27,"n_pointer_only_licence":21,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":8,"n_unverified":28,"n_samples":36,"n_pointer_only_licence":30,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}