{"url":"/task/lipreading","name":"Lipreading","slug":"lipreading","description_markdown":"Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing. \r\n\r\nDeep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks.  It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc. \r\n\r\nThe primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":103,"papers_with_code":36,"benchmarks":8,"benchmark_tables_in_archive":8,"benchmark_tables_shown":8,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":8,"subtasks":1,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/lipreading-on-lrs2","slug":"lipreading-on-lrs2","dataset":"LRS2","dataset_url":"/dataset/lrs2","rows_in_archive":25,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"Auto-AVSR","paper_title":"Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels","paper_url":"/paper/auto-avsr-audio-visual-speech-recognition","paper_date":"2023-03-25","arxiv_id":"2303.14307","code_links":[{"title":"mpc001/auto_avsr","url":"https://github.com/mpc001/auto_avsr"},{"title":"umbertocappellazzo/llama-avsr","url":"https://github.com/umbertocappellazzo/llama-avsr"}],"syntology":{"n":6,"n_ran":0,"n_unverified":6,"n_pointer_only":0}}},{"leaderboard":"/sota/lipreading-on-lrs3-ted","slug":"lipreading-on-lrs3-ted","dataset":"LRS3-TED","dataset_url":"/dataset/lrs3-ted","rows_in_archive":23,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"LP + Conformer","paper_title":"Conformers are All You Need for Visual Speech Recognition","paper_url":"/paper/conformers-are-all-you-need-for-visual-speech","paper_date":"2023-02-17","arxiv_id":"2302.10915","code_links":[],"syntology":null}},{"leaderboard":"/sota/lipreading-on-lip-reading-in-the-wild","slug":"lipreading-on-lip-reading-in-the-wild","dataset":"Lip Reading in the Wild","dataset_url":"/dataset/lrw","rows_in_archive":22,"metrics":["Top-1 Accuracy"],"first_row_in_archive_order":{"model":"SyncVSR (Word Boundary)","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","paper_url":"/paper/syncvsr-data-efficient-visual-speech","paper_date":"2024-06-18","arxiv_id":"2406.12233","code_links":[{"title":"KAIST-AILab/SyncVSR","url":"https://github.com/KAIST-AILab/SyncVSR"}],"syntology":null}},{"leaderboard":"/sota/lipreading-on-lrw-1000","slug":"lipreading-on-lrw-1000","dataset":"CAS-VSR-W1k (LRW-1000)","dataset_url":"/dataset/lrw-1000","rows_in_archive":9,"metrics":["Top-1 Accuracy"],"first_row_in_archive_order":{"model":"SyncVSR (Word Boundary)","paper_title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","paper_url":"/paper/syncvsr-data-efficient-visual-speech","paper_date":"2024-06-18","arxiv_id":"2406.12233","code_links":[{"title":"KAIST-AILab/SyncVSR","url":"https://github.com/KAIST-AILab/SyncVSR"}],"syntology":null}},{"leaderboard":"/sota/lipreading-on-cmlr","slug":"lipreading-on-cmlr","dataset":"CMLR","dataset_url":null,"rows_in_archive":5,"metrics":["CER"],"first_row_in_archive_order":{"model":"CTC/Attention","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","paper_url":"/paper/visual-speech-recognition-for-multiple","paper_date":"2022-02-26","arxiv_id":"2202.13084","code_links":[{"title":"mpc001/Visual_Speech_Recognition_for_Multiple_Languages","url":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages"},{"title":"david-gimeno/lip-rtve","url":"https://github.com/david-gimeno/lip-rtve"}],"syntology":null}},{"leaderboard":"/sota/lipreading-on-grid-corpus-mixed-speech","slug":"lipreading-on-grid-corpus-mixed-speech","dataset":"GRID corpus (mixed-speech)","dataset_url":"/dataset/grid","rows_in_archive":5,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"CTC/Attention","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","paper_url":"/paper/visual-speech-recognition-for-multiple","paper_date":"2022-02-26","arxiv_id":"2202.13084","code_links":[{"title":"mpc001/Visual_Speech_Recognition_for_Multiple_Languages","url":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages"},{"title":"david-gimeno/lip-rtve","url":"https://github.com/david-gimeno/lip-rtve"}],"syntology":null}},{"leaderboard":"/sota/lipreading-on-lrw-1000-1","slug":"lipreading-on-lrw-1000-1","dataset":"LRW-1000","dataset_url":null,"rows_in_archive":4,"metrics":["Top-1 Accuracy"],"first_row_in_archive_order":{"model":"3D Conv + ResNet-18 + MS-TCN","paper_title":"Lipreading using Temporal Convolutional Networks","paper_url":"/paper/lipreading-using-temporal-convolutional","paper_date":"2020-01-23","arxiv_id":"2001.08702","code_links":[{"title":"mpc001/Lipreading_using_Temporal_Convolutional_Networks","url":"https://github.com/mpc001/Lipreading_using_Temporal_Convolutional_Networks"},{"title":"Yondijr/FlowerPower","url":"https://github.com/Yondijr/FlowerPower"}],"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}}},{"leaderboard":"/sota/lipreading-on-cas-vsr-s101","slug":"lipreading-on-cas-vsr-s101","dataset":"CAS-VSR-S101","dataset_url":"/dataset/cas-vsr-s101","rows_in_archive":1,"metrics":["Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"ES³ Base*","paper_title":"ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations","paper_url":"/paper/es3-evolving-self-supervised-learning-of","paper_date":"2024-01-01","arxiv_id":null,"code_links":[],"syntology":null}}],"datasets":[{"url":"/dataset/lrw","name":"LRW","full_name":"Lip Reading in the Wild","num_papers_in_archive":188},{"url":"/dataset/lrs2","name":"LRS2","full_name":"Lip Reading Sentences 2","num_papers_in_archive":115},{"url":"/dataset/lrs3-ted","name":"LRS3-TED","full_name":"","num_papers_in_archive":63},{"url":"/dataset/grid","name":"GRID Dataset","full_name":"","num_papers_in_archive":10},{"url":"/dataset/lrw-1000","name":"CAS-VSR-W1k (LRW-1000)","full_name":"CAS-VSR-W1k (LRW-1000)","num_papers_in_archive":9},{"url":"/dataset/glips","name":"GLips","full_name":"German Lips","num_papers_in_archive":6},{"url":"/dataset/cas-vsr-s101","name":"CAS-VSR-S101","full_name":"","num_papers_in_archive":1},{"url":"/dataset/miracl-vc1","name":"MIRACL-VC1","full_name":"MIRACL-VC1","num_papers_in_archive":1}],"subtasks":[{"url":"/task/landmark-based-lipreading","name":"Landmark-based Lipreading"}],"parent_tasks":[{"url":"/task/natural-language-transduction","name":"Natural Language Transduction"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":36,"tagged_in_all":103,"items":[{"url":"/paper/lipnet-end-to-end-sentence-level-lipreading","title":"LipNet: End-to-End Sentence-level Lipreading","date":"2016-11-05","arxiv_id":"1611.01599","repositories_listed":13,"syntology":{"n":15,"n_ran":0,"n_unverified":15,"n_pointer_only":0}},{"url":"/paper/deep-audio-visual-speech-recognition","title":"Deep Audio-Visual Speech Recognition","date":"2018-09-06","arxiv_id":"1809.02108","repositories_listed":4,"syntology":null},{"url":"/paper/combining-residual-networks-with-lstms-for","title":"Combining Residual Networks with LSTMs for Lipreading","date":"2017-03-12","arxiv_id":"1703.04105","repositories_listed":4,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/end-to-end-audio-visual-speech-recognition","title":"End-to-end Audio-visual Speech Recognition with Conformers","date":"2021-02-12","arxiv_id":"2102.06657","repositories_listed":3,"syntology":null},{"url":"/paper/auto-avsr-audio-visual-speech-recognition","title":"Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels","date":"2023-03-25","arxiv_id":"2303.14307","repositories_listed":2,"syntology":{"n":6,"n_ran":0,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/visual-speech-recognition-for-multiple","title":"Visual Speech Recognition for Multiple Languages in the Wild","date":"2022-02-26","arxiv_id":"2202.13084","repositories_listed":2,"syntology":null},{"url":"/paper/learning-audio-visual-speech-representation-1","title":"Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction","date":"2022-01-05","arxiv_id":"2201.02184","repositories_listed":2,"syntology":null},{"url":"/paper/discriminative-multi-modality-speech","title":"Discriminative Multi-modality Speech Recognition","date":"2020-05-12","arxiv_id":"2005.05592","repositories_listed":2,"syntology":null},{"url":"/paper/lipreading-using-temporal-convolutional","title":"Lipreading using Temporal Convolutional Networks","date":"2020-01-23","arxiv_id":"2001.08702","repositories_listed":2,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/lrw-1000-a-naturally-distributed-large-scale","title":"LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the Wild","date":"2018-10-16","arxiv_id":"1810.06990","repositories_listed":2,"syntology":null},{"url":"/paper/end-to-end-audiovisual-speech-recognition","title":"End-to-end Audiovisual Speech Recognition","date":"2018-02-18","arxiv_id":"1802.06424","repositories_listed":2,"syntology":null},{"url":"/paper/audio-visual-representation-learning-via","title":"Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models","date":"2025-02-09","arxiv_id":"2502.05766","repositories_listed":1,"syntology":null},{"url":"/paper/evaluation-of-end-to-end-continuous-spanish","title":"Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions","date":"2025-02-01","arxiv_id":"2502.00464","repositories_listed":1,"syntology":null},{"url":"/paper/unified-speech-recognition-a-single-model-for","title":"Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs","date":"2024-11-04","arxiv_id":"2411.02256","repositories_listed":1,"syntology":{"n":9,"n_ran":0,"n_unverified":9,"n_pointer_only":9}},{"url":"/paper/syncvsr-data-efficient-visual-speech","title":"SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization","date":"2024-06-18","arxiv_id":"2406.12233","repositories_listed":1,"syntology":null},{"url":"/paper/watch-your-mouth-silent-speech-recognition","title":"Watch Your Mouth: Silent Speech Recognition with Depth Sensing","date":"2024-05-11","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/audio-visual-speech-recognition-based-on","title":"Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems","date":"2024-05-09","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/where-visual-speech-meets-language-vsp-llm","title":"Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing","date":"2024-02-23","arxiv_id":"2402.15151","repositories_listed":1,"syntology":{"n":12,"n_ran":8,"n_unverified":4,"n_pointer_only":12}},{"url":"/paper/liplearner-customizable-silent-speech","title":"LipLearner: Customizable Silent Speech Interactions on Mobile Devices","date":"2023-02-12","arxiv_id":"2302.05907","repositories_listed":1,"syntology":null},{"url":"/paper/jointly-learning-visual-and-auditory-speech","title":"Jointly Learning Visual and Auditory Speech Representations from Raw Data","date":"2022-12-12","arxiv_id":"2212.06246","repositories_listed":1,"syntology":null},{"url":"/paper/relaxed-attention-for-transformer-models","title":"Relaxed Attention for Transformer Models","date":"2022-09-20","arxiv_id":"2209.09735","repositories_listed":1,"syntology":null},{"url":"/paper/training-strategies-for-improved-lip-reading","title":"Training Strategies for Improved Lip-reading","date":"2022-09-03","arxiv_id":"2209.01383","repositories_listed":1,"syntology":null},{"url":"/paper/bayesian-neural-network-language-modeling-for","title":"Bayesian Neural Network Language Modeling for Speech Recognition","date":"2022-08-28","arxiv_id":"2208.13259","repositories_listed":1,"syntology":null},{"url":"/paper/distinguishing-homophenes-using-multi-head-1","title":"Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading","date":"2022-04-04","arxiv_id":"2204.01725","repositories_listed":1,"syntology":null},{"url":"/paper/leveraging-uni-modal-self-supervised-learning-1","title":"Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech Recognition","date":"2022-02-24","arxiv_id":"2203.07996","repositories_listed":1,"syntology":{"n":6,"n_ran":0,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/robust-self-supervised-audio-visual-speech","title":"Robust Self-Supervised Audio-Visual Speech Recognition","date":"2022-01-05","arxiv_id":"2201.01763","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/lips-don-t-lie-a-generalisable-and-robust","title":"Lips Don't Lie: A Generalisable and Robust Approach to Face Forgery Detection","date":"2020-12-14","arxiv_id":"2012.07657","repositories_listed":1,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/learn-an-effective-lip-reading-model-without","title":"Learn an Effective Lip Reading Model without Pains","date":"2020-11-15","arxiv_id":"2011.07557","repositories_listed":1,"syntology":null},{"url":"/paper/towards-practical-lipreading-with-distilled","title":"Towards Practical Lipreading with Distilled and Efficient Models","date":"2020-07-13","arxiv_id":"2007.06504","repositories_listed":1,"syntology":null},{"url":"/paper/spotfast-networks-with-memory-augmented","title":"SpotFast Networks with Memory Augmented Lateral Transformers for Lipreading","date":"2020-05-21","arxiv_id":"2005.10903","repositories_listed":1,"syntology":null}],"syntology_records":9,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}