{"url":"/sota/lipreading-on-cmlr","task":{"name":"Lipreading","url":"/task/lipreading","note":null},"dataset":{"name":"CMLR","url":null},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"Lipreading is a process of extracting speech by watching lip movements of a speaker in the absence of sound. Humans lipread all the time without even noticing. It is a big part in communication albeit not as dominant as audio. It is a very helpful skill to learn especially for those who are hard of hearing. \r\n\r\nDeep Lipreading is the process of extracting speech from a video of a silent talking face using deep neural networks.  It is also known by few other names: Visual Speech Recognition (VSR), Machine Lipreading, Automatic Lipreading etc. \r\n\r\nThe primary methodology involves two stages: i) Extracting visual and temporal features from a sequence of image frames from a silent talking video ii) Processing the sequence of features into units of speech e.g. characters, words, phrases etc. We can find several implementations of this methodology either done in two separate stages or trained end-to-end in one go.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["CER"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"CER":"lower"}},"counts":{"rows":5,"rows_with_code":2,"rows_with_paper_page":5,"rows_dated":5,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"CTC/Attention","metrics":{"CER":"9.1%"},"uses_additional_data":false,"paper_date":"2022-02-26","paper":"/paper/visual-speech-recognition-for-multiple","paper_url":"https://arxiv.org/abs/2202.13084v2","paper_title":"Visual Speech Recognition for Multiple Languages in the Wild","code":"https://github.com/mpc001/Visual_Speech_Recognition_for_Multiple_Languages","n_code_links":2,"syntology":null},{"rank_in_archive_order":2,"model":"LIBS","metrics":{"CER":"31.27%"},"uses_additional_data":false,"paper_date":"2019-11-26","paper":"/paper/hearing-lips-improving-lip-reading-by","paper_url":"https://arxiv.org/abs/1911.11502v1","paper_title":"Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers","code":"https://github.com/zju-vipa/KamalEngine","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"CSSMCM","metrics":{"CER":"32.48%"},"uses_additional_data":false,"paper_date":"2019-08-14","paper":"/paper/a-cascade-sequence-to-sequence-model-for","paper_url":"https://arxiv.org/abs/1908.04917v2","paper_title":"A Cascade Sequence-to-Sequence Model for Chinese Mandarin Lip Reading","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":4,"model":"LipCH-Net","metrics":{"CER":"34.07%"},"uses_additional_data":false,"paper_date":"2019-08-14","paper":"/paper/a-cascade-sequence-to-sequence-model-for","paper_url":"https://arxiv.org/abs/1908.04917v2","paper_title":"A Cascade Sequence-to-Sequence Model for Chinese Mandarin Lip Reading","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":5,"model":"WAS","metrics":{"CER":"38.93%"},"uses_additional_data":false,"paper_date":"2019-08-14","paper":"/paper/a-cascade-sequence-to-sequence-model-for","paper_url":"https://arxiv.org/abs/1908.04917v2","paper_title":"A Cascade Sequence-to-Sequence Model for Chinese Mandarin Lip Reading","code":null,"n_code_links":0,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}