{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lrw-1000-a-naturally-distributed-large-scale","title":"LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the Wild","arxiv_id":"1810.06990","date":"2018-10-16","proceeding":null,"authors":["Shuang Yang","Yuan-Hang Zhang","Dalu Feng","Mingmin Yang","Chenhao Wang","Jing-Yun Xiao","Keyu Long","Shiguang Shan","Xilin Chen"],"abstract":"Large-scale datasets have successively proven their fundamental importance in\nseveral research fields, especially for early progress in some emerging topics.\nIn this paper, we focus on the problem of visual speech recognition, also known\nas lipreading, which has received increasing interest in recent years. We\npresent a naturally-distributed large-scale benchmark for lip reading in the\nwild, named LRW-1000, which contains 1,000 classes with 718,018 samples from\nmore than 2,000 individual speakers. Each class corresponds to the syllables of\na Mandarin word composed of one or several Chinese characters. To the best of\nour knowledge, it is currently the largest word-level lipreading dataset and\nalso the only public large-scale Mandarin lipreading dataset. This dataset aims\nat covering a \"natural\" variability over different speech modes and imaging\nconditions to incorporate challenges encountered in practical applications. It\nhas shown a large variation in this benchmark in several aspects, including the\nnumber of samples in each class, video resolution, lighting conditions, and\nspeakers' attributes such as pose, age, gender, and make-up. Besides providing\na detailed description of the dataset and its collection pipeline, we evaluate\nseveral typical popular lipreading methods and perform a thorough analysis of\nthe results from several aspects. The results demonstrate the consistency and\nchallenges of our dataset, which may open up some new promising directions for\nfuture work.","url_abs":"http://arxiv.org/abs/1810.06990v6","url_pdf":"http://arxiv.org/pdf/1810.06990v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lrw-1000-a-naturally-distributed-large-scale","repo_url":"https://github.com/Fengdalu/Lipreading-DenseNet3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"lrw-1000-a-naturally-distributed-large-scale","repo_url":"https://github.com/NirHeaven/D3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"lip-reading","task_name":"Lip Reading"},{"task_slug":"lipreading","task_name":"Lipreading"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"visual-speech-recognition","task_name":"Visual Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[{"slug":"lrw-1000","name":"CAS-VSR-W1k (LRW-1000)","full_name":"CAS-VSR-W1k (LRW-1000)"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/lipreading-on-lrw-1000-1","task":"Lipreading","dataset":"LRW-1000","model":"3D Conv + ResNet-34 + Bi-GRU","rank_in_archive_order":2,"of":4,"metrics":{"Top-1 Accuracy":"38.19%"},"uses_additional_data":false},{"leaderboard":"/sota/lipreading-on-lrw-1000-1","task":"Lipreading","dataset":"LRW-1000","model":"DenseNet3D + Bi-GRU","rank_in_archive_order":3,"of":4,"metrics":{"Top-1 Accuracy":"34.76%"},"uses_additional_data":false},{"leaderboard":"/sota/lipreading-on-lrw-1000-1","task":"Lipreading","dataset":"LRW-1000","model":"Multi-Tower LSTM-5","rank_in_archive_order":4,"of":4,"metrics":{"Top-1 Accuracy":"25.76%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.06990","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}