{"url":"/dataset/av-digits-database","name":"AV Digits Database","full_name":"AV Digits Database","description_markdown":"AV Digits Database is an audiovisual database which contains normal, whispered and silent speech. 53 participants were recorded from 3 different views (frontal, 45 and profile) pronouncing digits and phrases in three speech modes.\r\n\r\nThe database consists of two parts: digits and short phrases. In the first part, participants were asked to read 10 digits, from 0 to 9, in English in random order five times. In case of non-native English speakers this part was also repeated in the participant’s native language. In total, 53 participants (41 males and 12 females) from 16 nationalities, were recorded with a mean age and standard deviation of 26.7 and 4.3 years, respectively.\r\n\r\nIn the second part, participants were asked to read 10 short phrases. The phrases are the same as the ones used in the OuluVS2 database: “Excuse me”, “Goodbye”, “Hello”, “How are you”, “Nice to meet you”, “See you”, “I am sorry”,   “Thank you”, “Have a good time”, “You are welcome”. Again, each phrase was repeated five times in 3 different modes, neutral, whisper and silent speech. Thirty nine participants (32 males and 7 females) were recorded for this part with a mean age and standard deviation of 26.3 and 3.8 years, respectively.\r\n\r\nSource: [AV Digits Database](https://ibug-avs.eu/)","description_withheld":null,"homepage":"https://ibug-avs.eu/","introduced_date":"2018-02-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/visual-only-recognition-of-normal-whispered","title":"Visual-Only Recognition of Normal, Whispered and Silent Speech","first_author":"Stavros Petridis","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Audio","url":"/datasets/modality/audio"},{"name":"Speech","url":"/datasets/modality/speech"}],"tasks":[{"name":"Speech Recognition","url":"/task/speech-recognition","datasets_with_task":"/datasets/task/speech-recognition"},{"name":"Visual Speech Recognition","url":"/task/visual-speech-recognition","datasets_with_task":"/datasets/task/visual-speech-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["AV Digits Database"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}