{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/100000-podcasts-a-spoken-english-document","title":"100,000 Podcasts: A Spoken English Document Corpus","arxiv_id":null,"date":"2020-12-01","proceeding":"COLING 2020 8","authors":["Ann Clifton","Sravana Reddy","Yongze Yu","Aasish Pappu","Rezvaneh Rezapour","Hamed Bonab","Maria Eskevich","Gareth Jones","Jussi Karlgren","Ben Carterette","Rosie Jones"],"abstract":"Podcasts are a large and growing repository of spoken audio. As an audio format, podcasts are more varied in style and production type than broadcast news, contain more genres than typically studied in video data, and are more varied in style and format than previous corpora of conversations. When transcribed with automatic speech recognition they represent a noisy but fascinating collection of documents which can be studied through the lens of natural language processing, information retrieval, and linguistics. Paired with the audio files, they are also a resource for speech processing and the study of paralinguistic, sociolinguistic, and acoustic aspects of the domain. We introduce the Spotify Podcast Dataset, a new corpus of 100,000 podcasts. We demonstrate the complexity of the domain with a case study of two tasks: (1) passage search and (2) summarization. This is orders of magnitude larger than previous speech corpora used for search and summarization. Our results show that the size and variability of this corpus opens up new avenues for research.","url_abs":"https://aclanthology.org/2020.coling-main.519","url_pdf":"https://aclanthology.org/2020.coling-main.519.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-facial-landmark-localization","task_name":"3D Facial Landmark Localization"},{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"facial-expression-recognition","task_name":"Facial Expression Recognition (FER)"},{"task_slug":"highlight-detection","task_name":"Highlight Detection"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":null,"method_name":null},{"method_slug":null,"method_name":null},{"method_slug":null,"method_name":"The Ultimate Guide USA To Reaching Frontier  Reservations Number Explained"},{"method_slug":null,"method_name":null},{"method_slug":null,"method_name":null}],"datasets_introduced":[{"slug":"all-conference-alert","name":"All Conference Alert","full_name":"All Conference Alert"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-facial-landmark-localization-on-urban","task":"3D Facial Landmark Localization","dataset":"Urban Hyperspectral Image","model":"Lucky Brand 13","rank_in_archive_order":1,"of":1,"metrics":{"10°5 cm":"13.69"},"uses_additional_data":true},{"leaderboard":"/sota/facial-expression-recognition-fer-on-4","task":"Facial Expression Recognition (FER)","dataset":"^(#$!@#$)(()))******","model":"S","rank_in_archive_order":1,"of":1,"metrics":{"0..5sec":"sa"},"uses_additional_data":false},{"leaderboard":"/sota/highlight-detection-on-arabiska","task":"Highlight Detection","dataset":"arabiska","model":"Kenan Kanan","rank_in_archive_order":1,"of":1,"metrics":{"0..5sec":"https://youtu.be/pJ0auP7dbcY?si=vSiZevfJ57YUKC2q"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}