{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speaker-detection-in-the-wild-lessons-learned","title":"Speaker detection in the wild: Lessons learned from JSALT 2019","arxiv_id":"1912.00938","date":"2019-12-02","proceeding":null,"authors":["Paola Garcia","Jesus Villalba","Herve Bredin","Jun Du","Diego Castan","Alejandrina Cristia","Latane Bullock","Ling Guo","Koji Okabe","Phani Sankar Nidadavolu","Saurabh Kataria","Sizhu Chen","Leo Galmant","Marvin Lavechin","Lei Sun","Marie-Philippe Gill","Bar Ben-Yair","Sajjad Abdoli","Xin Wang","Wassim Bouaziz","Hadrien Titeux","Emmanuel Dupoux","Kong Aik Lee","Najim Dehak"],"abstract":"This paper presents the problems and solutions addressed at the JSALT workshop when using a single microphone for speaker detection in adverse scenarios. The main focus was to tackle a wide range of conditions that go from meetings to wild speech. We describe the research threads we explored and a set of modules that was successful for these scenarios. The ultimate goal was to explore speaker detection; but our first finding was that an effective diarization improves detection, and not having a diarization stage impoverishes the performance. All the different configurations of our research agree on this fact and follow a main backbone that includes diarization as a previous stage. With this backbone, we analyzed the following problems: voice activity detection, how to deal with noisy signals, domain mismatch, how to improve the clustering; and the overall impact of previous stages in the final speaker detection. In this paper, we show partial results for speaker diarizarion to have a better understanding of the problem and we present the final results for speaker detection.","url_abs":"https://arxiv.org/abs/1912.00938v1","url_pdf":"https://arxiv.org/pdf/1912.00938v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"links_only","authors_date_abstract":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license), from the Kaggle arXiv metadata snapshot of 2026-09-12"},"code_links":[{"paper_slug":"speaker-detection-in-the-wild-lessons-learned","repo_url":"https://github.com/MarvinLvn/voice-type-classifier","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}