{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-speech-representation-pooling","title":"Unsupervised Speech Representation Pooling Using Vector Quantization","arxiv_id":"2304.03940","date":"2023-04-08","proceeding":null,"authors":["Jeongkyun Park","Kwanghee Choi","Hyunjun Heo","Hyung-Min Park"],"abstract":"With the advent of general-purpose speech representations from large-scale self-supervised models, applying a single model to multiple downstream tasks is becoming a de-facto approach. However, the pooling problem remains; the length of speech representations is inherently variable. The naive average pooling is often used, even though it ignores the characteristics of speech, such as differently lengthed phonemes. Hence, we design a novel pooling method to squash acoustically similar representations via vector quantization, which does not require additional training, unlike attention-based pooling. Further, we evaluate various unsupervised pooling methods on various self-supervised models. We gather diverse methods scattered around speech and text to evaluate on various tasks: keyword spotting, speaker identification, intent classification, and emotion recognition. Finally, we quantitatively and qualitatively analyze our method, comparing it with supervised pooling methods.","url_abs":"https://arxiv.org/abs/2304.03940v1","url_pdf":"https://arxiv.org/pdf/2304.03940v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-speech-representation-pooling","repo_url":"https://github.com/IIP-Sogang/speech-pooling-benchmark","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"intent-classification","task_name":"Intent Classification"},{"task_slug":"keyword-spotting","task_name":"Keyword Spotting"},{"task_slug":"quantization","task_name":"Quantization"},{"task_slug":"speaker-identification","task_name":"Speaker Identification"},{"task_slug":"intent-classification-1","task_name":"intent-classification"}],"methods":[{"method_slug":"average-pooling","method_name":"Average Pooling"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}