Papers › Unsupervised Speech Representation Pooling Using Vector Quantization

Unsupervised Speech Representation Pooling Using Vector Quantization

8 Apr 2023arXiv:2304.03940archive 2025-07-28

Jeongkyun Park, Kwanghee Choi, Hyunjun Heo, Hyung-Min Park

With the advent of general-purpose speech representations from large-scale self-supervised models, applying a single model to multiple downstream tasks is becoming a de-facto approach. However, the pooling problem remains; the length of speech representations is inherently variable. The naive average pooling is often used, even though it ignores the characteristics of speech, such as differently lengthed phonemes. Hence, we design a novel pooling method to squash acoustically similar representations via vector quantization, which does not require additional training, unlike attention-based pooling. Further, we evaluate various unsupervised pooling methods on various self-supervised models. We gather diverse methods scattered around speech and text to evaluate on various tasks: keyword spotting, speaker identification, intent classification, and emotion recognition. Finally, we quantitatively and qualitatively analyze our method, comparing it with supervised pooling methods.

PaperPDFCode

Code

IIP-Sogang/speech-pooling-benchmark officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion RecognitionIntent ClassificationKeyword SpottingQuantizationSpeaker Identificationintent-classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Average Pooling

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections