{"url":"/method/wav2vec-u","slug":"wav2vec-u","name":"wav2vec-U","full_name":"wav2vec Unsupervised","full_name_withheld":false,"description_markdown":"**wav2vec-U** is an unsupervised method to train speech recognition models without any labeled data. It leverages self-supervised speech representations to segment unlabeled language and learn a mapping from these representations to phonemes via adversarial training. \r\n\r\nSpecifically, we learn self-supervised representations with wav2vec 2.0 on unlabeled speech audio, then identify clusters in the representations with k-means to segment the audio data. Next, we build segment representations by mean pooling the wav2vec 2.0 representations, performing [PCA](https://paperswithcode.com/method/pca) and a second mean pooling step between adjacent segments. This is input to the generator which outputs a phoneme sequence that is fed to the discriminator, similar to phonemized unlabeled text to perform adversarial training.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2105.11084v3","title":"Unsupervised Speech Recognition","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Speech Recognition","url":"/methods/category/speech-recognition","pwc_aliases":[]}],"n_papers_tagged":6,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/the-hidden-dance-of-phonemes-and-visage","title":"The Hidden Dance of Phonemes and Visage: Unveiling the Enigmatic Link between Phonemes and Facial Features","date":"2023-07-26","arxiv_id":"2307.13953","n_code_links":1,"syntology":null},{"paper":null,"title":"Unsupervised ASR via Cross-Lingual Pseudo-Labeling","date":"2023-05-19","arxiv_id":"2305.13330","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Unsupervised Speech Recognition with Diffusion GANs","date":"2023-03-23","arxiv_id":"2303.13559","n_code_links":0,"syntology":null},{"paper":"/paper/towards-end-to-end-unsupervised-speech","title":"Towards End-to-end Unsupervised Speech Recognition","date":"2022-04-05","arxiv_id":"2204.02492","n_code_links":1,"syntology":null},{"paper":null,"title":"Analyzing the Robustness of Unsupervised Speech Recognition","date":"2021-10-07","arxiv_id":"2110.03509","n_code_links":0,"syntology":null},{"paper":"/paper/unsupervised-speech-recognition","title":"Unsupervised Speech Recognition","date":"2021-05-24","arxiv_id":"2105.11084","n_code_links":4,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}}],"papers_shown":6,"tasks":[{"task":"/task/speech-recognition","name":"Speech Recognition","papers":5},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":5},{"task":"/task/unsupervised-speech-recognition","name":"Unsupervised Speech Recognition","papers":4},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":3},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":3},{"task":null,"name":"Generative Adversarial Network","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1}],"tasks_shown":8,"n_tasks":8,"usage_by_year":[{"year":"2021","papers":2},{"year":"2022","papers":1},{"year":"2023","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/wav2vec-u"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}