{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wav2pix-speech-conditioned-face-generation","title":"Wav2Pix: Speech-conditioned Face Generation using Generative Adversarial Networks","arxiv_id":"1903.10195","date":"2019-03-25","proceeding":null,"authors":["Amanda Duarte","Francisco Roldan","Miquel Tubau","Janna Escur","Santiago Pascual","Amaia Salvador","Eva Mohedano","Kevin McGuinness","Jordi Torres","Xavier Giro-i-Nieto"],"abstract":"Speech is a rich biometric signal that contains information about the\nidentity, gender and emotional state of the speaker. In this work, we explore\nits potential to generate face images of a speaker by conditioning a Generative\nAdversarial Network (GAN) with raw speech input. We propose a deep neural\nnetwork that is trained from scratch in an end-to-end fashion, generating a\nface directly from the raw speech waveform without any additional identity\ninformation (e.g reference image or one-hot encoding). Our model is trained in\na self-supervised approach by exploiting the audio and visual signals naturally\naligned in videos. With the purpose of training from video data, we present a\nnovel dataset collected for this work, with high-quality videos of youtubers\nwith notable expressiveness in both the speech and visual signals.","url_abs":"http://arxiv.org/abs/1903.10195v1","url_pdf":"http://arxiv.org/pdf/1903.10195v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wav2pix-speech-conditioned-face-generation","repo_url":"https://github.com/imatge-upc/wav2pix","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"wav2pix-speech-conditioned-face-generation","repo_url":"https://github.com/Aryan05/Generative-Modelling-of-Images-from-Speech_Speech2Face","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"wav2pix-speech-conditioned-face-generation","repo_url":"https://github.com/saiteja-talluri/Speech2Face","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"face-generation","task_name":"Face Generation"},{"task_slug":null,"task_name":"Generative Adversarial Network"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.10195","atlas_url":"https://app.syntology.ai/?focus=1903.10195","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}