{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/speech-text-dialog-pre-training-for-spoken","title":"Speech-Text Dialog Pre-training for Spoken Dialog Understanding with Explicit Cross-Modal Alignment","arxiv_id":"2305.11579","date":"2023-05-19","proceeding":null,"authors":["Tianshu Yu","Haoyu Gao","Ting-En Lin","Min Yang","Yuchuan Wu","Wentao Ma","Chao Wang","Fei Huang","Yongbin Li"],"abstract":"Recently, speech-text pre-training methods have shown remarkable success in many speech and natural language processing tasks. However, most previous pre-trained models are usually tailored for one or two specific tasks, but fail to conquer a wide range of speech-text tasks. In addition, existing speech-text pre-training methods fail to explore the contextual information within a dialogue to enrich utterance representations. In this paper, we propose Speech-text dialog Pre-training for spoken dialog understanding with ExpliCiT cRoss-Modal Alignment (SPECTRA), which is the first-ever speech-text dialog pre-training model. Concretely, to consider the temporality of speech modality, we design a novel temporal position prediction task to capture the speech-text alignment. This pre-training task aims to predict the start and end time of each textual word in the corresponding speech waveform. In addition, to learn the characteristics of spoken dialogs, we generalize a response selection task from textual dialog pre-training to speech-text dialog pre-training scenarios. Experimental results on four different downstream speech-text tasks demonstrate the superiority of SPECTRA in learning speech-text alignment and multi-turn dialog context.","url_abs":"https://arxiv.org/abs/2305.11579v2","url_pdf":"https://arxiv.org/pdf/2305.11579v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"speech-text-dialog-pre-training-for-spoken","repo_url":"https://github.com/alibabaresearch/damo-convai","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"emotion-recognition-in-conversation","task_name":"Emotion Recognition in Conversation"},{"task_slug":"multimodal-intent-recognition","task_name":"Multimodal Intent Recognition"},{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"cross-modal-alignment","task_name":"cross-modal alignment"}],"methods":[{"method_slug":"fail","method_name":"fail"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/emotion-recognition-in-conversation-on","task":"Emotion Recognition in Conversation","dataset":"IEMOCAP","model":"SPECTRA","rank_in_archive_order":59,"of":59,"metrics":{"Accuracy":"67.94"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-intent-recognition-on-mintrec","task":"Multimodal Intent Recognition","dataset":"MIntRec","model":"SPECTRA","rank_in_archive_order":3,"of":6,"metrics":{"Accuracy (20 classes)":"73.48"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosei-1","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSEI","model":"SPECTRA","rank_in_archive_order":4,"of":15,"metrics":{"Accuracy":"87.34"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosi","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSI","model":"SPECTRA","rank_in_archive_order":11,"of":12,"metrics":{"Acc-2":"87.5"},"uses_additional_data":true},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-mosi","task":"Multimodal Sentiment Analysis","dataset":"MOSI","model":"SPECTRA","rank_in_archive_order":2,"of":11,"metrics":{"Accuracy":"87.50"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.11579","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}