{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigating-speech-features-for-continuous","title":"Investigating Speech Features for Continuous Turn-Taking Prediction Using LSTMs","arxiv_id":"1806.11461","date":"2018-06-29","proceeding":null,"authors":["Matthew Roddy","Gabriel Skantze","Naomi Harte"],"abstract":"For spoken dialog systems to conduct fluid conversational interactions with\nusers, the systems must be sensitive to turn-taking cues produced by a user.\nModels should be designed so that effective decisions can be made as to when it\nis appropriate, or not, for the system to speak. Traditional end-of-turn\nmodels, where decisions are made at utterance end-points, are limited in their\nability to model fast turn-switches and overlap. A more flexible approach is to\nmodel turn-taking in a continuous manner using RNNs, where the system predicts\nspeech probability scores for discrete frames within a future window. The\ncontinuous predictions represent generalized turn-taking behaviors observed in\nthe training data and can be applied to make decisions that are not just\nlimited to end-of-turn detection. In this paper, we investigate optimal\nspeech-related feature sets for making predictions at pauses and overlaps in\nconversation. We find that while traditional acoustic features perform well,\npart-of-speech features generally perform worse than word features. We show\nthat our current models outperform previously reported baselines.","url_abs":"http://arxiv.org/abs/1806.11461v1","url_pdf":"http://arxiv.org/pdf/1806.11461v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigating-speech-features-for-continuous","repo_url":"https://github.com/mattroddy/lstm_turn_taking_prediction","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.11461","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}