{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semtalk-holistic-co-speech-motion-generation","title":"SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis","arxiv_id":"2412.16563","date":"2024-12-21","proceeding":null,"authors":["Xiangyue Zhang","Jianfang Li","Jiaxu Zhang","Ziqiang Dang","Jianqiang Ren","Liefeng Bo","Zhigang Tu"],"abstract":"A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generation with frame-level semantic emphasis. Our key insight is to separately learn general motions and sparse motions, and then adaptively fuse them. In particular, rhythmic consistency learning is explored to establish rhythm-related base motion, ensuring a coherent foundation that synchronizes gestures with the speech rhythm. Subsequently, textit{semantic emphasis learning is designed to generate semantic-aware sparse motion, focusing on frame-level semantic cues. Finally, to integrate sparse motion into the base motion and generate semantic-emphasized co-speech gestures, we further leverage a learned semantic score for adaptive synthesis. Qualitative and quantitative comparisons on two public datasets demonstrate that our method outperforms the state-of-the-art, delivering high-quality co-speech motion with enhanced semantic richness over a stable base motion.","url_abs":"https://arxiv.org/abs/2412.16563v2","url_pdf":"https://arxiv.org/pdf/2412.16563v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"gesture-generation","task_name":"Gesture Generation"},{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"rhythm","task_name":"Rhythm"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/gesture-generation-on-beat2","task":"Gesture Generation","dataset":"BEAT2","model":"SemTalk","rank_in_archive_order":3,"of":14,"metrics":{"FGD":"0.4278"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}