{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guided-attention-for-interpretable-motion","title":"Guided Attention for Interpretable Motion Captioning","arxiv_id":"2310.07324","date":"2023-10-11","proceeding":null,"authors":["Karim Radouane","Julien Lagarde","Sylvie Ranwez","Andon Tchechmedjiev"],"abstract":"Diverse and extensive work has recently been conducted on text-conditioned human motion generation. However, progress in the reverse direction, motion captioning, has seen less comparable advancement. In this paper, we introduce a novel architecture design that enhances text generation quality by emphasizing interpretability through spatio-temporal and adaptive attention mechanisms. To encourage human-like reasoning, we propose methods for guiding attention during training, emphasizing relevant skeleton areas over time and distinguishing motion-related words. We discuss and quantify our model's interpretability using relevant histograms and density distributions. Furthermore, we leverage interpretability to derive fine-grained information about human motion, including action localization, body part identification, and the distinction of motion-related words. Finally, we discuss the transferability of our approaches to other tasks. Our experiments demonstrate that attention guidance leads to interpretable captioning while enhancing performance compared to higher parameter-count, non-interpretable state-of-the-art systems. The code is available at: https://github.com/rd20karim/M2T-Interpretable.","url_abs":"https://arxiv.org/abs/2310.07324v2","url_pdf":"https://arxiv.org/pdf/2310.07324v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guided-attention-for-interpretable-motion","repo_url":"https://github.com/rd20karim/m2t-interpretable","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"motion-captioning","task_name":"Motion Captioning"},{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"spatio-temporal-video-grounding","task_name":"Spatio-Temporal Video Grounding"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/motion-captioning-on-humanml3d","task":"Motion Captioning","dataset":"HumanML3D","model":"ST-MLP","rank_in_archive_order":1,"of":4,"metrics":{"BERTScore":"40.3","BLEU-4":"25.0"},"uses_additional_data":false},{"leaderboard":"/sota/motion-captioning-on-kit-motion-language","task":"Motion Captioning","dataset":"KIT Motion-Language","model":"ST-MLP","rank_in_archive_order":2,"of":3,"metrics":{"BERTScore":"41.2","BLEU-4":"24.4"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}