{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transfusion-a-practical-and-effective","title":"TransFusion: A Practical and Effective Transformer-based Diffusion Model for 3D Human Motion Prediction","arxiv_id":"2307.16106","date":"2023-07-30","proceeding":null,"authors":["Sibo Tian","Minghui Zheng","Xiao Liang"],"abstract":"Predicting human motion plays a crucial role in ensuring a safe and effective human-robot close collaboration in intelligent remanufacturing systems of the future. Existing works can be categorized into two groups: those focusing on accuracy, predicting a single future motion, and those generating diverse predictions based on observations. The former group fails to address the uncertainty and multi-modal nature of human motion, while the latter group often produces motion sequences that deviate too far from the ground truth or become unrealistic within historical contexts. To tackle these issues, we propose TransFusion, an innovative and practical diffusion-based model for 3D human motion prediction which can generate samples that are more likely to happen while maintaining a certain level of diversity. Our model leverages Transformer as the backbone with long skip connections between shallow and deep layers. Additionally, we employ the discrete cosine transform to model motion sequences in the frequency space, thereby improving performance. In contrast to prior diffusion-based models that utilize extra modules like cross-attention and adaptive layer normalization to condition the prediction on past observed motion, we treat all inputs, including conditions, as tokens to create a more lightweight model compared to existing approaches. Extensive experimental studies are conducted on benchmark datasets to validate the effectiveness of our human motion prediction model.","url_abs":"https://arxiv.org/abs/2307.16106v1","url_pdf":"https://arxiv.org/pdf/2307.16106v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"transfusion-a-practical-and-effective","repo_url":"https://github.com/sibotian96/TransFusion","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"human-pose-forecasting","task_name":"Human Pose Forecasting"},{"task_slug":"human-motion-prediction","task_name":"Human motion prediction"},{"task_slug":"motion-prediction","task_name":"motion prediction"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"discrete-cosine-transform","method_name":"Discrete Cosine Transform"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-pose-forecasting-on-amass","task":"Human Pose Forecasting","dataset":"AMASS","model":"TransFusion","rank_in_archive_order":7,"of":11,"metrics":{"ADE":"0.508","APD":"8.853","FDE":"0.568"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-human36m","task":"Human Pose Forecasting","dataset":"Human3.6M","model":"TransFusion","rank_in_archive_order":32,"of":33,"metrics":{"ADE":"358","APD":"5975","FDE":"468","MMADE":"506","MMFDE":"539"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-humaneva-i","task":"Human Pose Forecasting","dataset":"HumanEva-I","model":"TransFusion","rank_in_archive_order":10,"of":11,"metrics":{"ADE@2000ms":"204","APD@2000ms":"1031","FDE@2000ms":"234","MMADE@2000ms":"408","MMFDE@2000ms":"427"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2307.16106","atlas_url":"https://app.syntology.ai/?focus=2307.16106","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}