{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/digital-life-project-autonomous-3d-characters","title":"Digital Life Project: Autonomous 3D Characters with Social Intelligence","arxiv_id":"2312.04547","date":"2023-12-07","proceeding":"CVPR 2024 1","authors":["Zhongang Cai","Jianping Jiang","Zhongfei Qing","Xinying Guo","Mingyuan Zhang","Zhengyu Lin","Haiyi Mei","Chen Wei","Ruisi Wang","Wanqi Yin","Xiangyu Fan","Han Du","Liang Pan","Peng Gao","Zhitao Yang","Yang Gao","Jiaqi Li","Tianxiang Ren","Yukun Wei","Xiaogang Wang","Chen Change Loy","Lei Yang","Ziwei Liu"],"abstract":"In this work, we present Digital Life Project, a framework utilizing language as the universal medium to build autonomous 3D characters, who are capable of engaging in social interactions and expressing with articulated body motions, thereby simulating life in a digital environment. Our framework comprises two primary components: 1) SocioMind: a meticulously crafted digital brain that models personalities with systematic few-shot exemplars, incorporates a reflection process based on psychology principles, and emulates autonomy by initiating dialogue topics; 2) MoMat-MoGen: a text-driven motion synthesis paradigm for controlling the character's digital body. It integrates motion matching, a proven industry technique to ensure motion quality, with cutting-edge advancements in motion generation for diversity. Extensive experiments demonstrate that each module achieves state-of-the-art performance in its respective domain. Collectively, they enable virtual characters to initiate and sustain dialogues autonomously, while evolving their socio-psychological states. Concurrently, these characters can perform contextually relevant bodily movements. Additionally, a motion captioning module further allows the virtual character to recognize and appropriately respond to human players' actions. Homepage: https://digital-life-project.com/","url_abs":"https://arxiv.org/abs/2312.04547v1","url_pdf":"https://arxiv.org/pdf/2312.04547v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"motion-captioning","task_name":"Motion Captioning"},{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"motion-synthesis","task_name":"Motion Synthesis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/motion-synthesis-on-interhuman","task":"Motion Synthesis","dataset":"InterHuman","model":"MoMat-MoGen","rank_in_archive_order":3,"of":10,"metrics":{"FID":"5.674","MMDist":"3.790","MModality":"1.295","R-Precision Top3":"0.666"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2312.04547","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}