{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/motionllm-understanding-human-behaviors-from","title":"MotionLLM: Understanding Human Behaviors from Human Motions and Videos","arxiv_id":"2405.20340","date":"2024-05-30","proceeding":null,"authors":["Ling-Hao Chen","Shunlin Lu","Ailing Zeng","Hao Zhang","Benyou Wang","Ruimao Zhang","Lei Zhang"],"abstract":"This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs designed for video-only or motion-only understanding, we argue that understanding human behavior necessitates joint modeling from both videos and motion sequences (e.g., SMPL sequences) to capture nuanced body part dynamics and semantics effectively. In light of this, we present MotionLLM, a straightforward yet effective framework for human motion understanding, captioning, and reasoning. Specifically, MotionLLM adopts a unified video-motion training strategy that leverages the complementary advantages of existing coarse video-text data and fine-grained motion-text data to glean rich spatial-temporal insights. Furthermore, we collect a substantial dataset, MoVid, comprising diverse videos, motions, captions, and instructions. Additionally, we propose the MoVid-Bench, with carefully manual annotations, for better evaluation of human behavior understanding on video and motion. Extensive experiments show the superiority of MotionLLM in the caption, spatial-temporal comprehension, and reasoning ability.","url_abs":"https://arxiv.org/abs/2405.20340v1","url_pdf":"https://arxiv.org/pdf/2405.20340v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"motionllm-understanding-human-behaviors-from","repo_url":"https://github.com/IDEA-Research/MotionLLM","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2405.20340","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.20340"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IDEA-Research/MotionLLM","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran_draft_wrong":1,"ran":2,"unverified":2},"by_repo_kind":{"official":{"samples":5,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"5728c74084e12d35","entry":"apply_rope","repo":"IDEA-Research/MotionLLM","repo_kind":"official","path":"lit_gpt/model.py","file_url":"https://github.com/IDEA-Research/MotionLLM/blob/HEAD/lit_gpt/model.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5728c74084e12d35"}},{"code_sha256_prefix":"57b678d36ceb14f0","entry":"build_rope_cache","repo":"IDEA-Research/MotionLLM","repo_kind":"official","path":"lit_llama/model.py","file_url":"https://github.com/IDEA-Research/MotionLLM/blob/HEAD/lit_llama/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"57b678d36ceb14f0"}},{"code_sha256_prefix":"688ce44db612be08","entry":"generate","repo":"IDEA-Research/MotionLLM","repo_kind":"official","path":"generate.py","file_url":"https://github.com/IDEA-Research/MotionLLM/blob/HEAD/generate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"688ce44db612be08"}},{"code_sha256_prefix":"ebb19c1f412b0e2a","entry":"apply_rope","repo":"IDEA-Research/MotionLLM","repo_kind":"official","path":"lit_llama/model.py","file_url":"https://github.com/IDEA-Research/MotionLLM/blob/HEAD/lit_llama/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ebb19c1f412b0e2a"}},{"code_sha256_prefix":"c0cabd60c81402a5","entry":"build_rope_cache","repo":"IDEA-Research/MotionLLM","repo_kind":"official","path":"lit_gpt/model.py","file_url":"https://github.com/IDEA-Research/MotionLLM/blob/HEAD/lit_gpt/model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c0cabd60c81402a5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}