{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vmbench-a-benchmark-for-perception-aligned","title":"VMBench: A Benchmark for Perception-Aligned Video Motion Generation","arxiv_id":"2503.10076","date":"2025-03-13","proceeding":null,"authors":["Xinrang Ling","Chen Zhu","Meiqi Wu","Hangyu Li","Xiaokun Feng","Cundian Yang","Aiming Hao","Jiashu Zhu","JiaHong Wu","Xiangxiang Chu"],"abstract":"Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human perceptions; 2) the existing motion prompts are limited. Based on these findings, we introduce VMBench--a comprehensive Video Motion Benchmark that has perception-aligned motion metrics and features the most diverse types of motion. VMBench has several appealing properties: 1) Perception-Driven Motion Evaluation Metrics, we identify five dimensions based on human perception in motion video assessment and develop fine-grained evaluation metrics, providing deeper insights into models' strengths and weaknesses in motion quality. 2) Meta-Guided Motion Prompt Generation, a structured method that extracts meta-information, generates diverse motion prompts with LLMs, and refines them through human-AI validation, resulting in a multi-level prompt library covering six key dynamic scene dimensions. 3) Human-Aligned Validation Mechanism, we provide human preference annotations to validate our benchmarks, with our metrics achieving an average 35.3% improvement in Spearman's correlation over baseline methods. This is the first time that the quality of motion in videos has been evaluated from the perspective of human perception alignment. Additionally, we will soon release VMBench at https://github.com/GD-AIGC/VMBench, setting a new standard for evaluating and advancing motion generation models.","url_abs":"https://arxiv.org/abs/2503.10076v1","url_pdf":"https://arxiv.org/pdf/2503.10076v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vmbench-a-benchmark-for-perception-aligned","repo_url":"https://github.com/gd-aigc/vmbench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"motion-generation","task_name":"Motion Generation"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":null,"method_name":"Library"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.10076","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.10076"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/gd-aigc/vmbench","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/AMAP-ML/VMBench","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":6},"by_repo_kind":{"found_in_text":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c5ea12db107359b3","entry":"extract_frames_from_video","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"temporal_coherence_score.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/temporal_coherence_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c5ea12db107359b3"}},{"code_sha256_prefix":"dccf6af3b73d6aad","entry":"get_artifacts_frames","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"motion_smoothness_score.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/motion_smoothness_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"dccf6af3b73d6aad"}},{"code_sha256_prefix":"29416043c7035e7c","entry":"get_loss_scale_for_deepspeed","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"VideoMAEv2/engine_for_finetuning.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/VideoMAEv2/engine_for_finetuning.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"29416043c7035e7c"}},{"code_sha256_prefix":"07eeaf67c614f428","entry":"object_info_to_dict","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"temporal_coherence_score.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/temporal_coherence_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"07eeaf67c614f428"}},{"code_sha256_prefix":"6f8f7b238244ce38","entry":"set_threshold","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"motion_smoothness_score.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/motion_smoothness_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6f8f7b238244ce38"}},{"code_sha256_prefix":"c7976ea27edc377a","entry":"train_class_batch","repo":"AMAP-ML/VMBench","repo_kind":"found_in_text","path":"VideoMAEv2/engine_for_finetuning.py","file_url":"https://github.com/AMAP-ML/VMBench/blob/HEAD/VideoMAEv2/engine_for_finetuning.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c7976ea27edc377a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}