{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-the-effect-of-reinforcement","title":"Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1","arxiv_id":"2503.24376","date":"2025-03-31","proceeding":null,"authors":["Yi Chen","Yuying Ge","Rui Wang","Yixiao Ge","Lu Qiu","Ying Shan","Xihui Liu"],"abstract":"Recent advancements in Chain of Thought (COT) generation have significantly improved the reasoning capabilities of Large Language Models (LLMs), with reinforcement learning (RL) emerging as an effective post-training approach. Multimodal Large Language Models (MLLMs) inherit this reasoning potential but remain underexplored in tasks requiring both perception and logical reasoning. To address this, we introduce SEED-Bench-R1, a benchmark designed to systematically evaluate post-training methods for MLLMs in video understanding. It includes intricate real-world videos and complex everyday planning tasks in the format of multiple-choice questions, requiring sophisticated perception and reasoning. SEED-Bench-R1 assesses generalization through a three-level hierarchy: in-distribution, cross-environment, and cross-environment-task scenarios, equipped with a large-scale training dataset with easily verifiable ground-truth answers. Using Qwen2-VL-Instruct-7B as a base model, we compare RL with supervised fine-tuning (SFT), demonstrating RL's data efficiency and superior performance on both in-distribution and out-of-distribution tasks, even outperforming SFT on general video understanding benchmarks like LongVideoBench. Our detailed analysis reveals that RL enhances visual perception but often produces less logically coherent reasoning chains. We identify key limitations such as inconsistent reasoning and overlooked visual cues, and suggest future improvements in base model reasoning, reward modeling, and RL robustness against noisy signals.","url_abs":"https://arxiv.org/abs/2503.24376v1","url_pdf":"https://arxiv.org/pdf/2503.24376v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-the-effect-of-reinforcement","repo_url":"https://github.com/tencentarc/seed-bench-r1","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"logical-reasoning","task_name":"Logical Reasoning"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[{"method_slug":"base","method_name":"BASE"},{"method_slug":"sft","method_name":"SFT"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.24376","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.24376"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tencentarc/seed-bench-r1","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_honours":3,"unverified":6},"by_repo_kind":{"official":{"samples":9,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6e45201fa27cb24a","entry":"ceil_by_factor","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6e45201fa27cb24a"}},{"code_sha256_prefix":"8155263d7ff19bb3","entry":"floor_by_factor","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8155263d7ff19bb3"}},{"code_sha256_prefix":"e252767324188623","entry":"round_by_factor","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process_egoplan.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e252767324188623"}},{"code_sha256_prefix":"06c3869c64af2266","entry":"convert_example","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"src/open_r1_egoplan/sft.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/src/open_r1_egoplan/sft.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"06c3869c64af2266"}},{"code_sha256_prefix":"196b59b21bda8cbe","entry":"create_dataset_from_jsonl_simple","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"src/open_r1_egoplan/grpo.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/src/open_r1_egoplan/grpo.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"196b59b21bda8cbe"}},{"code_sha256_prefix":"b239a84dd5aac76c","entry":"create_dataset_from_jsonl_simple","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"src/open_r1_egoplan/sft.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/src/open_r1_egoplan/sft.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b239a84dd5aac76c"}},{"code_sha256_prefix":"1fd4dec0e26848af","entry":"format_reward","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"src/open_r1_egoplan/grpo.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/src/open_r1_egoplan/grpo.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1fd4dec0e26848af"}},{"code_sha256_prefix":"c1df6681680a10d2","entry":"make_conversation_egoplan","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"infer.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/infer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c1df6681680a10d2"}},{"code_sha256_prefix":"58ba34587ddd45bf","entry":"make_conversation_longvideobench","repo":"tencentarc/seed-bench-r1","repo_kind":"official","path":"infer.py","file_url":"https://github.com/tencentarc/seed-bench-r1/blob/HEAD/infer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"58ba34587ddd45bf"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}