{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ttrl-test-time-reinforcement-learning","title":"TTRL: Test-Time Reinforcement Learning","arxiv_id":"2504.16084","date":"2025-04-22","proceeding":null,"authors":["Yuxin Zuo","Kaiyan Zhang","Shang Qu","Li Sheng","Xuekai Zhu","Biqing Qi","Youbang Sun","Ganqu Cui","Ning Ding","BoWen Zhou"],"abstract":"This paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 159% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks, and highlight TTRL's potential for broader tasks and domains. GitHub: https://github.com/PRIME-RL/TTRL","url_abs":"https://arxiv.org/abs/2504.16084v1","url_pdf":"https://arxiv.org/pdf/2504.16084v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ttrl-test-time-reinforcement-learning","repo_url":"https://github.com/prime-rl/ttrl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"ttrl-test-time-reinforcement-learning","repo_url":"https://github.com/tsinghuac3i/awesome-rl-reasoning-recipes","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"ttrl-test-time-reinforcement-learning","repo_url":"https://github.com/qingyangzhang/empo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"math","task_name":"Math"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2504.16084","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.16084"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/prime-rl/ttrl","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qingyangzhang/empo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tsinghuac3i/awesome-rl-reasoning-recipes","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":7,"unverified":9},"by_repo_kind":{"official":{"samples":16,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f220595adfc5e83c","entry":"compute_ce_dpo_loss_rm","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f220595adfc5e83c"}},{"code_sha256_prefix":"4d91e4383716e42c","entry":"compute_detach_dpo_loss_rm","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4d91e4383716e42c"}},{"code_sha256_prefix":"f19825983a523003","entry":"compute_online_dpo_loss","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f19825983a523003"}},{"code_sha256_prefix":"41b6e81dc3ecba79","entry":"compute_onlinedpo_pref","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"41b6e81dc3ecba79"}},{"code_sha256_prefix":"9c496dae615ba2a8","entry":"generate_config_from_args","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/model_merger/base_model_merger.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/model_merger/base_model_merger.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9c496dae615ba2a8"}},{"code_sha256_prefix":"757d109d90121142","entry":"get_kl_controller","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"757d109d90121142"}},{"code_sha256_prefix":"c30b2a96103abd2c","entry":"reward_func","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/recipe/r1/reward_score.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/recipe/r1/reward_score.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c30b2a96103abd2c"}},{"code_sha256_prefix":"a5a4390c25526df6","entry":"apply_original_gt","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/trainer/ppo/ttrl_utils.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/trainer/ppo/ttrl_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a5a4390c25526df6"}},{"code_sha256_prefix":"bb175bff1060d928","entry":"handle_base","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/grader.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/grader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bb175bff1060d928"}},{"code_sha256_prefix":"2a180d3156997924","entry":"is_digit","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/grader.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/grader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2a180d3156997924"}},{"code_sha256_prefix":"3f78ce44e20c50d1","entry":"mathd_normalize_answer","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/math_utils.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/math_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3f78ce44e20c50d1"}},{"code_sha256_prefix":"1e5adbd8b9d51c04","entry":"normalize","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/grader.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/grader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1e5adbd8b9d51c04"}},{"code_sha256_prefix":"31113f93a6e0e98a","entry":"normalize_answer","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/math_normalize.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/math_normalize.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"31113f93a6e0e98a"}},{"code_sha256_prefix":"5662e5ee755b541c","entry":"normalize_final_answer","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/math_utils.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/math_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5662e5ee755b541c"}},{"code_sha256_prefix":"50e288f719695c06","entry":"select_top_k_per_prompt","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/trainer/ppo/ttrl_utils.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/trainer/ppo/ttrl_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"50e288f719695c06"}},{"code_sha256_prefix":"bc7d12a7bef2d7e8","entry":"timeout_ours","repo":"prime-rl/ttrl","repo_kind":"official","path":"verl/verl/utils/reward_score/ttrl_math/math_utils.py","file_url":"https://github.com/prime-rl/ttrl/blob/HEAD/verl/verl/utils/reward_score/ttrl_math/math_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bc7d12a7bef2d7e8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}