{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2607-28077","title":"LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models","arxiv_id":"2607.28077","date":"2026-07-30","proceeding":null,"authors":["Shuang Liang","Haoyang Zhou","Yifan Gong","Guowei Wang","Xiting Wang"],"abstract":"Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore--Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6\\% and 3.7\\% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at https://github.com/ShuangLiangX/LEEPS.","url_abs":"https://arxiv.org/abs/2607.28077","url_pdf":"https://arxiv.org/pdf/2607.28077","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2607.28077","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2607.28077"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/ShuangLiangX/LEEPS","reach":null}],"summary":{"ran_violates":4,"ran_honours":1,"ran_draft_wrong":3,"ran":1,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":10,"ran":9,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f894b12627feeb71","entry":"_as_bool","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f894b12627feeb71"}},{"code_sha256_prefix":"928b8813c0c6ff67","entry":"_as_float","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"928b8813c0c6ff67"}},{"code_sha256_prefix":"57ae4bb54b785f88","entry":"_cfg_get","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"57ae4bb54b785f88"}},{"code_sha256_prefix":"4644280371daaab5","entry":"_normalize_id","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4644280371daaab5"}},{"code_sha256_prefix":"d3cf54397bde9ee3","entry":"_slice_value","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d3cf54397bde9ee3"}},{"code_sha256_prefix":"9b03d151644f93ab","entry":"_tasksampler_cfg","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9b03d151644f93ab"}},{"code_sha256_prefix":"ccd102431be91db5","entry":"_to_embedding_matrix","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ccd102431be91db5"}},{"code_sha256_prefix":"f20078df7a8b992b","entry":"_to_numpy_indices","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f20078df7a8b992b"}},{"code_sha256_prefix":"4bb9a4bd0cf9085d","entry":"l2_normalize","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4bb9a4bd0cf9085d"}},{"code_sha256_prefix":"4537898ad5f25bd4","entry":"KNNExploreExploitSampler","repo":"ShuangLiangX/LEEPS","repo_kind":"found_in_text","path":"recipe/leeps/task_sampler.py","file_url":"https://github.com/ShuangLiangX/LEEPS/blob/HEAD/recipe/leeps/task_sampler.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4537898ad5f25bd4"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}