{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cascade-reward-sampling-for-efficient","title":"Cascade Reward Sampling for Efficient Decoding-Time Alignment","arxiv_id":"2406.16306","date":"2024-06-24","proceeding":null,"authors":["Bolian Li","Yifan Wang","Anamika Lochab","Ananth Grama","Ruqi Zhang"],"abstract":"Aligning large language models (LLMs) with human preferences is essential for their applications. Recently, decoding-time alignment has emerged as an effective plug-and-play technique that avoids fine-tuning model parameters. This approach retains the general utility of pretrained LLMs but often suffers from significant inefficiencies during decoding, primarily due to wasted token generation and excessive reward evaluations. To address these challenges, we introduce Cascade Reward Sampling (CARDS) to resolve both efficiency bottlenecks in decoding-time alignment. Specifically, we develop a segment-level rejection sampling algorithm that minimizes redundant computations of both LLMs and reward models (RMs). Central to CARDS is an uncertainty-based segmentation mechanism, which ensures the accuracy of RMs evaluations on incomplete segments. Furthermore, we provide a detailed analysis of reward scores on segments to elucidate the improved alignment performance. Experimental results demonstrate that CARDS significantly improves decoding efficiency, alignment quality, and general utility compared to existing decoding-time alignment methods, achieving approximately a 70% reduction in decoding time and over 90% win-ties in utility and safety benchmarks.","url_abs":"https://arxiv.org/abs/2406.16306v2","url_pdf":"https://arxiv.org/pdf/2406.16306v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cascade-reward-sampling-for-efficient","repo_url":"https://github.com/lblaoke/CARDS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2406.16306","atlas_url":"https://app.syntology.ai/?focus=2406.16306","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.16306"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lblaoke/CARDS","reach":{"status":"ok"}}],"summary":{"ran":7,"unverified":3},"by_repo_kind":{"official":{"samples":10,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"9d741ca160b48466","entry":"clean","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/eval_wintie.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/eval_wintie.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9d741ca160b48466"}},{"code_sha256_prefix":"3cf2a6616e939e43","entry":"clean","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/metric.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/metric.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3cf2a6616e939e43"}},{"code_sha256_prefix":"cf6e7912b92ce301","entry":"extract_content","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/eval_ai.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/eval_ai.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cf6e7912b92ce301"}},{"code_sha256_prefix":"84540b0af0e4d7ea","entry":"gpt4_eval","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/eval_wintie.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/eval_wintie.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"84540b0af0e4d7ea"}},{"code_sha256_prefix":"3ab2f8d7ec77d909","entry":"load_responses","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/eval_ai.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/eval_ai.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3ab2f8d7ec77d909"}},{"code_sha256_prefix":"e24a02f57af303dc","entry":"parse_json","repo":"lblaoke/CARDS","repo_kind":"official","path":"data_loader.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/data_loader.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e24a02f57af303dc"}},{"code_sha256_prefix":"514c234b6e525f37","entry":"parse_plain","repo":"lblaoke/CARDS","repo_kind":"official","path":"data_loader.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/data_loader.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"514c234b6e525f37"}},{"code_sha256_prefix":"93c198d58d919707","entry":"compute_diversity","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/metric.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/metric.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"93c198d58d919707"}},{"code_sha256_prefix":"6c1c5b07ef997d49","entry":"compute_rep_n","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/metric.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/metric.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6c1c5b07ef997d49"}},{"code_sha256_prefix":"160c4a9ab3f8bda1","entry":"extract_out","repo":"lblaoke/CARDS","repo_kind":"official","path":"evaluation/measure_reward.py","file_url":"https://github.com/lblaoke/CARDS/blob/HEAD/evaluation/measure_reward.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"160c4a9ab3f8bda1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}