{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2604-16242","title":"Detecting and Suppressing Reward Hacking with Gradient Fingerprints","arxiv_id":"2604.16242","date":"2026-04-17","proceeding":null,"authors":["Songtao Wang","Quang Hieu Pham","Fangcong Yin","Xinpeng Wang","Jocelyn Qiaochu Chen","Greg Durrett","Xi Ye"],"abstract":"Reinforcement learning with verifiable rewards (RLVR) typically optimizes for outcome rewards without imposing constraints on intermediate reasoning. This leaves training susceptible to reward hacking, where models exploit loopholes (e.g., spurious patterns in training data) in the reward function to achieve high scores without solving the intended task. These reward-hacking behaviors are often implicit, as the intermediate chain-of-thought (CoT) may appear plausible on the surface, limiting the effectiveness of purely text-based monitoring. We propose Gradient Fingerprint (GRIFT), a method for detecting reward hacking using models' internal computations. Given a prompt and a model-generated CoT, GRIFT computes gradients of the CoT conditioned on the prompt and compresses them into a compact representation, which is then used to assess whether the CoT reflects reward hacking behavior. Across verifiable reasoning benchmarks spanning math, code, and logical reasoning, GRIFT substantially outperforms strong baselines, including CoT Monitor and TRACE, achieving over 25% relative improvement in detecting reward hacking behavior. Moreover, integrating GRIFT into the rejection fine-tuning pipeline for reasoning tasks reduces reward hacking and improves performance on the true task objective. Our results highlight a promising direction of leveraging gradient level representations for assessing the quality of CoT reasoning traces. Our code is available at: https://github.com/songtao-x/reward_hack.","url_abs":"https://arxiv.org/abs/2604.16242","url_pdf":"https://arxiv.org/pdf/2604.16242","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2604.16242","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2604.16242"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/songtao-x/reward_hack","reach":null}],"summary":{"ran":1,"ran_fixture":1,"ran_draft_wrong":3,"ran_honours":2,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":8,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"5772557597b71ecb","entry":"QADatasetTorch","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5772557597b71ecb"}},{"code_sha256_prefix":"500f9ad2df12528c","entry":"flatten_lora_grads","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"500f9ad2df12528c"}},{"code_sha256_prefix":"20587dd88a857397","entry":"grad_hidden_to_fixed_vector","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"20587dd88a857397"}},{"code_sha256_prefix":"8b0e21f297af88b4","entry":"iter_lora_trainable_params","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8b0e21f297af88b4"}},{"code_sha256_prefix":"ebf641f187e6a57d","entry":"make_dataloader","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ebf641f187e6a57d"}},{"code_sha256_prefix":"207968a0d373e9f2","entry":"make_dense_pi","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"207968a0d373e9f2"}},{"code_sha256_prefix":"d6e393d3e4896c4e","entry":"masked_mean_pool_token_vectors","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d6e393d3e4896c4e"}},{"code_sha256_prefix":"72eaebdb315b454f","entry":"get_gradients_over_dataset","repo":"songtao-x/reward_hack","repo_kind":"found_in_text","path":"arlsat/icl/gradient/gradient_h.py","file_url":"https://github.com/songtao-x/reward_hack/blob/HEAD/arlsat/icl/gradient/gradient_h.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"72eaebdb315b454f"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6264},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}