{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2510-10963","title":"APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport","arxiv_id":"2510.10963","date":"2025-10-13","proceeding":"EMNLP","authors":["Zhuo Li","Yuege Feng","Dandan Guo","Jinpeng Hu","Anningzhe Gao","Xiang Wan"],"abstract":"The reward model (RM) plays a crucial role in aligning Large Language Models (LLMs) with human preferences through Reinforcement Learning, where the Bradley-Terry (BT) objective has been recognized as simple yet powerful, specifically for pairwise preference learning. However, BT-based RMs often struggle to effectively distinguish between similar preference responses, leading to insufficient separation between preferred and non-preferred outputs. Consequently, they may easily overfit easy samples and cannot generalize well to Out-Of-Distribution (OOD) samples, resulting in suboptimal performance. To address these challenges, this paper introduces an effective enhancement to BT-based RMs through an adaptive margin mechanism. Specifically, we design to dynamically adjust the RM focus on more challenging samples through margins, based on both semantic similarity and model-predicted reward differences, which is approached from a distributional perspective solvable with Optimal Transport (OT). By incorporating these factors into a principled OT cost matrix design, our adaptive margin enables the RM to better capture distributional differences between chosen and rejected responses, yielding significant improvements in performance, convergence speed, and generalization capabilities. Experimental results across multiple benchmarks demonstrate that our method outperforms several existing RM techniques, showcasing enhanced performance in both In-Distribution (ID) and OOD settings. Moreover, RLHF experiments support our practical effectiveness in better aligning LLMs with human preferences. Our code is available at https://github.com/BIRlz/APLOT","url_abs":"https://arxiv.org/abs/2510.10963","url_pdf":"https://arxiv.org/pdf/2510.10963","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2510.10963","atlas_url":"https://app.syntology.ai/?focus=2510.10963","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2510.10963"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/BIRlz/APLOT","reach":{"status":"ok"}}],"summary":{"unverified":9},"by_repo_kind":{"found_in_text":{"samples":9,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"47ee1e29f8388bb2","entry":"build_dataset","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/load_datasets.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/load_datasets.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"47ee1e29f8388bb2"}},{"code_sha256_prefix":"d5c8de0d3609ddc2","entry":"build_dataset_SK","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/load_datasets.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/load_datasets.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d5c8de0d3609ddc2"}},{"code_sha256_prefix":"c7dd18f13397a20b","entry":"build_dataset_UF","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/load_datasets.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/load_datasets.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c7dd18f13397a20b"}},{"code_sha256_prefix":"7710f1274273fc6c","entry":"compute_metrics","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/utils.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7710f1274273fc6c"}},{"code_sha256_prefix":"9eb0f896e194e1da","entry":"cosine_similarity_matrix","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/aplot_trainer.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/aplot_trainer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9eb0f896e194e1da"}},{"code_sha256_prefix":"aed395f45af87c9c","entry":"get_trainable_weights","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/utils.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"aed395f45af87c9c"}},{"code_sha256_prefix":"f8e7d855924e7e64","entry":"is_lora_model","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/utils.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f8e7d855924e7e64"}},{"code_sha256_prefix":"97f43d7c928e992e","entry":"post_process_RBv1","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/aplot_trainer.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/aplot_trainer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"97f43d7c928e992e"}},{"code_sha256_prefix":"66d572c3fb9e65bc","entry":"post_process_RMBench","repo":"BIRlz/APLOT","repo_kind":"found_in_text","path":"reward_models/aplot_trainer.py","file_url":"https://github.com/BIRlz/APLOT/blob/HEAD/reward_models/aplot_trainer.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"66d572c3fb9e65bc"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_api"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9623,"papers_checked":6885},"entries":[],"not_placed":{"boards":1,"rejected_by_independent_check":0,"refused_by_a_rule":1,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}