{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rankalign-a-ranking-view-of-the-generator","title":"RankAlign: A Ranking View of the Generator-Validator Gap in Large Language Models","arxiv_id":"2504.11381","date":"2025-04-15","proceeding":null,"authors":["Juan Diego Rodriguez","Wenxuan Ding","Katrin Erk","Greg Durrett"],"abstract":"Although large language models (LLMs) have become generally more capable and accurate across many tasks, some fundamental sources of unreliability remain in their behavior. One key limitation is their inconsistency at reporting the the same information when prompts are changed. In this paper, we consider the discrepancy between a model's generated answer and their own verification of that answer, the generator-validator gap. We define this gap in a more stringent way than prior work: we expect correlation of scores from a generator and a validator over the entire set of candidate answers. We show that according to this measure, a large gap exists in various settings, including question answering, lexical semantics tasks, and next-word prediction. We then propose RankAlign, a ranking-based training method, and show that it significantly closes the gap by 31.8% on average, surpassing all baseline methods. Moreover, this approach generalizes well to out-of-domain tasks and lexical items.","url_abs":"https://arxiv.org/abs/2504.11381v1","url_pdf":"https://arxiv.org/pdf/2504.11381v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rankalign-a-ranking-view-of-the-generator","repo_url":"https://github.com/juand-r/rankalign","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2504.11381","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.11381"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/juand-r/rankalign","reach":null}],"summary":{"ran_honours":2,"ran_violates":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"b3f3784ee305065e","entry":"alpha_fun_1","repo":"juand-r/rankalign","repo_kind":"official","path":"scripts/ranking_loss_ref.py","file_url":"https://github.com/juand-r/rankalign/blob/HEAD/scripts/ranking_loss_ref.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b3f3784ee305065e"}},{"code_sha256_prefix":"4d21c17a8da02faa","entry":"get_alpha","repo":"juand-r/rankalign","repo_kind":"official","path":"scripts/ranking_loss_ref.py","file_url":"https://github.com/juand-r/rankalign/blob/HEAD/scripts/ranking_loss_ref.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d21c17a8da02faa"}},{"code_sha256_prefix":"9cb246c6ad8ea6f5","entry":"good_pair","repo":"juand-r/rankalign","repo_kind":"official","path":"scripts/ranking_loss_ref.py","file_url":"https://github.com/juand-r/rankalign/blob/HEAD/scripts/ranking_loss_ref.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9cb246c6ad8ea6f5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}