{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/validating-llm-as-a-judge-systems-in-the","title":"Validating LLM-as-a-Judge Systems in the Absence of Gold Labels","arxiv_id":"2503.05965","date":"2025-03-07","proceeding":null,"authors":["Luke Guerdan","Solon Barocas","Kenneth Holstein","Hanna Wallach","Zhiwei Steven Wu","Alexandra Chouldechova"],"abstract":"The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, has come to play a critical role in scaling and standardizing GenAI evaluations. To validate judge systems, evaluators collect multiple human ratings for each item in a validation corpus, and then aggregate the ratings into a single, per-item gold label rating. High agreement rates between these gold labels and judge system ratings are then taken as a sign of good judge system performance. In many cases, however, items or rating criteria may be ambiguous, or there may be principled disagreement among human raters. In such settings, gold labels may not exist for many of the items. In this paper, we introduce a framework for LLM-as-a-judge validation in the absence of gold labels. We present a theoretical analysis drawing connections between different measures of judge system performance under different rating elicitation and aggregation schemes. We also demonstrate empirically that existing validation approaches can select judge systems that are highly suboptimal, performing as much as 34% worse than the systems selected by alternative approaches that we describe. Based on our findings, we provide concrete recommendations for developing more reliable approaches to LLM-as-a-judge validation.","url_abs":"https://arxiv.org/abs/2503.05965v2","url_pdf":"https://arxiv.org/pdf/2503.05965v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.05965","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.05965"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/lguerdan/indeterminacy","reach":null}],"summary":{"ran_honours":1,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"cbf9ee53e8bdafde","entry":"construct_rs_option_lookup","repo":"lguerdan/indeterminacy","repo_kind":"found_in_text","path":"core/validator.py","file_url":"https://github.com/lguerdan/indeterminacy/blob/HEAD/core/validator.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cbf9ee53e8bdafde"}},{"code_sha256_prefix":"29cd5a4c17c5b78e","entry":"RatingModel","repo":"lguerdan/indeterminacy","repo_kind":"found_in_text","path":"core/validator.py","file_url":"https://github.com/lguerdan/indeterminacy/blob/HEAD/core/validator.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"29cd5a4c17c5b78e"}},{"code_sha256_prefix":"69468298a541caac","entry":"Validator","repo":"lguerdan/indeterminacy","repo_kind":"found_in_text","path":"core/validator.py","file_url":"https://github.com/lguerdan/indeterminacy/blob/HEAD/core/validator.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"69468298a541caac"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}