{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/empirical-evaluation-of-uncertainty","title":"Empirical evaluation of Uncertainty Quantification in Retrieval-Augmented Language Models for Science","arxiv_id":"2311.09358","date":"2023-11-15","proceeding":null,"authors":["Sridevi Wagle","Sai Munikoti","Anurag Acharya","Sara Smith","Sameera Horawalavithana"],"abstract":"Large language models (LLMs) have shown remarkable achievements in natural language processing tasks, producing high-quality outputs. However, LLMs still exhibit limitations, including the generation of factually incorrect information. In safety-critical applications, it is important to assess the confidence of LLM-generated content to make informed decisions. Retrieval Augmented Language Models (RALMs) is relatively a new area of research in NLP. RALMs offer potential benefits for scientific NLP tasks, as retrieved documents, can serve as evidence to support model-generated content. This inclusion of evidence enhances trustworthiness, as users can verify and explore the retrieved documents to validate model outputs. Quantifying uncertainty in RALM generations further improves trustworthiness, with retrieved text and confidence scores contributing to a comprehensive and reliable model for scientific applications. However, there is limited to no research on UQ for RALMs, particularly in scientific contexts. This study aims to address this gap by conducting a comprehensive evaluation of UQ in RALMs, focusing on scientific tasks. This research investigates how uncertainty scores vary when scientific knowledge is incorporated as pretraining and retrieval data and explores the relationship between uncertainty scores and the accuracy of model-generated outputs. We observe that an existing RALM finetuned with scientific knowledge as the retrieval data tends to be more confident in generating predictions compared to the model pretrained only with scientific knowledge. We also found that RALMs are overconfident in their predictions, making inaccurate predictions more confidently than accurate ones. Scientific knowledge provided either as pretraining or retrieval corpus does not help alleviate this issue. We released our code, data and dashboards at https://github.com/pnnl/EXPERT2.","url_abs":"https://arxiv.org/abs/2311.09358v1","url_pdf":"https://arxiv.org/pdf/2311.09358v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"empirical-evaluation-of-uncertainty","repo_url":"https://github.com/pnnl/expert2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-2-Clause"}}],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"uncertainty-quantification","task_name":"Uncertainty Quantification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2311.09358","atlas_url":"https://app.syntology.ai/?focus=2311.09358","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.09358"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pnnl/expert2","reach":{"status":"ok","spdx":"BSD-2-Clause"}}],"summary":{"ran":4,"unverified":2},"by_repo_kind":{"official":{"samples":6,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3205c4709d8055e2","entry":"get_alo_accuracy","repo":"pnnl/expert2","repo_kind":"official","path":"model/metrics_eval.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/metrics_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3205c4709d8055e2"}},{"code_sha256_prefix":"f4091badea744020","entry":"get_em_accuracy","repo":"pnnl/expert2","repo_kind":"official","path":"model/metrics_eval.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/metrics_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"f4091badea744020"}},{"code_sha256_prefix":"274a5f5175c05893","entry":"get_entropy","repo":"pnnl/expert2","repo_kind":"official","path":"model/evaluation_scripts/custom_evaluate.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/evaluation_scripts/custom_evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"274a5f5175c05893"}},{"code_sha256_prefix":"f60ebd1fa96c5391","entry":"get_pm_accuracy","repo":"pnnl/expert2","repo_kind":"official","path":"model/metrics_eval.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/metrics_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"f60ebd1fa96c5391"}},{"code_sha256_prefix":"210d1520a7ea9d7f","entry":"get_argument_value","repo":"pnnl/expert2","repo_kind":"official","path":"model/finetune_qa.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/finetune_qa.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"210d1520a7ea9d7f"}},{"code_sha256_prefix":"3d952765274002ec","entry":"set_parser_options","repo":"pnnl/expert2","repo_kind":"official","path":"model/finetune_qa.py","file_url":"https://github.com/pnnl/expert2/blob/HEAD/model/finetune_qa.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3d952765274002ec"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}