{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-correctness-and-faithfulness-of","title":"Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering","arxiv_id":"2307.16877","date":"2023-07-31","proceeding":null,"authors":["Vaibhav Adlakha","Parishad BehnamGhader","Xing Han Lu","Nicholas Meade","Siva Reddy"],"abstract":"Retriever-augmented instruction-following models are attractive alternatives to fine-tuned approaches for information-seeking tasks such as question answering (QA). By simply prepending retrieved documents in its input along with an instruction, these models can be adapted to various information domains and tasks without additional fine-tuning. While the model responses tend to be natural and fluent, the additional verbosity makes traditional QA evaluation metrics such as exact match (EM) and F1 unreliable for accurately quantifying model performance. In this work, we investigate the performance of instruction-following models across three information-seeking QA tasks. We use both automatic and human evaluation to evaluate these models along two dimensions: 1) how well they satisfy the user's information need (correctness), and 2) whether they produce a response based on the provided knowledge (faithfulness). Guided by human evaluation and analysis, we highlight the shortcomings of traditional metrics for both correctness and faithfulness. We then propose simple token-overlap based and model-based metrics that reflect the true performance of these models. Our analysis reveals that instruction-following models are competitive, and sometimes even outperform fine-tuned models for correctness. However, these models struggle to stick to the provided knowledge and often hallucinate in their responses. We hope our work encourages a more holistic evaluation of instruction-following models for QA. Our code and data is available at https://github.com/McGill-NLP/instruct-qa","url_abs":"https://arxiv.org/abs/2307.16877v2","url_pdf":"https://arxiv.org/pdf/2307.16877v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-correctness-and-faithfulness-of","repo_url":"https://github.com/mcgill-nlp/instruct-qa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2307.16877","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.16877"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mcgill-nlp/instruct-qa","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"235ff5681c4da416","entry":"generate_experiment_id","repo":"mcgill-nlp/instruct-qa","repo_kind":"official","path":"instruct_qa/experiment_utils.py","file_url":"https://github.com/mcgill-nlp/instruct-qa/blob/HEAD/instruct_qa/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"235ff5681c4da416"}},{"code_sha256_prefix":"cc3246edc7751304","entry":"normalize_passage","repo":"mcgill-nlp/instruct-qa","repo_kind":"official","path":"instruct_qa/collections/dpr_wiki_collection.py","file_url":"https://github.com/mcgill-nlp/instruct-qa/blob/HEAD/instruct_qa/collections/dpr_wiki_collection.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cc3246edc7751304"}},{"code_sha256_prefix":"91a9d14080102d04","entry":"parse_experiment_id","repo":"mcgill-nlp/instruct-qa","repo_kind":"official","path":"instruct_qa/experiment_utils.py","file_url":"https://github.com/mcgill-nlp/instruct-qa/blob/HEAD/instruct_qa/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"91a9d14080102d04"}},{"code_sha256_prefix":"90f97c979eeb879e","entry":"wget","repo":"mcgill-nlp/instruct-qa","repo_kind":"official","path":"instruct_qa/experiment_utils.py","file_url":"https://github.com/mcgill-nlp/instruct-qa/blob/HEAD/instruct_qa/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"90f97c979eeb879e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}