{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attributed-question-answering-evaluation-and","title":"Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models","arxiv_id":"2212.08037","date":"2022-12-15","proceeding":null,"authors":["Bernd Bohnet","Vinh Q. Tran","Pat Verga","Roee Aharoni","Daniel Andor","Livio Baldini Soares","Massimiliano Ciaramita","Jacob Eisenstein","Kuzman Ganchev","Jonathan Herzig","Kai Hui","Tom Kwiatkowski","Ji Ma","Jianmo Ni","Lierni Sestorain Saralegui","Tal Schuster","William W. Cohen","Michael Collins","Dipanjan Das","Donald Metzler","Slav Petrov","Kellie Webster"],"abstract":"Large language models (LLMs) have shown impressive results while requiring little or no direct supervision. Further, there is mounting evidence that LLMs may have potential in information-seeking scenarios. We believe the ability of an LLM to attribute the text that it generates is likely to be crucial in this setting. We formulate and study Attributed QA as a key first step in the development of attributed LLMs. We propose a reproducible evaluation framework for the task and benchmark a broad set of architectures. We take human annotations as a gold standard and show that a correlated automatic metric is suitable for development. Our experimental work gives concrete answers to two key questions (How to measure attribution?, and How well do current state-of-the-art methods perform on attribution?), and give some hints as to how to address a third (How to build LLMs with attribution?).","url_abs":"https://arxiv.org/abs/2212.08037v2","url_pdf":"https://arxiv.org/pdf/2212.08037v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attributed-question-answering-evaluation-and","repo_url":"https://github.com/google-research-datasets/attributed-qa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.08037","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2212.08037"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google-research-datasets/attributed-qa","reach":null}],"summary":{"ran_draft_wrong":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b31bec242d56bfde","entry":"format_passage_for_ais","repo":"google-research-datasets/attributed-qa","repo_kind":"official","path":"evaluation.py","file_url":"https://github.com/google-research-datasets/attributed-qa/blob/HEAD/evaluation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b31bec242d56bfde"}},{"code_sha256_prefix":"5e0a93b8b904b193","entry":"format_passage_for_autoais","repo":"google-research-datasets/attributed-qa","repo_kind":"official","path":"evaluation.py","file_url":"https://github.com/google-research-datasets/attributed-qa/blob/HEAD/evaluation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5e0a93b8b904b193"}},{"code_sha256_prefix":"98770e95557bde8e","entry":"read_wikipedia","repo":"google-research-datasets/attributed-qa","repo_kind":"official","path":"evaluation.py","file_url":"https://github.com/google-research-datasets/attributed-qa/blob/HEAD/evaluation.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"98770e95557bde8e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}