{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/expertqa-expert-curated-questions-and","title":"ExpertQA: Expert-Curated Questions and Attributed Answers","arxiv_id":"2309.07852","date":"2023-09-14","proceeding":null,"authors":["Chaitanya Malaviya","Subin Lee","Sihao Chen","Elizabeth Sieber","Mark Yatskar","Dan Roth"],"abstract":"As language models are adopted by a more sophisticated and diverse set of users, the importance of guaranteeing that they provide factually correct information supported by verifiable sources is critical across fields of study. This is especially the case for high-stakes fields, such as medicine and law, where the risk of propagating false information is high and can lead to undesirable societal consequences. Previous work studying attribution and factuality has not focused on analyzing these characteristics of language model outputs in domain-specific scenarios. In this work, we conduct human evaluation of responses from a few representative systems along various axes of attribution and factuality, by bringing domain experts in the loop. Specifically, we collect expert-curated questions from 484 participants across 32 fields of study, and then ask the same experts to evaluate generated responses to their own questions. In addition, we ask experts to improve upon responses from language models. The output of our analysis is ExpertQA, a high-quality long-form QA dataset with 2177 questions spanning 32 fields, along with verified answers and attributions for claims in the answers.","url_abs":"https://arxiv.org/abs/2309.07852v2","url_pdf":"https://arxiv.org/pdf/2309.07852v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"expertqa-expert-curated-questions-and","repo_url":"https://github.com/chaitanyamalaviya/expertqa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"expertqa-expert-curated-questions-and","repo_url":"https://github.com/ai21labs/factor","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"expertqa-expert-curated-questions-and","repo_url":"https://github.com/yangrui525/kg-rank","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"expertqa-expert-curated-questions-and","repo_url":"https://github.com/youngerous/qtree","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2309.07852","atlas_url":"https://app.syntology.ai/?focus=2309.07852","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.07852"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai21labs/factor","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chaitanyamalaviya/expertqa","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yangrui525/kg-rank","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/youngerous/qtree","reach":{"status":"ok"}}],"summary":{"ran_draft_wrong":3,"ran_honours":1,"ran_violates":1},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9a3be6f12cbbe3b2","entry":"find_urls","repo":"chaitanyamalaviya/expertqa","repo_kind":"official","path":"modeling/response_collection/fetch_bingchat_responses.py","file_url":"https://github.com/chaitanyamalaviya/expertqa/blob/HEAD/modeling/response_collection/fetch_bingchat_responses.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9a3be6f12cbbe3b2"}},{"code_sha256_prefix":"f66345e65929ad41","entry":"format_example_for_autoais","repo":"chaitanyamalaviya/expertqa","repo_kind":"official","path":"modeling/auto_attribution/autoais.py","file_url":"https://github.com/chaitanyamalaviya/expertqa/blob/HEAD/modeling/auto_attribution/autoais.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f66345e65929ad41"}},{"code_sha256_prefix":"b8be2a13d4b17f23","entry":"format_passage_for_autoais","repo":"chaitanyamalaviya/expertqa","repo_kind":"official","path":"modeling/auto_attribution/autoais.py","file_url":"https://github.com/chaitanyamalaviya/expertqa/blob/HEAD/modeling/auto_attribution/autoais.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b8be2a13d4b17f23"}},{"code_sha256_prefix":"38a0c176b9e9285a","entry":"get_entail_label_ids","repo":"chaitanyamalaviya/expertqa","repo_kind":"official","path":"modeling/auto_attribution/autoais.py","file_url":"https://github.com/chaitanyamalaviya/expertqa/blob/HEAD/modeling/auto_attribution/autoais.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"38a0c176b9e9285a"}},{"code_sha256_prefix":"6c4b02266edc91ba","entry":"get_score","repo":"chaitanyamalaviya/expertqa","repo_kind":"official","path":"modeling/fact_score/factscore.py","file_url":"https://github.com/chaitanyamalaviya/expertqa/blob/HEAD/modeling/fact_score/factscore.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6c4b02266edc91ba"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}