{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-faithfulness-and-factuality-in-abstractive","title":"On Faithfulness and Factuality in Abstractive Summarization","arxiv_id":"2005.00661","date":"2020-05-02","proceeding":"ACL 2020 6","authors":["Joshua Maynez","Shashi Narayan","Bernd Bohnet","Ryan Mcdonald"],"abstract":"It is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story generation. In this paper we have analyzed limitations of these models for abstractive document summarization and found that these models are highly prone to hallucinate content that is unfaithful to the input document. We conducted a large scale human evaluation of several neural abstractive summarization systems to better understand the types of hallucinations they produce. Our human annotators found substantial amounts of hallucinated content in all model generated summaries. However, our analysis does show that pretrained models are better summarizers not only in terms of raw metrics, i.e., ROUGE, but also in generating faithful and factual summaries as evaluated by humans. Furthermore, we show that textual entailment measures better correlate with faithfulness than standard metrics, potentially leading the way to automatic evaluation metrics as well as training and decoding criteria.","url_abs":"https://arxiv.org/abs/2005.00661v1","url_pdf":"https://arxiv.org/pdf/2005.00661v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-faithfulness-and-factuality-in-abstractive","repo_url":"https://github.com/google-research-datasets/xsum_hallucination_annotations","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"on-faithfulness-and-factuality-in-abstractive","repo_url":"https://github.com/tagoyal/factuality-datasets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"abstractive-text-summarization","task_name":"Abstractive Text Summarization"},{"task_slug":"document-summarization","task_name":"Document Summarization"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"},{"task_slug":"story-generation","task_name":"Story Generation"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2005.00661","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2005.00661"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google-research-datasets/xsum_hallucination_annotations","reach":{"status":"ok"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/EdinburghNLP/XSum","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tagoyal/factuality-datasets","reach":null}],"summary":{"unverified":2},"by_repo_kind":{"found_in_text":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d10052dcb1c9b468","entry":"get_training_stats","repo":"EdinburghNLP/XSum","repo_kind":"found_in_text","path":"XSum-ConvS2S/singleprocess_train.py","file_url":"https://github.com/EdinburghNLP/XSum/blob/HEAD/XSum-ConvS2S/singleprocess_train.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d10052dcb1c9b468"}},{"code_sha256_prefix":"1f42a3abd8d7fae5","entry":"get_valid_stats","repo":"EdinburghNLP/XSum","repo_kind":"found_in_text","path":"XSum-ConvS2S/singleprocess_train.py","file_url":"https://github.com/EdinburghNLP/XSum/blob/HEAD/XSum-ConvS2S/singleprocess_train.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1f42a3abd8d7fae5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}