{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/are-machine-rationales-not-useful-to-humans","title":"Are Machine Rationales (Not) Useful to Humans? Measuring and Improving Human Utility of Free-Text Rationales","arxiv_id":"2305.07095","date":"2023-05-11","proceeding":null,"authors":["Brihi Joshi","Ziyi Liu","Sahana Ramnath","Aaron Chan","Zhewei Tong","Shaoliang Nie","Qifan Wang","Yejin Choi","Xiang Ren"],"abstract":"Among the remarkable emergent capabilities of large language models (LMs) is free-text rationalization; beyond a certain scale, large LMs are capable of generating seemingly useful rationalizations, which in turn, can dramatically enhance their performances on leaderboards. This phenomenon raises a question: can machine generated rationales also be useful for humans, especially when lay humans try to answer questions based on those machine rationales? We observe that human utility of existing rationales is far from satisfactory, and expensive to estimate with human studies. Existing metrics like task performance of the LM generating the rationales, or similarity between generated and gold rationales are not good indicators of their human utility. While we observe that certain properties of rationales like conciseness and novelty are correlated with their human utility, estimating them without human involvement is challenging. We show that, by estimating a rationale's helpfulness in answering similar unseen instances, we can measure its human utility to a better extent. We also translate this finding into an automated score, GEN-U, that we propose, which can help improve LMs' ability to generate rationales with better human utility, while maintaining most of its task performance. Lastly, we release all code and collected data with this project.","url_abs":"https://arxiv.org/abs/2305.07095v1","url_pdf":"https://arxiv.org/pdf/2305.07095v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"are-machine-rationales-not-useful-to-humans","repo_url":"https://github.com/ink-usc/rationalehumanutility","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.07095","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.07095"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ink-usc/rationalehumanutility","reach":null}],"summary":{"ran_draft_wrong":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"d32d5bfe045233e9","entry":"_parse_I_RO_response","repo":"ink-usc/rationalehumanutility","repo_kind":"official","path":"src/utils/gpt3_utils.py","file_url":"https://github.com/ink-usc/rationalehumanutility/blob/HEAD/src/utils/gpt3_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d32d5bfe045233e9"}},{"code_sha256_prefix":"6139790c77878a35","entry":"_parse_label_text","repo":"ink-usc/rationalehumanutility","repo_kind":"official","path":"src/utils/gpt3_utils.py","file_url":"https://github.com/ink-usc/rationalehumanutility/blob/HEAD/src/utils/gpt3_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6139790c77878a35"}},{"code_sha256_prefix":"d6bd64117606b3ed","entry":"output_label_mapper","repo":"ink-usc/rationalehumanutility","repo_kind":"official","path":"src/utils/gpt3_utils.py","file_url":"https://github.com/ink-usc/rationalehumanutility/blob/HEAD/src/utils/gpt3_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d6bd64117606b3ed"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}