{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fb-bench-a-fine-grained-multi-task-benchmark","title":"FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback","arxiv_id":"2410.09412","date":"2024-10-12","proceeding":null,"authors":["Youquan Li","Miao Zheng","Fan Yang","Guosheng Dong","Bin Cui","WeiPeng Chen","Zenan Zhou","Wentao Zhang"],"abstract":"Human feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues, the user inputs are often independent, neglecting the nuanced and complex nature of human feedback within real-world usage scenarios. To fill this research gap, we introduce FB-Bench, a fine-grained, multi-task benchmark designed to evaluate LLMs' responsiveness to human feedback in real-world usage scenarios. Drawing from the two main interaction scenarios, FB-Bench comprises 734 meticulously curated samples, encompassing eight task types, five deficiency types of response, and nine feedback types. We extensively evaluate a broad array of popular LLMs, revealing significant variations in their performance across different interaction scenarios. Further analysis indicates that task, human feedback, and deficiencies of previous responses can also significantly impact LLMs' responsiveness. Our findings underscore both the strengths and limitations of current models, providing valuable insights and directions for future research. Both the toolkits and the dataset of FB-Bench are available at https://github.com/PKU-Baichuan-MLSystemLab/FB-Bench.","url_abs":"https://arxiv.org/abs/2410.09412v1","url_pdf":"https://arxiv.org/pdf/2410.09412v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fb-bench-a-fine-grained-multi-task-benchmark","repo_url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.09412","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.09412"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench","reach":null}],"summary":{"ran_draft_wrong":2,"ran_honours":2},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"c71ce00637f6ad07","entry":"check_gpt_judge","repo":"pku-baichuan-mlsystemlab/fb-bench","repo_kind":"official","path":"gen_judgment.py","file_url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench/blob/HEAD/gen_judgment.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c71ce00637f6ad07"}},{"code_sha256_prefix":"3b67de747085ba07","entry":"get_avg_score","repo":"pku-baichuan-mlsystemlab/fb-bench","repo_kind":"official","path":"show_results.py","file_url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench/blob/HEAD/show_results.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3b67de747085ba07"}},{"code_sha256_prefix":"d222b6a7d91f5a42","entry":"get_avg_score_of_nest_list","repo":"pku-baichuan-mlsystemlab/fb-bench","repo_kind":"official","path":"show_results.py","file_url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench/blob/HEAD/show_results.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d222b6a7d91f5a42"}},{"code_sha256_prefix":"f695219565a0c2ea","entry":"split_checklist","repo":"pku-baichuan-mlsystemlab/fb-bench","repo_kind":"official","path":"gen_judgment.py","file_url":"https://github.com/pku-baichuan-mlsystemlab/fb-bench/blob/HEAD/gen_judgment.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f695219565a0c2ea"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}