{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leveraging-large-language-models-for-1","title":"Leveraging Large Language Models for Automated Dialogue Analysis","arxiv_id":"2309.06490","date":"2023-09-12","proceeding":null,"authors":["Sarah E. Finch","Ellie S. Paek","Jinho D. Choi"],"abstract":"Developing high-performing dialogue systems benefits from the automatic identification of undesirable behaviors in system responses. However, detecting such behaviors remains challenging, as it draws on a breadth of general knowledge and understanding of conversational practices. Although recent research has focused on building specialized classifiers for detecting specific dialogue behaviors, the behavior coverage is still incomplete and there is a lack of testing on real-world human-bot interactions. This paper investigates the ability of a state-of-the-art large language model (LLM), ChatGPT-3.5, to perform dialogue behavior detection for nine categories in real human-bot dialogues. We aim to assess whether ChatGPT can match specialized models and approximate human performance, thereby reducing the cost of behavior detection tasks. Our findings reveal that neither specialized models nor ChatGPT have yet achieved satisfactory results for this task, falling short of human performance. Nevertheless, ChatGPT shows promising potential and often outperforms specialized detection models. We conclude with an in-depth examination of the prevalent shortcomings of ChatGPT, offering guidance for future research to enhance LLM capabilities.","url_abs":"https://arxiv.org/abs/2309.06490v1","url_pdf":"https://arxiv.org/pdf/2309.06490v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"leveraging-large-language-models-for-1","repo_url":"https://github.com/emorynlp/gpt-abceval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"general-knowledge","task_name":"General Knowledge"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.06490","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.06490"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/emorynlp/gpt-abceval","reach":{"status":"ok"}}],"summary":{"ran":2,"unverified":3},"by_repo_kind":{"official":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"5147c781ddd64cf8","entry":"bootstrap_ci","repo":"emorynlp/gpt-abceval","repo_kind":"official","path":"gpt_interface/metric_utils.py","file_url":"https://github.com/emorynlp/gpt-abceval/blob/HEAD/gpt_interface/metric_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5147c781ddd64cf8"}},{"code_sha256_prefix":"2c3c80cf9b4ef1fd","entry":"gpt","repo":"emorynlp/gpt-abceval","repo_kind":"official","path":"gpt_interface/gpt_utils.py","file_url":"https://github.com/emorynlp/gpt-abceval/blob/HEAD/gpt_interface/gpt_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2c3c80cf9b4ef1fd"}},{"code_sha256_prefix":"97d1a59c464a8626","entry":"get_overall_stats","repo":"emorynlp/gpt-abceval","repo_kind":"official","path":"gpt_interface/analyze_results.py","file_url":"https://github.com/emorynlp/gpt-abceval/blob/HEAD/gpt_interface/analyze_results.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"97d1a59c464a8626"}},{"code_sha256_prefix":"53e63d2b469268ca","entry":"is_labelled","repo":"emorynlp/gpt-abceval","repo_kind":"official","path":"gpt_interface/load_data.py","file_url":"https://github.com/emorynlp/gpt-abceval/blob/HEAD/gpt_interface/load_data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"53e63d2b469268ca"}},{"code_sha256_prefix":"b5b8238cdd662cba","entry":"num_tokens_from_messages","repo":"emorynlp/gpt-abceval","repo_kind":"official","path":"gpt_interface/gpt_utils.py","file_url":"https://github.com/emorynlp/gpt-abceval/blob/HEAD/gpt_interface/gpt_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b5b8238cdd662cba"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}