{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ai-hospital-interactive-evaluation-and","title":"AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator","arxiv_id":"2402.09742","date":"2024-02-15","proceeding":null,"authors":["Zhihao Fan","Jialong Tang","Wei Chen","Siyuan Wang","Zhongyu Wei","Jun Xi","Fei Huang","Jingren Zhou"],"abstract":"Artificial intelligence has significantly advanced healthcare, particularly through large language models (LLMs) that excel in medical question answering benchmarks. However, their real-world clinical application remains limited due to the complexities of doctor-patient interactions. To address this, we introduce \\textbf{AI Hospital}, a multi-agent framework simulating dynamic medical interactions between \\emph{Doctor} as player and NPCs including \\emph{Patient}, \\emph{Examiner}, \\emph{Chief Physician}. This setup allows for realistic assessments of LLMs in clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and NPCs to evaluate LLMs' performance in symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance diagnostic accuracy through iterative discussions. Despite improvements, current LLMs exhibit significant performance gaps in multi-turn interactions compared to one-step approaches. Our findings highlight the need for further research to bridge these gaps and improve LLMs' clinical diagnostic capabilities. Our data, code, and experimental results are all open-sourced at \\url{https://github.com/LibertFan/AI_Hospital}.","url_abs":"https://arxiv.org/abs/2402.09742v4","url_pdf":"https://arxiv.org/pdf/2402.09742v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ai-hospital-interactive-evaluation-and","repo_url":"https://github.com/LibertFan/AI_Hospital","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":null,"task_name":"Medical Question Answering"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[{"slug":"mvme","name":"MVME","full_name":"Multi-View Medical Evaluation Benchmark"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2402.09742","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}