{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/clue-a-clinical-language-understanding","title":"Does Biomedical Training Lead to Better Medical Performance?","arxiv_id":"2404.04067","date":"2024-04-05","proceeding":null,"authors":["Amin Dada","Marie Bauer","Amanda Butler Contreras","Osman Alperen Koraş","Constantin Marc Seibold","Kaleb E Smith","Jens Kleesiek"],"abstract":"Large Language Models (LLMs) are expected to significantly contribute to patient care, diagnostics, and administrative processes. Emerging biomedical LLMs aim to address healthcare-specific challenges, including privacy demands and computational constraints. Assessing the models' suitability for this sensitive application area is of the utmost importance. However, biomedical training has not been systematically evaluated on medical tasks. This study investigates the effect of biomedical training in the context of six practical medical tasks evaluating $25$ models. In contrast to previous evaluations, our results reveal a performance decline in nine out of twelve biomedical models after fine-tuning, particularly on tasks involving hallucinations, ICD10 coding, and instruction adherence. General-domain models like Meta-Llama-3.1-70B-Instruct outperformed their biomedical counterparts, indicating a trade-off between domain-specific fine-tuning and general medical task performance. We open-source all evaluation scripts and datasets at https://github.com/TIO-IKIM/CLUE to support further research in this critical area.","url_abs":"https://arxiv.org/abs/2404.04067v4","url_pdf":"https://arxiv.org/pdf/2404.04067v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"clue-a-clinical-language-understanding","repo_url":"https://github.com/tio-ikim/clue","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.04067","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.04067"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tio-ikim/clue","reach":null}],"summary":{"ran_violates":1,"ran_draft_wrong":4},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0c7db73dea0a25ae","entry":"compute_metrics","repo":"tio-ikim/clue","repo_kind":"official","path":"eval/eval_LongHealth.py","file_url":"https://github.com/tio-ikim/clue/blob/HEAD/eval/eval_LongHealth.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0c7db73dea0a25ae"}},{"code_sha256_prefix":"92987fdeaa63d462","entry":"get_correct_answer","repo":"tio-ikim/clue","repo_kind":"official","path":"eval/eval_LongHealth.py","file_url":"https://github.com/tio-ikim/clue/blob/HEAD/eval/eval_LongHealth.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"92987fdeaa63d462"}},{"code_sha256_prefix":"70cb661f76d88166","entry":"parse_model_choice","repo":"tio-ikim/clue","repo_kind":"official","path":"eval/eval_LongHealth.py","file_url":"https://github.com/tio-ikim/clue/blob/HEAD/eval/eval_LongHealth.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"70cb661f76d88166"}},{"code_sha256_prefix":"b80a70a12a04cb9d","entry":"postprocess_llm_results","repo":"tio-ikim/clue","repo_kind":"official","path":"data/MeDiSumQA/generate_qas.py","file_url":"https://github.com/tio-ikim/clue/blob/HEAD/data/MeDiSumQA/generate_qas.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b80a70a12a04cb9d"}},{"code_sha256_prefix":"72277b101129d205","entry":"postprocess_qas","repo":"tio-ikim/clue","repo_kind":"official","path":"data/MeDiSumQA/generate_qas.py","file_url":"https://github.com/tio-ikim/clue/blob/HEAD/data/MeDiSumQA/generate_qas.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"72277b101129d205"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}