{"url":"/task/medical-code-prediction","name":"Medical Code Prediction","slug":"medical-code-prediction","description_markdown":"Context: Prediction of medical codes from clinical notes is both a practical and essential need for every healthcare delivery organization within current medical systems. Automating annotation will save significant time and excessive effort by human coders today. A new milestone will mark a meaningful step toward fully Autonomous Medical Coding in machines reaching parity with human coders' performance in medical code prediction.\r\n\r\nQuestion: What exactly is the medical code prediction problem?\r\n\r\nAnswer: Clinical notes contain much information about what precisely happened during the patient's entire stay. And those clinical notes (e.g., discharge summary) is typically long, loosely structured, consists of medical domain language, and sometimes riddled with spelling errors. So, it's a highly multi-label classification problem, and the forthcoming ICD-11 standard will add more complexity to the problem! The medical code prediction problem is to annotate this clinical note with multiple codes subset from nearly 70K total codes (in the current ICD-10 system, for example).","categories":[{"name":"Medical","url":"/area/medical"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":27,"papers_with_code":16,"benchmarks":7,"benchmark_tables_in_archive":7,"benchmark_tables_shown":7,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":7,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/medical-code-prediction-on-mimic-iii","slug":"medical-code-prediction-on-mimic-iii","dataset":"MIMIC-III","dataset_url":"/dataset/mimic-iii","rows_in_archive":18,"metrics":["Micro-F1","Macro-F1","Micro-AUC","Macro-AUC","Precision@5","Precision@8","Precision@15","mAP"],"first_row_in_archive_order":{"model":"GKI-ICD","paper_title":"A General Knowledge Injection Framework for ICD Coding","paper_url":"/paper/a-general-knowledge-injection-framework-for-1","paper_date":"2025-05-24","arxiv_id":"2505.18708","code_links":[{"title":"xuzhang0112/GKI-ICD","url":"https://github.com/xuzhang0112/GKI-ICD"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd-10","slug":"medical-code-prediction-on-mimic-iv-icd-10","dataset":"MIMIC-IV ICD-10","dataset_url":"/dataset/mimic-iv-icd-10","rows_in_archive":6,"metrics":["Precision@8","F1 Macro","F1 Micro","Precision@15","R-Prec","mAP","Exact Match Ratio","AUC Macro","AUC Micro"],"first_row_in_archive_order":{"model":"PLM-ICD","paper_title":"Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study","paper_url":"/paper/automated-medical-coding-on-mimic-iii-and","paper_date":"2023-04-21","arxiv_id":"2304.10909","code_links":[{"title":"joakimedin/medical-coding-reproducibility","url":"https://github.com/joakimedin/medical-coding-reproducibility"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd-9","slug":"medical-code-prediction-on-mimic-iv-icd-9","dataset":"MIMIC-IV ICD-9","dataset_url":"/dataset/mimic-iv-icd-9","rows_in_archive":6,"metrics":["AUC Macro","AUC Micro","Exact Match Ratio","F1 Macro","F1 Micro","Precision@15","Precision@8","R-Prec","mAP"],"first_row_in_archive_order":{"model":"PLM-ICD","paper_title":"Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study","paper_url":"/paper/automated-medical-coding-on-mimic-iii-and","paper_date":"2023-04-21","arxiv_id":"2304.10909","code_links":[{"title":"joakimedin/medical-coding-reproducibility","url":"https://github.com/joakimedin/medical-coding-reproducibility"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd-10-1","slug":"medical-code-prediction-on-mimic-iv-icd-10-1","dataset":"MIMIC-IV-ICD-10-full","dataset_url":"/dataset/mimic-iv-icd-10-full","rows_in_archive":5,"metrics":["Macro-AUC","Micro-AUC","Macro-F1","Micro-F1","Precision@8"],"first_row_in_archive_order":{"model":"MSMN","paper_title":"Mimic-IV-ICD: A new benchmark for eXtreme MultiLabel Classification","paper_url":"/paper/mimic-iv-icd-a-new-benchmark-for-extreme-1","paper_date":"2023-04-27","arxiv_id":"2304.13998","code_links":[{"title":"thomasnguyen92/MIMIC-IV-ICD-data-processing","url":"https://github.com/thomasnguyen92/MIMIC-IV-ICD-data-processing"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd10","slug":"medical-code-prediction-on-mimic-iv-icd10","dataset":"MIMIC-IV-ICD10-top50","dataset_url":"/dataset/mimic-iv-icd10-top50","rows_in_archive":5,"metrics":["F1 (micro)","F1 (macro)","AUC (Micro)","AUC (Macro)","Precision@5"],"first_row_in_archive_order":{"model":"MSMN","paper_title":"Mimic-IV-ICD: A new benchmark for eXtreme MultiLabel Classification","paper_url":"/paper/mimic-iv-icd-a-new-benchmark-for-extreme-1","paper_date":"2023-04-27","arxiv_id":"2304.13998","code_links":[{"title":"thomasnguyen92/MIMIC-IV-ICD-data-processing","url":"https://github.com/thomasnguyen92/MIMIC-IV-ICD-data-processing"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd9","slug":"medical-code-prediction-on-mimic-iv-icd9","dataset":"MIMIC-IV-ICD9-top50","dataset_url":"/dataset/mimic-iv-icd9-top50","rows_in_archive":5,"metrics":["AUC Macro","AUC Micro","F1 Macro","F1 Micro","Precision @5"],"first_row_in_archive_order":{"model":"MSMN","paper_title":"Mimic-IV-ICD: A new benchmark for eXtreme MultiLabel Classification","paper_url":"/paper/mimic-iv-icd-a-new-benchmark-for-extreme-1","paper_date":"2023-04-27","arxiv_id":"2304.13998","code_links":[{"title":"thomasnguyen92/MIMIC-IV-ICD-data-processing","url":"https://github.com/thomasnguyen92/MIMIC-IV-ICD-data-processing"}],"syntology":null}},{"leaderboard":"/sota/medical-code-prediction-on-mimic-iv-icd9-full","slug":"medical-code-prediction-on-mimic-iv-icd9-full","dataset":"MIMIC-IV-ICD9-full","dataset_url":"/dataset/mimic-iv-icd9-full","rows_in_archive":5,"metrics":["Macro AUC","Micro AUC","F1 Macro","F1 Micro","Precision@8"],"first_row_in_archive_order":{"model":"MSMN","paper_title":"Mimic-IV-ICD: A new benchmark for eXtreme MultiLabel Classification","paper_url":"/paper/mimic-iv-icd-a-new-benchmark-for-extreme-1","paper_date":"2023-04-27","arxiv_id":"2304.13998","code_links":[{"title":"thomasnguyen92/MIMIC-IV-ICD-data-processing","url":"https://github.com/thomasnguyen92/MIMIC-IV-ICD-data-processing"}],"syntology":null}}],"datasets":[{"url":"/dataset/mimic-iii","name":"MIMIC-III","full_name":"The Medical Information Mart for Intensive Care III","num_papers_in_archive":1041},{"url":"/dataset/mimic-iv-icd-10","name":"MIMIC-IV ICD-10","full_name":"","num_papers_in_archive":3},{"url":"/dataset/mimic-iv-icd-10-full","name":"MIMIC-IV-ICD-10-full","full_name":"","num_papers_in_archive":3},{"url":"/dataset/mimic-iv-icd9-full","name":"MIMIC-IV-ICD9-full","full_name":"","num_papers_in_archive":3},{"url":"/dataset/mimic-iv-icd-9","name":"MIMIC-IV ICD-9","full_name":"","num_papers_in_archive":1},{"url":"/dataset/mimic-iv-icd10-top50","name":"MIMIC-IV-ICD10-top50","full_name":"","num_papers_in_archive":1},{"url":"/dataset/mimic-iv-icd9-top50","name":"MIMIC-IV-ICD9-top50","full_name":"MIMIC-IV-ICD9-top50","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/multi-label-classification","name":"Multi-Label Classification"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":16,"of":16,"tagged_in_all":27,"items":[{"url":"/paper/icd-coding-from-clinical-text-using-multi","title":"ICD Coding from Clinical Text Using Multi-Filter Residual Convolutional Neural Network","date":"2019-11-25","arxiv_id":"1912.00862","repositories_listed":3,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/explainable-prediction-of-medical-codes-from","title":"Explainable Prediction of Medical Codes from Clinical Text","date":"2018-02-15","arxiv_id":"1802.05695","repositories_listed":3,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/an-unsupervised-approach-to-achieve","title":"An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare Records","date":"2024-06-13","arxiv_id":"2406.08958","repositories_listed":2,"syntology":null},{"url":"/paper/multi-task-balanced-and-recalibrated-network","title":"Multitask Balanced and Recalibrated Network for Medical Code Prediction","date":"2021-09-06","arxiv_id":"2109.02418","repositories_listed":2,"syntology":null},{"url":"/paper/an-explainable-cnn-approach-for-medical-codes","title":"An Explainable CNN Approach for Medical Codes Prediction from Clinical Text","date":"2021-01-14","arxiv_id":"2101.11430","repositories_listed":2,"syntology":null},{"url":"/paper/explainable-automated-coding-of-clinical","title":"Explainable Automated Coding of Clinical Notes using Hierarchical Label-wise Attention Networks and Label Embedding Initialisation","date":"2020-10-29","arxiv_id":"2010.15728","repositories_listed":2,"syntology":null},{"url":"/paper/a-label-attention-model-for-icd-coding-from","title":"A Label Attention Model for ICD Coding from Clinical Text","date":"2020-07-13","arxiv_id":"2007.06351","repositories_listed":2,"syntology":{"n":5,"n_ran":2,"n_unverified":3,"n_pointer_only":1}},{"url":"/paper/mimic-iii-a-freely-accessible-critical-care","title":"MIMIC-III, a freely accessible critical care database","date":"2016-05-24","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/automated-medical-coding-on-mimic-iii-and","title":"Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study","date":"2023-04-21","arxiv_id":"2304.10909","repositories_listed":1,"syntology":null},{"url":"/paper/knowledge-injected-prompt-based-fine-tuning","title":"Knowledge Injected Prompt Based Fine-tuning for Multi-label Few-shot ICD Coding","date":"2022-10-07","arxiv_id":"2210.03304","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/automatic-icd-coding-exploiting-discourse","title":"Automatic ICD Coding Exploiting Discourse Structure and Reconciled Code Embeddings","date":"2022-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/hicu-leveraging-hierarchy-for-curriculum","title":"HiCu: Leveraging Hierarchy for Curriculum Learning in Automated ICD Coding","date":"2022-08-03","arxiv_id":"2208.02301","repositories_listed":1,"syntology":null},{"url":"/paper/code-synonyms-do-matter-multiple-synonyms-1","title":"Code Synonyms Do Matter: Multiple Synonyms Matching Network for Automatic ICD Coding","date":"2022-03-03","arxiv_id":"2203.01515","repositories_listed":1,"syntology":null},{"url":"/paper/modeling-diagnostic-label-correlation-for-1","title":"Modeling Diagnostic Label Correlation for Automatic ICD Coding","date":"2021-06-24","arxiv_id":"2106.12800","repositories_listed":1,"syntology":null},{"url":"/paper/medical-code-prediction-from-discharge","title":"Medical Code Prediction from Discharge Summary: Document to Sequence BERT using Sequence Attention","date":"2021-06-15","arxiv_id":"2106.07932","repositories_listed":1,"syntology":null},{"url":"/paper/multitask-recalibrated-aggregation-network","title":"Multitask Recalibrated Aggregation Network for Medical Code Prediction","date":"2021-04-02","arxiv_id":"2104.00952","repositories_listed":1,"syntology":null}],"syntology_records":4,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}