{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/machine-learned-solutions-for-three-stages-of","title":"Machine-learned solutions for three stages of clinical information extraction: the state of the art at i2b2 2010","arxiv_id":null,"date":"2011-05-12","proceeding":"JAMIA 2011 5","authors":["Berry de Bruijn","Colin Cherry","Svetlana Kiritchenko","Joel Martin","Xiaodan Zhu"],"abstract":"Objective: As clinical text mining continues to mature, its\r\npotential as an enabling technology for innovations in\r\npatient care and clinical research is becoming a reality. A\r\ncritical part of that process is rigid benchmark testing of\r\nnatural language processing methods on realistic clinical\r\nnarrative. In this paper, the authors describe the design\r\nand performance of three state-of-the-art text-mining\r\napplications from the National Research Council of\r\nCanada on evaluations within the 2010 i2b2 challenge.\r\nDesign: The three systems perform three key steps in\r\nclinical information extraction: (1) extraction of medical\r\nproblems, tests, and treatments, from discharge\r\nsummaries and progress notes; (2) classification of\r\nassertions made on the medical problems; (3)\r\nclassification of relations between medical concepts.\r\nMachine learning systems performed these tasks using\r\nlarge-dimensional bags of features, as derived from both\r\nthe text itself and from external sources: UMLS, cTAKES,\r\nand Medline.\r\nMeasurements: Performance was measured per\r\nsubtask, using micro-averaged F-scores, as calculated by\r\ncomparing system annotations with ground-truth\r\nannotations on a test set.\r\nResults: The systems ranked high among all submitted\r\nsystems in the competition, with the following F-scores:\r\nconcept extraction 0.8523 (ranked first); assertion\r\ndetection 0.9362 (ranked first); relationship detection\r\n0.7313 (ranked second).\r\nConclusion: For all tasks, we found that the introduction\r\nof a wide range of features was crucial to success.\r\nImportantly, our choice of machine learning algorithms\r\nallowed us to be versatile in our feature design, and to\r\nintroduce a large number of features without overfitting\r\nand without encountering computing-resource\r\nbottlenecks.","url_abs":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3168309/","url_pdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3168309/pdf/amiajnl-2011-000150.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"clinical-concept-extraction","task_name":"Clinical Concept Extraction"},{"task_slug":"relationship-detection","task_name":"Relationship Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/clinical-concept-extraction-on-2010-i2b2va","task":"Clinical Concept Extraction","dataset":"2010 i2b2/VA","model":"deBruijn et al. (System 1.1)","rank_in_archive_order":5,"of":5,"metrics":{"Exact Span F1":"85.23"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}