{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/create-cohort-retrieval-enhanced-by-analysis","title":"CREATE: Cohort Retrieval Enhanced by Analysis of Text from Electronic Health Records using OMOP Common Data Model","arxiv_id":"1901.07601","date":"2019-01-22","proceeding":null,"authors":["Liu Sijia","Wang Yanshan","Wen Andrew","Wang Liwei","Hong Na","Shen Feichen","Bedrick Steven","Hersh William","Liu Hongfang"],"abstract":"Background: Widespread adoption of electronic health records (EHRs) has\nenabled secondary use of EHR data for clinical research and healthcare\ndelivery. Natural language processing (NLP) techniques have shown promise in\ntheir capability to extract the embedded information in unstructured clinical\ndata, and information retrieval (IR) techniques provide flexible and scalable\nsolutions that can augment the NLP systems for retrieving and ranking relevant\nrecords. Methods: In this paper, we present the implementation of Cohort\nRetrieval Enhanced by Analysis of Text from EHRs (CREATE), a cohort retrieval\nsystem that can execute textual cohort selection queries on both structured and\nunstructured EHR data. CREATE is a proof-of-concept system that leverages a\ncombination of structured queries and IR techniques on NLP results to improve\ncohort retrieval performance while adopting the Observational Medical Outcomes\nPartnership (OMOP) Common Data Model (CDM) to enhance model portability. The\nNLP component empowered by cTAKES is used to extract CDM concepts from textual\nqueries. We design a hierarchical index in Elasticsearch to support CDM concept\nsearch utilizing IR techniques and frameworks. Results: Our case study on 5\ncohort identification queries evaluated using the IR metric, P@5 (Precision at\n5) at both the patient-level and document-level, demonstrates that CREATE\nachieves an average P@5 of 0.90, which outperforms systems using only\nstructured data or only unstructured data with average P@5s of 0.54 and 0.74,\nrespectively.","url_abs":"http://arxiv.org/abs/1901.07601v1","url_pdf":"http://arxiv.org/pdf/1901.07601v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"create-cohort-retrieval-enhanced-by-analysis","repo_url":"https://github.com/OHNLPIR/OMOP_CDM_IO","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}