{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/relevant-word-order-vectorization-for","title":"Relevant Word Order Vectorization for Improved Natural Language Processing in Electronic Healthcare Records","arxiv_id":"1812.02627","date":"2018-12-06","proceeding":null,"authors":["Jeffrey Thompson","Jinxiang Hu","Dinesh Pal Mudaranthakam","David Streeter","Lisa Neums","Michele Park","Devin C. Koestler","Byron Gajewski","Matthew S. Mayo"],"abstract":"Objective: Electronic health records (EHR) represent a rich resource for\nconducting observational studies, supporting clinical trials, and more.\nHowever, much of the relevant information is stored in an unstructured format\nthat makes it difficult to use. Natural language processing approaches that\nattempt to automatically classify the data depend on vectorization algorithms\nthat impose structure on the text, but these algorithms were not designed for\nthe unique characteristics of EHR. Here, we propose a new algorithm for\nstructuring so-called free-text that may help researchers make better use of\nEHR. We call this method Relevant Word Order Vectorization (RWOV).\n  Materials and Methods: As a proof-of-concept, we attempted to classify the\nhormone receptor status of breast cancer patients treated at the University of\nKansas Medical Center during a recent year, from the unstructured text of\npathology reports. Our approach attempts to account for the semi-structured way\nthat healthcare providers often enter information. We compared this approach to\nthe ngrams and word2vec methods.\n  Results: Our approach resulted in the most consistently high accuracy, as\nmeasured by F1 score and area under the receiver operating characteristic curve\n(AUC).\n  Discussion: Our results suggest that methods of structuring free text that\ntake into account its context may show better performance, and that our\napproach is promising.\n  Conclusion: By using a method that accounts for the fact that healthcare\nproviders tend to use certain key words repetitively and that the order of\nthese key words is important, we showed improved performance over methods that\ndo not.","url_abs":"http://arxiv.org/abs/1812.02627v1","url_pdf":"http://arxiv.org/pdf/1812.02627v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"relevant-word-order-vectorization-for","repo_url":"https://github.com/jeffreyat/RWOV","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}