{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/awe-cm-vectors-augmenting-word-embeddings","title":"AWE-CM Vectors: Augmenting Word Embeddings with a Clinical Metathesaurus","arxiv_id":"1712.01460","date":"2017-12-05","proceeding":null,"authors":["Willie Boag","Hassan Kané"],"abstract":"In recent years, word embeddings have been surprisingly effective at\ncapturing intuitive characteristics of the words they represent. These vectors\nachieve the best results when training corpora are extremely large, sometimes\nbillions of words. Clinical natural language processing datasets, however, tend\nto be much smaller. Even the largest publicly-available dataset of medical\nnotes is three orders of magnitude smaller than the dataset of the oft-used\n\"Google News\" word vectors. In order to make up for limited training data\nsizes, we encode expert domain knowledge into our embeddings. Building on a\nprevious extension of word2vec, we show that generalizing the notion of a\nword's \"context\" to include arbitrary features creates an avenue for encoding\ndomain knowledge into word embeddings. We show that the word vectors produced\nby this method outperform their text-only counterparts across the board in\ncorrelation with clinical experts.","url_abs":"http://arxiv.org/abs/1712.01460v1","url_pdf":"http://arxiv.org/pdf/1712.01460v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"awe-cm-vectors-augmenting-word-embeddings","repo_url":"https://github.com/wboag/awecm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}