{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/using-natural-language-processing-techniques","title":"Using natural language processing techniques to extract information on the properties and functionalities of energetic materials from large text corpora","arxiv_id":"1903.00415","date":"2019-03-01","proceeding":null,"authors":["Daniel C. Elton","Dhruv Turakhia","Nischal Reddy","Zois Boukouvalas","Mark D. Fuge","Ruth M. Doherty","Peter W. Chung"],"abstract":"The number of scientific journal articles and reports being published about\nenergetic materials every year is growing exponentially, and therefore\nextracting relevant information and actionable insights from the latest\nresearch is becoming a considerable challenge. In this work we explore how\ntechniques from natural language processing and machine learning can be used to\nautomatically extract chemical insights from large collections of documents. We\nfirst describe how to download and process documents from a variety of sources\n- journal articles, conference proceedings (including NTREM), the US Patent &\nTrademark Office, and the Defense Technical Information Center archive on\narchive.org. We present a custom NLP pipeline which uses open source NLP tools\nto identify the names of chemical compounds and relates them to function words\n(\"underwater\", \"rocket\", \"pyrotechnic\") and property words (\"elastomer\",\n\"non-toxic\"). After explaining how word embeddings work we compare the utility\nof two popular word embeddings - word2vec and GloVe. Chemical-chemical and\nchemical-application relationships are obtained by doing computations with word\nvectors. We show that word embeddings capture latent information about\nenergetic materials, so that related materials appear close together in the\nword embedding space.","url_abs":"http://arxiv.org/abs/1903.00415v1","url_pdf":"http://arxiv.org/pdf/1903.00415v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"using-natural-language-processing-techniques","repo_url":"https://github.com/pprzetacznik/patent-parsing-tools","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[{"method_slug":"glove","method_name":"GloVe"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}