{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/probing-biomedical-embeddings-from-language","title":"Probing Biomedical Embeddings from Language Models","arxiv_id":"1904.02181","date":"2019-04-03","proceeding":"WS 2019 6","authors":["Qiao Jin","Bhuwan Dhingra","William W. Cohen","Xinghua Lu"],"abstract":"Contextualized word embeddings derived from pre-trained language models (LMs)\nshow significant improvements on downstream NLP tasks. Pre-training on\ndomain-specific corpora, such as biomedical articles, further improves their\nperformance. In this paper, we conduct probing experiments to determine what\nadditional information is carried intrinsically by the in-domain trained\ncontextualized embeddings. For this we use the pre-trained LMs as fixed feature\nextractors and restrict the downstream task models to not have additional\nsequence modeling layers. We compare BERT, ELMo, BioBERT and BioELMo, a\nbiomedical version of ELMo trained on 10M PubMed abstracts. Surprisingly, while\nfine-tuned BioBERT is better than BioELMo in biomedical NER and NLI tasks, as a\nfixed feature extractor BioELMo outperforms BioBERT in our probing tasks. We\nuse visualization and nearest neighbor analysis to show that better encoding of\nentity-type and relational information leads to this superiority.","url_abs":"http://arxiv.org/abs/1904.02181v1","url_pdf":"http://arxiv.org/pdf/1904.02181v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"probing-biomedical-embeddings-from-language","repo_url":"https://github.com/Andy-jqa/bioelmo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"cg","task_name":"NER"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bert","method_name":"BERT"},{"method_slug":"bilstm","method_name":"BiLSTM"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"elmo","method_name":"ELMo"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-linear-decay","method_name":"Linear Warmup With Linear Decay"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"},{"method_slug":"weight-decay","method_name":"Weight Decay"},{"method_slug":"wordpiece","method_name":"WordPiece"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1904.02181","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}