Papers › Sub-word information in pre-trained biomedical word representations: evaluation and...

Sub-word information in pre-trained biomedical word representations: evaluation and hyper-parameter optimization

1 Jul 2018WS 2018 7archive 2025-07-28

Dieter Galea, Ivan Laponogov, Kirill Veselkov

Word2vec embeddings are limited to computing vectors for in-vocabulary terms and do not take into account sub-word information. Character-based representations, such as fastText, mitigate such limitations. We optimize and compare these representations for the biomedical domain. fastText was found to consistently outperform word2vec in named entity recognition tasks for entities such as chemicals and genes. This is likely due to gained information from computed out-of-vocabulary term vectors, as well as the word compositionality of such entities. Contrastingly, performance varied on intrinsic datasets. Optimal hyper-parameters were intrinsic dataset-dependent, likely due to differences in term types distributions. This indicates embeddings should be chosen based on the task at hand. We therefore provide a number of optimized hyper-parameter sets and pre-trained word2vec and fastText models, available on \url{https://github.com/dterg/bionlp-embed}.

PaperPDFCode

Code

dterg/bionlp-embed officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Feature EngineeringNamed Entity RecognitionNamed Entity Recognition (NER)Word Embeddingsnamed-entity-recognition

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

fastText

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections