Papers › Neural sentence embedding models for semantic similarity estimation in the biomedical domain
Neural sentence embedding models for semantic similarity estimation in the biomedical domain
Kathrin Blagec, Hong Xu, Asan Agibetov, Matthias Samwald
BACKGROUND: In this study, we investigated the efficacy of current state-of-the-art neural sentence embedding models for semantic similarity estimation of sentences from biomedical literature. We trained different neural embedding models on 1.7 million articles from the PubMed Open Access dataset, and evaluated them based on a biomedical benchmark set containing 100 sentence pairs annotated by human experts and a smaller contradiction subset derived from the original benchmark set. RESULTS: With a Pearson correlation of 0.819, our best unsupervised model based on the Paragraph Vector Distributed Memory algorithm outperforms previous state-of-the-art results achieved on the BIOSSES biomedical benchmark set. Moreover, our proposed supervised model that combines different string-based similarity metrics with a neural embedding model surpasses previous ontology-dependent supervised state-of-the-art approaches in terms of Pearson's r (r=0.871) on the biomedical benchmark set. In contrast to the promising results for the original benchmark, we found our best models' performance on the smaller contradiction subset to be poor. CONCLUSIONS: In this study we highlighted the value of neural network-based models for semantic similarity estimation in the biomedical domain by showing that they can keep up with and even surpass previous state-of-the-art approaches for semantic similarity estimation that depend on the availability of laboriously curated ontologies when evaluated on a biomedical benchmark set. Capturing contradictions and negations in biomedical sentences, however, emerged as an essential area for further work.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Sentence Embeddings For Biomedical Texts | BIOSSES | Supervised combination of: Jaccard, Q-gram, sent2vec, Paragraph vector DM, skip-thoughts, fastText | Pearson Correlation | 0.871 | #1 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Unsupervised combination (mean) of: Jaccard, q-gram, Paragraph vector (PV-DBOW) and sent2vec | Pearson Correlation | 0.846 | #2 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Paragraph vector (PV-DM) | Pearson Correlation | 0.819 | #3 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Paragraph vector (PV-DBOW) | Pearson Correlation | 0.804 | #5 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Sent2vec | Pearson Correlation | 0.798 | #6 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | fastText (skip-gram, max pooling) | Pearson Correlation | 0.766 | #9 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Q-gram (q = 3) | Pearson Correlation | 0.723 | #10 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | Skip-thoughts | Pearson Correlation | 0.485 | #11 of 14 | Archive leaderboard | report |
| Sentence Embeddings For Biomedical Texts | BIOSSES | fastText (CBOW, max pooling) | Pearson Correlation | 0.253 | #14 of 14 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections