Datasets › BIOSSES
BIOSSES (Biomedical Semantic Similarity Estimation System)
The BIOSSES data set comprises total 100 sentence pairs all of which were selected from the "TAC2 Biomedical Summarization Track Training Data Set" .
The sentence pairs were evaluated by five different human experts that judged their similarity and gave scores in a range [0-4]. Our guideline was prepared based on SemEval 2012 Task 6 Guideline.
Image source: BIOSSES
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Sentence Embeddings For Biomedical Texts | BIOSSES | Supervised combination of: Jaccard, Q-gram, sent2vec, Paragraph vector DM, skip-thoughts, fastText Pearson Correlation 0.871 | Neural sentence embedding models for semantic similarity... | kathrinblagec/neural-sentence-embedding-models-for-biomedical-applications | 14 | Compare |
| Semantic Similarity | BIOSSES | BioLinkBERT (large) Pearson Correlation 0.9363 | LinkBERT: Pretraining Language Models with Document Links | michiyasunaga/LinkBERT | 3 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 38. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| LinkBERT: Pretraining Language Models with Document Links | 1 | 2 | 29 Mar 2022 | ran 0 of 14 samples (14 unverified) |
| Neural sentence embedding models for semantic similarity estimation in the biomedical domain | 1 | 9 | 1 Oct 2021 | not harvested |
| Transfer Learning in Biomedical Natural Language Processing: An Evaluation of BERT and ELMo on Ten Benchmarking Datasets | 4 | 1 | 13 Jun 2019 | ran 0 of 2 samples (2 unverified) |
| BioSentVec: creating sentence embeddings for biomedical texts | 4 | 4 | 22 Oct 2018 | not harvested |
| BIOSSES: A Semantic Sentence Similarity Estimation System for the Biomedical Domain | 0 | 1 | 15 Jul 2017 | not harvested |
Dataset loaders archive 2025-07-28
3 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- BIOSSES
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections