Datasets › STS Benchmark

STS Benchmark (Semantic Textual Similarity)

archive 2025-07-28

STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums.

Source: STS Benchmark

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Semantic Textual Similarity STS Benchmark MT-DNN-SMART Pearson Correlation 0.929 SMART: Robust and Efficient Fine-Tuning for Pre-trained... namisan/mt-dnn +5 66 Compare
Semantic Textual Similarity STS13 AnglE-LLaMA-7B Spearman Correlation 0.9058 AnglE-optimized Text Embeddings SeanLee97/AnglE +1 22 Compare
Semantic Textual Similarity STS14 AnglE-LLaMA-13B Spearman Correlation 0.8689 AnglE-optimized Text Embeddings SeanLee97/AnglE +1 21 Compare
Semantic Textual Similarity STS12 PromptEOL+CSE+OPT-13B Spearman Correlation 0.8020 Scaling Sentence Embeddings with Large Language Models kongds/scaling_sentemb 20 Compare
Semantic Textual Similarity STS15 PromptEOL+CSE+LLaMA-30B Spearman Correlation 0.9004 Scaling Sentence Embeddings with Large Language Models kongds/scaling_sentemb 20 Compare
Semantic Textual Similarity STS16 AnglE-LLaMA-7B-v2 Spearman Correlation 0.8700 AnglE-optimized Text Embeddings SeanLee97/AnglE +1 20 Compare

Papers archive 2025-07-28

30 shown of 40 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 45. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity 1 1 2 Apr 2024 not harvested
Def2Vec: Extensible Word Embeddings from Dictionary Definitions 2 1 16 Dec 2023 not harvested
AnglE-optimized Text Embeddings 2 15 22 Sep 2023 not harvested
Scaling Sentence Embeddings with Large Language Models 1 18 31 Jul 2023 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale 4 1 15 Aug 2022 ran 2 of 5 samples (3 unverified)
Adversarial Self-Attention for Language Understanding 1 2 25 Jun 2022 not harvested
DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings 1 10 21 Apr 2022 ran 2 of 4 samples (2 unverified)
Improved Universal Sentence Embeddings with Prompt-based Contrastive Learning and Energy-based Learning 1 6 14 Mar 2022 not harvested
MNet-Sim: A Multi-layered Semantic Similarity Network to Evaluate Sentence Similarity 0 1 9 Nov 2021 not harvested
Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations 1 28 27 Sep 2021 not harvested
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2 1 23 Jun 2021 ran 7 of 10 samples (3 unverified)
FNet: Mixing Tokens with Fourier Transforms 12 1 9 May 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
SimCSE: Simple Contrastive Learning of Sentence Embeddings 23 9 18 Apr 2021 ran 17 of 30 samples (13 unverified; 19 pointer-only for licence)
Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders 1 12 16 Apr 2021 ran 0 of 1 samples (1 unverified)
How to Train BERT with an Academic Budget 4 1 15 Apr 2021 not harvested
Generating Datasets with Pretrained Language Models 2 7 15 Apr 2021 ran 0 of 3 samples (3 unverified)
CLEAR: Contrastive Learning for Sentence Representation 0 1 31 Dec 2020 not harvested
RealFormer: Transformer Likes Residual Attention 5 1 21 Dec 2020 not harvested
On the Sentence Embeddings from Pre-trained Language Models 3 6 2 Nov 2020 ran 3 of 3 samples (0 unverified)
A Statistical Framework for Low-bitwidth Training of Deep Neural Networks 2 1 27 Oct 2020 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
An Unsupervised Sentence Embedding Method by Mutual Information Maximization 1 6 25 Sep 2020 not harvested
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization 6 3 8 Nov 2019 ran 6 of 8 samples (2 unverified; 1 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 6 23 Oct 2019 ran 2 of 31 samples (29 unverified)
Q8BERT: Quantized 8Bit BERT 5 1 14 Oct 2019 ran 3 of 11 samples (8 unverified; 3 pointer-only for licence)
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 37 1 2 Oct 2019 ran 19 of 27 samples (8 unverified)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
TinyBERT: Distilling BERT for Natural Language Understanding 10 1 23 Sep 2019 ran 0 of 4 samples (4 unverified; 4 pointer-only for licence)

The full list of 40 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • STS Benchmark Dev
  • STS16
  • STS15
  • STS14
  • STS13
  • STS12
  • STS-B
  • STS Benchmark

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections