Datasets › SICK

SICK (Sentences Involving Compositional Knowledge)

Introduced by Marco Marelli et al. in A SICK cure for the evaluation of compositional distributional semantic models1 Jan 2014 archive 2025-07-28

The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics. It includes a large number of sentence pairs that are rich in the lexical, syntactic and semantic phenomena. Each pair of sentences is annotated in two dimensions: relatedness and entailment. The relatedness score ranges from 1 to 5, and Pearson’s r is used for evaluation; the entailment relation is categorical, consisting of entailment, contradiction, and neutral. There are 4439 pairs in the train split, 495 in the trial split used for development and 4906 in the test split. The sentence pairs are generated from image and video caption datasets before being paired up using some algorithm.

Source: Multi-Label Transfer Learning for Multi-Relational Semantic Similarity Image Source: https://www.researchgate.net/figure/Example-of-SICK-dataset-sentence-expansion-process-14_fig1_344863619

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Semantic Textual Similarity SICK PromCSE-RoBERTa-large (0.355B) Spearman Correlation 0.8243 Improved Universal Sentence Embeddings with Prompt-based... yjiangcm/promcse 22 Compare
Tabular Data Generation SICK GReaT DT Accuracy 97.72 Language Models are Realistic Tabular Data Generators kathrinse/be_great 6 Compare
Semantic Similarity SICK Dependency Tree-LSTM (Tai et al., 2015) MSE 0.2532 Improved Semantic Representations From Tree-Structured... dmlc/dgl +15 5 Compare
Semantic Textual Similarity SICK-R AnglE-LLaMA-7B Spearman Correlation 0.8094 AnglE-optimized Text Embeddings SeanLee97/AnglE +1 2 Compare
Natural Language Inference SICK NeuralLog 1:1 Accuracy 0.903 NeuralLog: Natural Language Inference with Joint Neural... eric11eca/NeuralLog 1 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 348. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Tabular Data Generation using Binary Diffusion 1 1 20 Sep 2024 ran 10 of 10 samples (0 unverified)
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity 1 1 2 Apr 2024 not harvested
AnglE-optimized Text Embeddings 2 2 22 Sep 2023 not harvested
Scaling Sentence Embeddings with Large Language Models 1 3 31 Jul 2023 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Language Models are Realistic Tabular Data Generators 1 2 12 Oct 2022 ran 0 of 6 samples (6 unverified)
Improved Universal Sentence Embeddings with Prompt-based Contrastive Learning and Energy-based Learning 1 1 14 Mar 2022 not harvested
Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations 1 5 27 Sep 2021 not harvested
NeuralLog: Natural Language Inference with Joint Neural and Logical Reasoning 1 1 29 May 2021 ran 0 of 6 samples (6 unverified)
SimCSE: Simple Contrastive Learning of Sentence Embeddings 23 1 18 Apr 2021 ran 17 of 30 samples (13 unverified; 19 pointer-only for licence)
Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders 1 2 16 Apr 2021 ran 0 of 1 samples (1 unverified)
Generating Datasets with Pretrained Language Models 2 2 15 Apr 2021 ran 0 of 3 samples (3 unverified)
On the Sentence Embeddings from Pre-trained Language Models 3 1 2 Nov 2020 ran 3 of 3 samples (0 unverified)
An Unsupervised Sentence Embedding Method by Mutual Information Maximization 1 1 25 Sep 2020 not harvested
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks 64 5 27 Aug 2019 ran 20 of 58 samples (38 unverified; 11 pointer-only for licence)
Modeling Tabular data using Conditional GAN 9 3 1 Jul 2019 ran 2 of 3 samples (1 unverified)
Efficient Vector Representation for Documents through Corruption 1 1 8 Jul 2017 not harvested
Skip-Thought Vectors 16 1 22 Jun 2015 ran 0 of 3 samples (3 unverified)
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks 16 3 28 Feb 2015 ran 6 of 15 samples (9 unverified; 6 pointer-only for licence)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC-SA 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SICK
  • SICK-R

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections