Datasets › SICK
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics. It includes a large number of sentence pairs that are rich in the lexical, syntactic and semantic phenomena. Each pair of sentences is annotated in two dimensions: relatedness and entailment. The relatedness score ranges from 1 to 5, and Pearson’s r is used for evaluation; the entailment relation is categorical, consisting of entailment, contradiction, and neutral. There are 4439 pairs in the train split, 495 in the trial split used for development and 4906 in the test split. The sentence pairs are generated from image and video caption datasets before being paired up using some algorithm.
Source: Multi-Label Transfer Learning for Multi-Relational Semantic Similarity Image Source: https://www.researchgate.net/figure/Example-of-SICK-dataset-sentence-expansion-process-14_fig1_344863619
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Semantic Textual Similarity | SICK | PromCSE-RoBERTa-large (0.355B) Spearman Correlation 0.8243 | Improved Universal Sentence Embeddings with Prompt-based... | yjiangcm/promcse | 22 | Compare |
| Tabular Data Generation | SICK | GReaT DT Accuracy 97.72 | Language Models are Realistic Tabular Data Generators | kathrinse/be_great | 6 | Compare |
| Semantic Similarity | SICK | Dependency Tree-LSTM (Tai et al., 2015) MSE 0.2532 | Improved Semantic Representations From Tree-Structured... | dmlc/dgl +15 | 5 | Compare |
| Semantic Textual Similarity | SICK-R | AnglE-LLaMA-7B Spearman Correlation 0.8094 | AnglE-optimized Text Embeddings | SeanLee97/AnglE +1 | 2 | Compare |
| Natural Language Inference | SICK | NeuralLog 1:1 Accuracy 0.903 | NeuralLog: Natural Language Inference with Joint Neural... | eric11eca/NeuralLog | 1 | Compare |
Papers archive 2025-07-28
18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 348. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
2 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- SICK
- SICK-R
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections