Datasets › Quora Question Pairs

Quora Question Pairs

archive 2025-07-28

Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.

Source: Bilateral Multi-Perspective Matching for Natural Language Sentences

Benchmarks archive 2025-07-28

All 8 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 45 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 55. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
BM25S: Orders of magnitude faster lexical search via eager sparse scoring 3 5 4 Jul 2024 ran 10 of 22 samples (12 unverified)
Memory-efficient Stochastic methods for Memory-based Transformers 1 2 14 Nov 2023 not harvested
SplitEE: Early Exit in Deep Neural Networks with Split Computing 1 1 17 Sep 2023 not harvested
Adversarial Self-Attention for Language Understanding 1 2 25 Jun 2022 not harvested
Hierarchical Sketch Induction for Paraphrase Generation 1 1 7 Mar 2022 ran 3 of 3 samples (0 unverified)
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language 12 1 7 Feb 2022 ran 0 of 6 samples (6 unverified)
Charformer: Fast Character Transformers via Gradient-based Subword Tokenization 2 1 23 Jun 2021 ran 7 of 10 samples (3 unverified)
Factorising Meaning and Form for Intent-Preserving Paraphrasing 1 1 31 May 2021 not harvested
FNet: Mixing Tokens with Fourier Transforms 12 1 9 May 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Entailment as Few-Shot Learner 3 1 29 Apr 2021 ran 1 of 3 samples (2 unverified)
How to Train BERT with an Academic Budget 4 1 15 Apr 2021 not harvested
CLEAR: Contrastive Learning for Sentence Representation 0 1 31 Dec 2020 not harvested
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning 2 1 22 Dec 2020 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
RealFormer: Transformer Likes Residual Attention 5 1 21 Dec 2020 not harvested
Self-Explaining Structures Improve NLP Models 1 1 3 Dec 2020 not harvested
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
SqueezeBERT: What can computer vision teach NLP about efficient neural networks? 6 1 19 Jun 2020 ran 0 of 1 samples (1 unverified)
DeBERTa: Decoding-enhanced BERT with Disentangled Attention 14 1 5 Jun 2020 ran 4 of 13 samples (9 unverified; 3 pointer-only for licence)
What Do Questions Exactly Ask? MFAE: Duplicate Question Identification with Multi-Fusion Asking Emphasis 1 2 7 May 2020 not harvested
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators 19 1 23 Mar 2020 ran 26 of 40 samples (14 unverified; 10 pointer-only for licence)
TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding 0 1 16 Mar 2020 not harvested
Multi-task Sentence Encoding Model for Semantic Retrieval in Question Answering Systems 0 1 18 Nov 2019 not harvested
SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization 6 3 8 Nov 2019 ran 6 of 8 samples (2 unverified; 1 pointer-only for licence)
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer 57 5 23 Oct 2019 ran 2 of 31 samples (29 unverified)
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 37 1 2 Oct 2019 ran 19 of 27 samples (8 unverified)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
TinyBERT: Distilling BERT for Natural Language Understanding 10 1 23 Sep 2019 ran 0 of 4 samples (4 unverified; 4 pointer-only for licence)
StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding 0 1 13 Aug 2019 not harvested
Simple and Effective Text Matching with Richer Alignment Features 3 2 1 Aug 2019 ran 1 of 1 samples (0 unverified)
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding 3 2 29 Jul 2019 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)

The full list of 45 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (non-commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Quora Question Pairs
  • Quora Question Pairs Dev
  • qqp

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections