Datasets › RRS Ranking Test
RRS Ranking Test (Restoration-200k for Response Selection with Ranking Test Set)
| Train | Validation | Test | Ranking Test | |
|---|---|---|---|---|
| size | 0.4M | 50K | 5K | 800 |
| pos:neg | 1:1 | 1:9 | 1.2:8.8 | - |
| avg turns | 5.0 | 5.0 | 5.0 | 5.0 |
Ranking test set contains the high-quality responses that selected by some baselines, and their correlation with the conversation context are carefully annotated by 8 professional annotators (the average annotation scores are saved for ranking). For ranking test set, the metrics should be NDCG@3 and NDCG@5, since the correlation scores are provided. More details are available in the Appendix of the paper.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Conversational Response Selection | RRS Ranking Test | Poly-encoder NDCG@3 0.679 | Poly-encoders: Transformer Architectures and... | sfzhou5678/PolyEncoder +6 | 4 | Compare |
Papers archive 2025-07-28
4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 5. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Fine-grained Post-training for Improving Retrieval-based Dialogue Systems | 1 | 1 | 24 May 2021 | not harvested |
| Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots | 2 | 1 | 7 Apr 2020 | not harvested |
| An Effective Domain Adaptive Post-Training Method for BERT in Response Selection | 1 | 1 | 13 Aug 2019 | not harvested |
| Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring | 7 | 1 | 22 Apr 2019 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- RRS Ranking Test
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections