Datasets › Douban
Douban (Douban Conversation Corpus)
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot. The statistics of Douban Conversation Corpus are shown in the following table.
| Train | Val | Test | |
|---|---|---|---|
| session-response pairs | 1m | 50k | 10k |
| Avg. positive response per session | 1 | 1 | 1.18 |
| Fless Kappa | N\A | N\A | 0.41 |
| Min turn per session | 3 | 3 | 3 |
| Max ture per session | 98 | 91 | 45 |
| Average turn per session | 6.69 | 6.75 | 5.95 |
| Average Word per utterance | 18.56 | 18.50 | 20.74 |
The test data contains 1000 dialogue context, and for each context we create 10 responses as candidates. We recruited three labelers to judge if a candidate is a proper response to the session. A proper response means the response can naturally reply to the message given the context. Each pair received three labels and the majority of the labels was taken as the final decision.
As far as we known, this is the first human-labeled test set for retrieval-based chatbots. The entire corpus link https://www.dropbox.com/s/90t0qtji9ow20ca/DoubanConversaionCorpus.zip?dl=0
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Conversational Response Selection | Douban | SEMSOL(W/o utterances) MAP 0.651 | Knowledge-aware response selection with semantics... | losmes/SemSol | 16 | Compare |
| Recommendation Systems | Douban Monti | GLocal-K RMSE 0.7208 | GLocal-K: Global and Local Kernels for Recommender Systems | usydnlp/Glocal_K +2 | 8 | Compare |
| Recommendation Systems | Douban | I-CFN RMSE 0.6911 | Hybrid Recommender System based on Autoencoders | fstrub95/Autoencoders_cf +3 | 7 | Compare |
| Link Prediction | Douban | HSRL (DW) AUC 84.2 | Learning Topological Representation for Networks via... | fuguoji/HSRL | 2 | Compare |
Papers archive 2025-07-28
29 shown of 29 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 81. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Douban
- Douban Monti
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections