{"url":"/dataset/douban","name":"Douban","full_name":"Douban Conversation Corpus","description_markdown":"We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot. The statistics of Douban Conversation Corpus are shown in the following table. \r\n\r\n|      |Train|Val| Test         | \r\n| ------------- |:-------------:|:-------------:|:-------------:|\r\n| session-response pairs  | 1m|50k| 10k |\r\n| Avg. positive response per session     | 1|1| 1.18    | \r\n| Fless Kappa | N\\A|N\\A|0.41      | \r\n| Min turn per session | 3|3| 3      | \r\n| Max ture per session | 98|91|45    | \r\n| Average turn per session | 6.69|6.75|5.95    | \r\n| Average Word per utterance | 18.56|18.50|20.74   | \r\n\r\n\r\nThe test data contains 1000 dialogue context, and for each context we create 10 responses as candidates. We recruited three labelers to judge if a candidate is a proper response to the session. A proper response means the response can naturally reply to the message given the context. Each pair received three labels and the majority of the labels was taken as the final decision.\r\n\r\n<br>\r\nAs far as we known, this is the first human-labeled test set for retrieval-based chatbots. The entire corpus link https://www.dropbox.com/s/90t0qtji9ow20ca/DoubanConversaionCorpus.zip?dl=0","description_withheld":null,"homepage":"https://github.com/MarkWuNLP/MultiTurnResponseSelection","introduced_date":"2016-12-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/sequential-matching-network-a-new","title":"Sequential Matching Network: A New Architecture for Multi-turn Response Selection in Retrieval-based Chatbots","first_author":"Yu Wu","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Link Prediction","url":"/task/link-prediction","datasets_with_task":"/datasets/task/link-prediction"},{"name":"Recommendation Systems","url":"/task/recommendation-systems","datasets_with_task":"/datasets/task/recommendation-systems"},{"name":"Conversational Response Selection","url":"/task/conversational-response-selection","datasets_with_task":"/datasets/task/conversational-response-selection"}],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["Douban","Douban Monti"],"data_loaders":[],"num_papers_in_archive":81,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/conversational-response-selection-on-douban-1","task":"Conversational Response Selection","dataset_variant":"Douban","rows":16,"metrics":["MAP","MRR","P@1","R10@1","R10@2","R10@5"],"first_row_in_archive_order":{"model":"SEMSOL(W/o utterances)","paper":"/paper/knowledge-aware-response-selection-with","metrics":{"MAP":"0.651","MRR":"0.687","P@1":"0.510","R10@1":"0.328","R10@2":"0.552","R10@5":"0.877"},"code_links":[{"title":"losmes/SemSol","url":"https://github.com/losmes/SemSol"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/recommendation-systems-on-douban-monti","task":"Recommendation Systems","dataset_variant":"Douban Monti","rows":8,"metrics":["RMSE"],"first_row_in_archive_order":{"model":"GLocal-K","paper":"/paper/glocal-k-global-and-local-kernels-for","metrics":{"RMSE":"0.7208"},"code_links":[{"title":"usydnlp/Glocal_K","url":"https://github.com/usydnlp/Glocal_K"},{"title":"fleanend/TorchGlocalK","url":"https://github.com/fleanend/TorchGlocalK"},{"title":"fleanend/NeuralRecommender","url":"https://github.com/fleanend/NeuralRecommender"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/collaborative-filtering-on-douban","task":"Recommendation Systems","dataset_variant":"Douban","rows":7,"metrics":["RMSE","NDCG","Recall@20","AUC","HR@10","HR@100","PSP@10","nDCG@10","nDCG@100"],"first_row_in_archive_order":{"model":"I-CFN","paper":"/paper/hybrid-recommender-system-based-on","metrics":{"RMSE":"0.6911"},"code_links":[{"title":"fstrub95/Autoencoders_cf","url":"https://github.com/fstrub95/Autoencoders_cf"},{"title":"SJD1882/Big-Data-Recommender-Systems","url":"https://github.com/SJD1882/Big-Data-Recommender-Systems"},{"title":"jowoojun/collaborative_filtering_keras","url":"https://github.com/jowoojun/collaborative_filtering_keras"},{"title":"Recvani/benchmark","url":"https://github.com/Recvani/benchmark"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/link-prediction-on-douban","task":"Link Prediction","dataset_variant":"Douban","rows":2,"metrics":["AUC"],"first_row_in_archive_order":{"model":"HSRL (DW)","paper":"/paper/learning-topological-representation-for","metrics":{"AUC":"84.2"},"code_links":[{"title":"fuguoji/HSRL","url":"https://github.com/fuguoji/HSRL"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/knowledge-aware-response-selection-with","title":"Knowledge-aware response selection with semantics underlying multi-turn open-domain conversations","date":"2023-07-27","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/infinite-recommendation-networks-a-data","title":"Infinite Recommendation Networks: A Data-Centric Approach","date":"2022-06-03","rows_on_this_dataset":1,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":6,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/a-federated-graph-neural-network-framework","title":"A federated graph neural network framework for privacy-preserving personalization","date":"2022-06-02","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/glocal-k-global-and-local-kernels-for","title":"GLocal-K: Global and Local Kernels for Recommender Systems","date":"2021-08-27","rows_on_this_dataset":1,"code_links":3,"syntology":null},{"paper":"/paper/global-selector-a-new-benchmark-dataset-and","title":"Uni-Encoder: A Fast and Accurate Response Selection Paradigm for Generation-Based Dialogue Systems","date":"2021-06-02","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/fine-grained-post-training-for-improving","title":"Fine-grained Post-training for Improving Retrieval-based Dialogue Systems","date":"2021-05-24","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/fedgnn-federated-graph-neural-network-for","title":"FedGNN: Federated Graph Neural Network for Privacy-Preserving Recommendation","date":"2021-02-09","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/dialogue-response-selection-with-hierarchical","title":"Dialogue Response Selection with Hierarchical Curriculum Learning","date":"2020-12-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/interpretable-recommender-system-with","title":"Interpretable Recommender System With Heterogeneous Information: A Geometric Deep Learning Perspective","date":"2020-09-20","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/do-response-selection-models-really-know-what","title":"Do Response Selection Models Really Know What's Next? Utterance Manipulation Strategies for Multi-turn Response Selection","date":"2020-09-10","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/speaker-aware-bert-for-multi-turn-response","title":"Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots","date":"2020-04-07","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/multi-hop-selector-network-for-multi-turn","title":"Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based Chatbots","date":"2019-11-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/scalable-probabilistic-matrix-factorization","title":"Scalable Probabilistic Matrix Factorization with Graph-Based Priors","date":"2019-08-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/domain-adaptive-training-bert-for-response","title":"An Effective Domain Adaptive Post-Training Method for BERT in Response Selection","date":"2019-08-13","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/one-time-of-interaction-may-not-be-enough-go","title":"One Time of Interaction May Not Be Enough: Go Deep with an Interaction-over-Interaction Network for Response Selection in Dialogues","date":"2019-07-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/inductive-graph-pattern-learning-for","title":"Inductive Matrix Completion Based on Graph Neural Networks","date":"2019-04-26","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":1,"samples_unverified":1,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/190501969","title":"Poly-encoders: Transformer Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring","date":"2019-04-22","rows_on_this_dataset":1,"code_links":7,"syntology":null},{"paper":"/paper/session-based-social-recommendation-via","title":"Session-based Social Recommendation via Dynamic Graph Attention Networks","date":"2019-02-25","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":1,"samples_unverified":12,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-topological-representation-for","title":"Learning Topological Representation for Networks via Hierarchical Sampling","date":"2019-02-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/representation-learning-for-heterogeneous","title":"Representation Learning for Heterogeneous Information Networks via Embedding Events","date":"2019-01-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/interactive-matching-network-for-multi-turn","title":"Interactive Matching Network for Multi-Turn Response Selection in Retrieval-Based Chatbots","date":"2019-01-07","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/multi-turn-response-selection-for-chatbots","title":"Multi-Turn Response Selection for Chatbots with Deep Attention Matching Network","date":"2018-07-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/modeling-multi-turn-conversation-with-deep","title":"Modeling Multi-turn Conversation with Deep Utterance Aggregation","date":"2018-06-24","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/deep-models-of-interactions-across-sets","title":"Deep Models of Interactions Across Sets","date":"2018-03-07","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/graph-convolutional-matrix-completion","title":"Graph Convolutional Matrix Completion","date":"2017-06-07","rows_on_this_dataset":1,"code_links":17,"syntology":null},{"paper":"/paper/geometric-matrix-completion-with-recurrent","title":"Geometric Matrix Completion with Recurrent Multi-Graph Neural Networks","date":"2017-04-22","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/sequential-matching-network-a-new","title":"Sequential Matching Network: A New Architecture for Multi-turn Response Selection in Retrieval-based Chatbots","date":"2016-12-06","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":0,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hybrid-recommender-system-based-on","title":"Hybrid Recommender System based on Autoencoders","date":"2016-06-24","rows_on_this_dataset":2,"code_links":4,"syntology":null},{"paper":"/paper/collaborative-filtering-with-graph","title":"Collaborative Filtering with Graph Information: Consistency and Scalable Methods","date":"2015-12-01","rows_on_this_dataset":2,"code_links":2,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":4,"samples_harvested":31,"samples_ran":8,"samples_unverified":23,"pointer_only_for_licence":2,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}