{"url":"/dataset/ubuntu-irc-1","name":"Ubuntu IRC","full_name":null,"description_markdown":"The **Ubuntu IRC dataset** is a valuable resource for research in natural language understanding and dialogue systems. Let me provide you with some details:\r\n\r\n1. **Ubuntu Dialogue Corpus**:\r\n    - This dataset contains **almost 1 million multi-turn dialogues**, comprising over **7 million utterances** and **100 million words**.\r\n    - It serves as a unique resource for building dialogue managers based on neural language models that can leverage large amounts of unlabeled data³.\r\n    - The dialogues are sourced from **IRC (Internet Relay Chat)** conversations related to Ubuntu, a popular open-source operating system.\r\n    - Researchers can use this corpus to explore various aspects of dialogue understanding and generation.\r\n\r\n2. **Specifics of the Ubuntu IRC Dataset**:\r\n    - The dataset includes **77,563 annotated messages** from IRC.\r\n    - Most of these messages originate from the **Ubuntu IRC Logs for the #ubuntu channel**.\r\n    - Additionally, a smaller subset is a re-annotation of data from the **#linux channel**, which was originally collected by Elsner and Charniak in 2008².\r\n    - You can find this dataset in the **kummerfeld/data folder** of the repository².\r\n\r\nIn summary, the Ubuntu IRC dataset provides a rich collection of dialogues that researchers can use to advance the field of natural language processing and dialogue modeling. 🌐🗣️\r\n\r\n(1) The Ubuntu Dialogue Corpus: A Large Dataset for Research in .... https://arxiv.org/abs/1506.08909.\r\n(2) GitHub - amarazad/DSRNet: Meta-Context Transformers for Domain-Specific .... https://github.com/amarazad/DSRNet.\r\n(3) Ubuntu Dialogue Corpus | Kaggle. https://www.kaggle.com/datasets/rtatman/ubuntu-dialogue-corpus.\r\n(4) GitHub - jkkummerfeld/irc-disentanglement: Dataset and model for .... https://github.com/jkkummerfeld/irc-disentanglement.","description_withheld":null,"homepage":"https://github.com/jkkummerfeld/irc-disentanglement","introduced_date":"2018-10-25","introduced_date_note":null,"introduced_by":{"paper":"/paper/analyzing-assumptions-in-conversation","title":"A Large-Scale Corpus for Conversation Disentanglement","first_author":"Jonathan K. Kummerfeld","url":null},"license":null,"modalities":[],"tasks":[{"name":"Conversational Response Selection","url":"/task/conversational-response-selection","datasets_with_task":"/datasets/task/conversational-response-selection"}],"languages":[],"variants":["Ubuntu IRC"],"data_loaders":[],"num_papers_in_archive":24,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/conversational-response-selection-on-ubuntu-3","task":"Conversational Response Selection","dataset_variant":"Ubuntu IRC","rows":5,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"MPC-BERT","paper":"/paper/mpc-bert-a-pre-trained-language-model-for","metrics":{"Accuracy":"63.64"},"code_links":[{"title":"JasonForJoy/MPC-BERT","url":"https://github.com/JasonForJoy/MPC-BERT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/language-modelling-on-ubuntu-irc","task":"Language Modelling","dataset_variant":"Ubuntu IRC","rows":1,"metrics":["BPB"],"first_row_in_archive_order":{"model":"Gopher","paper":"/paper/scaling-language-models-methods-analysis-1","metrics":{"BPB":"1.09"},"code_links":[{"title":"allenai/dolma","url":"https://github.com/allenai/dolma"},{"title":"rvlopes/gloria","url":"https://github.com/rvlopes/gloria"},{"title":"bramiozo/PubScience","url":"https://github.com/bramiozo/PubScience"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/scaling-language-models-methods-analysis-1","title":"Scaling Language Models: Methods, Analysis & Insights from Training Gopher","date":"2021-12-08","rows_on_this_dataset":1,"code_links":3,"syntology":null},{"paper":"/paper/mpc-bert-a-pre-trained-language-model-for","title":"MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation Understanding","date":"2021-06-03","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/speaker-aware-bert-for-multi-turn-response","title":"Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots","date":"2020-04-07","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/addressee-and-response-selection-in-multi","title":"Addressee and Response Selection in Multi-Party Conversations with Speaker Interaction RNNs","date":"2017-09-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/addressee-and-response-selection-for-multi","title":"Addressee and Response Selection for Multi-Party Conversation","date":"2016-11-01","rows_on_this_dataset":2,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}