Datasets › Ubuntu IRC
Ubuntu IRC
The Ubuntu IRC dataset is a valuable resource for research in natural language understanding and dialogue systems. Let me provide you with some details:
-
Ubuntu Dialogue Corpus:
- This dataset contains almost 1 million multi-turn dialogues, comprising over 7 million utterances and 100 million words.
- It serves as a unique resource for building dialogue managers based on neural language models that can leverage large amounts of unlabeled data³.
- The dialogues are sourced from IRC (Internet Relay Chat) conversations related to Ubuntu, a popular open-source operating system.
- Researchers can use this corpus to explore various aspects of dialogue understanding and generation.
-
Specifics of the Ubuntu IRC Dataset:
- The dataset includes 77,563 annotated messages from IRC.
- Most of these messages originate from the Ubuntu IRC Logs for the #ubuntu channel.
- Additionally, a smaller subset is a re-annotation of data from the #linux channel, which was originally collected by Elsner and Charniak in 2008².
- You can find this dataset in the kummerfeld/data folder of the repository².
In summary, the Ubuntu IRC dataset provides a rich collection of dialogues that researchers can use to advance the field of natural language processing and dialogue modeling. 🌐🗣️
(1) The Ubuntu Dialogue Corpus: A Large Dataset for Research in .... https://arxiv.org/abs/1506.08909. (2) GitHub - amarazad/DSRNet: Meta-Context Transformers for Domain-Specific .... https://github.com/amarazad/DSRNet. (3) Ubuntu Dialogue Corpus | Kaggle. https://www.kaggle.com/datasets/rtatman/ubuntu-dialogue-corpus. (4) GitHub - jkkummerfeld/irc-disentanglement: Dataset and model for .... https://github.com/jkkummerfeld/irc-disentanglement.
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Conversational Response Selection | Ubuntu IRC | MPC-BERT Accuracy 63.64 | MPC-BERT: A Pre-Trained Language Model for Multi-Party... | JasonForJoy/MPC-BERT | 5 | Compare |
| Language Modelling | Ubuntu IRC | Gopher BPB 1.09 | Scaling Language Models: Methods, Analysis & Insights... | allenai/dolma +2 | 1 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 24. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Scaling Language Models: Methods, Analysis & Insights from Training Gopher | 3 | 1 | 8 Dec 2021 | not harvested |
| MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation Understanding | 1 | 1 | 3 Jun 2021 | not harvested |
| Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based Chatbots | 2 | 1 | 7 Apr 2020 | not harvested |
| Addressee and Response Selection in Multi-Party Conversations with Speaker Interaction RNNs | 1 | 1 | 12 Sep 2017 | not harvested |
| Addressee and Response Selection for Multi-Party Conversation | 1 | 2 | 1 Nov 2016 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Ubuntu IRC
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections