Datasets › Ubuntu IRC

Ubuntu IRC

Introduced by Jonathan K. Kummerfeld et al. in A Large-Scale Corpus for Conversation Disentanglement25 Oct 2018 archive 2025-07-28

The Ubuntu IRC dataset is a valuable resource for research in natural language understanding and dialogue systems. Let me provide you with some details:

  1. Ubuntu Dialogue Corpus:

    • This dataset contains almost 1 million multi-turn dialogues, comprising over 7 million utterances and 100 million words.
    • It serves as a unique resource for building dialogue managers based on neural language models that can leverage large amounts of unlabeled data³.
    • The dialogues are sourced from IRC (Internet Relay Chat) conversations related to Ubuntu, a popular open-source operating system.
    • Researchers can use this corpus to explore various aspects of dialogue understanding and generation.
  2. Specifics of the Ubuntu IRC Dataset:

    • The dataset includes 77,563 annotated messages from IRC.
    • Most of these messages originate from the Ubuntu IRC Logs for the #ubuntu channel.
    • Additionally, a smaller subset is a re-annotation of data from the #linux channel, which was originally collected by Elsner and Charniak in 2008².
    • You can find this dataset in the kummerfeld/data folder of the repository².

In summary, the Ubuntu IRC dataset provides a rich collection of dialogues that researchers can use to advance the field of natural language processing and dialogue modeling. 🌐🗣️

(1) The Ubuntu Dialogue Corpus: A Large Dataset for Research in .... https://arxiv.org/abs/1506.08909. (2) GitHub - amarazad/DSRNet: Meta-Context Transformers for Domain-Specific .... https://github.com/amarazad/DSRNet. (3) Ubuntu Dialogue Corpus | Kaggle. https://www.kaggle.com/datasets/rtatman/ubuntu-dialogue-corpus. (4) GitHub - jkkummerfeld/irc-disentanglement: Dataset and model for .... https://github.com/jkkummerfeld/irc-disentanglement.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 24. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Ubuntu IRC

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections