Papers › TL;DR: Mining Reddit to Learn Automatic Summarization

TL;DR: Mining Reddit to Learn Automatic Summarization

1 Sep 2017WS 2017 9archive 2025-07-28

Michael V{\"o}lske, Martin Potthast, Shahbaz Syed, Benno Stein

Recent advances in automatic text summarization have used deep neural networks to generate high-quality abstractive summaries, but the performance of these models strongly depends on large amounts of suitable training data. We propose a new method for mining social media for author-provided summaries, taking advantage of the common practice of appending a {``}TL;DR{''} to long posts. A case study using a large Reddit crawl yields the Webis-TLDR-17 dataset, complementing existing corpora primarily from the news genre. Our technique is likely applicable to other social media sites and general web crawls.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Abstractive Text SummarizationDocument SummarizationText Summarization

Datasets

Introduced by this paper, per the archive.

Webis-TLDR-17 Corpus

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections