Papers › MuLD: The Multitask Long Document Benchmark

MuLD: The Multitask Long Document Benchmark

15 Feb 2022LREC 2022 6arXiv:2202.07362archive 2025-07-28

G Thomas Hudson, Noura Al Moubayed

The impressive progress in NLP techniques has been driven by the development of multi-task benchmarks such as GLUE and SuperGLUE. While these benchmarks focus on tasks for one or two input sentences, there has been exciting work in designing efficient techniques for processing much longer inputs. In this paper, we present MuLD: a new long document benchmark consisting of only documents over 10,000 tokens. By modifying existing NLP tasks, we create a diverse benchmark which requires models to successfully model long-term dependencies in the text. We evaluate how existing models perform, and find that our benchmark is much more challenging than their `short document' equivalents. Furthermore, by evaluating both regular and efficient transformers, we show that models with increased context length are better able to solve the tasks presented, suggesting that future improvements in these models are vital for solving similar long document problems. We release the data and code for baselines to encourage further research on efficient NLP models.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ghomashudson/muld officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringStyle change detectionSummarizationText ClassificationTranslation

Datasets

Introduced by this paper, per the archive.

MuLD

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering MuLD (HotpotQA) Longformer BLEU-1 30.38 #1 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) Longformer BLEU-4 16.76 #1 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) Longformer METEOR 4.98 #1 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) Longformer Rouge-L 30.49 #1 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) T5 BLEU-1 28.11 #2 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) T5 BLEU-4 13.63 #2 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) T5 METEOR 4.46 #2 of 2 Archive leaderboard report
Question Answering MuLD (HotpotQA) T5 Rouge-L 27.61 #2 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) Longformer BLEU-1 19.84 #1 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) Longformer BLEU-4 62 #1 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) Longformer METEOR 4.52 #1 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) Longformer Rouge-L 22.09 #1 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) T5 BLEU-1 17.67 #2 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) T5 BLEU-4 55 #2 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) T5 METEOR 3.36 #2 of 2 Archive leaderboard report
Question Answering MuLD (NarrativeQA) T5 Rouge-L 19.03 #2 of 2 Archive leaderboard report
Summarization MuLD (VLSP) Longformer BLEU-1 46.74 #1 of 2 Archive leaderboard report
Summarization MuLD (VLSP) Longformer BLEU-4 3.05 #1 of 2 Archive leaderboard report
Summarization MuLD (VLSP) Longformer METEOR 9.58 #1 of 2 Archive leaderboard report
Summarization MuLD (VLSP) Longformer Rouge-L 19.52 #1 of 2 Archive leaderboard report
Summarization MuLD (VLSP) T5 BLEU-1 28.85 #2 of 2 Archive leaderboard report
Summarization MuLD (VLSP) T5 BLEU-4 84 #2 of 2 Archive leaderboard report
Summarization MuLD (VLSP) T5 METEOR 7.98 #2 of 2 Archive leaderboard report
Summarization MuLD (VLSP) T5 Rouge-L 16.55 #2 of 2 Archive leaderboard report
Text Classification MuLD (Character Type) Longformer F1 82.58 #1 of 2 Archive leaderboard report
Text Classification MuLD (Character Type) T5 F1 54.01 #2 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) T5 BLEU-1 34.07 #1 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) T5 BLEU-4 1.63 #1 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) T5 METEOR 38.53 #1 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) T5 Rouge-L 35.35 #1 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) Longformer BLEU-1 22.74 #2 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) Longformer BLEU-4 20 #2 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) Longformer METEOR 22.95 #2 of 2 Archive leaderboard report
Translation MuLD (OpenSubtitles) Longformer Rouge-L 22.17 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections