Datasets › SCROLLS

SCROLLS (Standardized CompaRison Over Long Language Sequences)

Introduced by Uri Shaham et al. in SCROLLS: Standardized CompaRison Over Long Language Sequences10 Jan 2022 archive 2025-07-28

** SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts**. SCROLLS contains summarization, question answering, and natural language inference tasks, covering multiple domains, including literature, science, business, and entertainment. The dataset is made available in a unified text-to-text format and host a live leaderboard to facilitate research on model architecture and pretraining methods.

The SCROLLS benchmark contains the datasets GovReport, SummScreenFD, QMSum, QASPER, NarrativeQA, QuALITY and ContractNLI.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Long-range modeling SCROLLS CoLT5 XL Avg. 43.51 CoLT5: Faster Long-Range Transformers with Conditional... — 13 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 42. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CoLT5: Faster Long-Range Transformers with Conditional Computation 0 1 17 Mar 2023 ran 0 of 9 samples (9 unverified)
Adapting Pretrained Text-to-Text Models for Long Text Sequences 1 1 21 Sep 2022 not harvested
Investigating Efficiently Extending Transformers for Long Input Summarization 2 2 8 Aug 2022 not harvested
Efficient Long-Text Understanding with Short-Text Models 1 1 1 Aug 2022 not harvested
UL2: Unifying Language Learning Paradigms 2 2 10 May 2022 ran 0 of 16 samples (16 unverified)
SCROLLS: Standardized CompaRison Over Long Language Sequences 2 3 10 Jan 2022 ran 2 of 8 samples (6 unverified)
LongT5: Efficient Text-To-Text Transformer for Long Sequences 4 3 15 Dec 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SCROLLS

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections