Datasets › SCROLLS
SCROLLS (Standardized CompaRison Over Long Language Sequences)
** SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts**. SCROLLS contains summarization, question answering, and natural language inference tasks, covering multiple domains, including literature, science, business, and entertainment. The dataset is made available in a unified text-to-text format and host a live leaderboard to facilitate research on model architecture and pretraining methods.
The SCROLLS benchmark contains the datasets GovReport, SummScreenFD, QMSum, QASPER, NarrativeQA, QuALITY and ContractNLI.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Long-range modeling | SCROLLS | CoLT5 XL Avg. 43.51 | CoLT5: Faster Long-Range Transformers with Conditional... | — | 13 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 42. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| CoLT5: Faster Long-Range Transformers with Conditional Computation | 0 | 1 | 17 Mar 2023 | ran 0 of 9 samples (9 unverified) |
| Adapting Pretrained Text-to-Text Models for Long Text Sequences | 1 | 1 | 21 Sep 2022 | not harvested |
| Investigating Efficiently Extending Transformers for Long Input Summarization | 2 | 2 | 8 Aug 2022 | not harvested |
| Efficient Long-Text Understanding with Short-Text Models | 1 | 1 | 1 Aug 2022 | not harvested |
| UL2: Unifying Language Learning Paradigms | 2 | 2 | 10 May 2022 | ran 0 of 16 samples (16 unverified) |
| SCROLLS: Standardized CompaRison Over Long Language Sequences | 2 | 3 | 10 Jan 2022 | ran 2 of 8 samples (6 unverified) |
| LongT5: Efficient Text-To-Text Transformer for Long Sequences | 4 | 3 | 15 Dec 2021 | ran 1 of 1 samples (0 unverified; 1 pointer-only for licence) |
Dataset loaders archive 2025-07-28
2 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- SCROLLS
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections