Browse State-of-the-Art › Long-range modeling
Long-range modeling
69 papers with code · 2 benchmarks · 4 datasets archive 2025-07-28
A new task for testing the long-sequence modeling capabilities and efficiency of language models.
Image credit: SCROLLS: Standardized CompaRison Over Long Language Sequences
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SCROLLS (13 rows) | CoLT5 XL | CoLT5: Faster Long-Range Transformers with Conditional Computation | — | Syntology ran 0 of 9 samples · 9 unverified | Compare |
| LRA (7 rows) | S5 | Simplified State Space Layers for Sequence Modeling | code | Syntology ran 4 of 19 samples · 15 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 69 papers with code (95 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 Dec 2023 35 repositories listed Syntology ran 18 of 62 samples · 44 unverified · 28 pointer-only (licence)Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.
-
31 Oct 2021 8 repositories listed Syntology ran 28 of 55 samples · 27 unverified · 3 pointer-only (licence)A central goal of sequence modeling is designing a single principled model that can address sequence data across a range of modalities and tasks, particularly on long-range dependencies.
-
21 Sep 2022 7 repositories listed Syntology ran 12 of 13 samples · 1 unverified · 8 pointer-only (licence)The design choices in the Transformer attention mechanism, including weak inductive bias and quadratic computational complexity, have limited its application for modeling long sequences.
-
9 Aug 2022 6 repositories listed Syntology ran 4 of 19 samples · 15 unverified · 3 pointer-only (licence)Models using structured state space sequence (S4) layers have achieved state-of-the-art performance on long-range sequence modeling tasks.
-
8 Nov 2020 5 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedIn the recent months, a wide spectrum of efficient, fast Transformers have been proposed to tackle this problem, more often than not claiming superior or comparable model quality to vanilla Transformer models.
-
15 Dec 2021 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent work has shown that either (1) increasing the input length or (2) increasing model size can improve the performance of Transformer-based neural models.
-
9 Apr 2024 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Recent advancements in anomaly detection have seen the efficacy of CNN- and transformer-based approaches.
-
28 Dec 2022 3 repositories listed Syntology ran 7 of 15 samples · 8 unverifiedFirst, we use synthetic language modeling tasks to understand the gap between SSMs and attention.
-
31 Mar 2020 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedSpatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics.
-
12 Mar 2024 2 repositories listed Syntology ran 7 of 11 samples · 4 unverified · 11 pointer-only (licence)In this paper, we introduce a Large Kernel Vision Mamba U-shape Network, or LKM-UNet, for medical image segmentation.
-
25 Jan 2024 2 repositories listedCapturing voxel-wise spatial correspondence across distinct modalities is crucial for medical image analysis.
-
30 Nov 2023 2 repositories listedThe recent success of multiple neural architectures like CNNs, Transformers, and MLP-Mixers motivated us to look for similarities and differences between them.
-
12 May 2023 2 repositories listedAnd based on this attention, a network called T-former is designed for image inpainting.
-
8 Aug 2022 2 repositories listedWhile large pretrained Transformer models have proven highly capable at tackling natural language tasks, handling long sequence inputs continues to be a significant challenge.
-
21 Jul 2022 2 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWeakly Supervised Object Localization (WSOL), which aims to localize objects by only using image-level labels, has attracted much attention because of its low annotation cost in real applications.
-
23 Jun 2022 2 repositories listedOn the other hand, a recent variant of S4 called DSS showed that restricting the state matrix to be fully diagonal can still preserve the performance of the original model when using a specific initialization based on…
-
10 May 2022 2 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedOur model also achieve strong results at in-context learning, outperforming 175B GPT-3 on zero-shot SuperGLUE and tripling the performance of T5-XXL on one-shot summarization.
-
22 Apr 2022 2 repositories listedThe overall computing cost of the new building block is as low as O(N logN).
-
27 Mar 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Modeling long range dependencies in sequential data is a fundamental step towards attaining human-level performance in many modalities such as text, vision, audio and video.
-
10 Jan 2022 2 repositories listed Syntology ran 2 of 8 samples · 6 unverifiedNLP benchmarks have largely focused on short texts, such as sentences and paragraphs, even though long texts comprise a considerable amount of natural language in the wild.
-
15 Jul 2025 1 repository listed Syntology ran 2 of 7 samples · 5 unverifiedAchieving equity in healthcare accessibility requires lightweight yet high-performance solutions for medical image segmentation, particularly in resource-limited settings.
-
6 Jul 2025 1 repository listed Syntology ran 0 of 10 samples · 10 unverifiedWe present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for camera-LiDAR multi-modal 3D object detection.
-
25 May 2025 1 repository listedMost publicly available medical segmentation datasets are only partially labeled, with annotations provided for a subset of anatomical structures.
-
22 May 2025 1 repository listed Syntology ran 2 of 5 samples · 3 unverifiedMasked language models (MLMs) allow bidirectional understanding but are inefficient, as only masked tokens contribute to the loss per step.
-
16 May 2025 1 repository listedThe point cloud classification tasks face the dual challenge of efficiently extracting local geometric features while maintaining model complexity.
-
vGamba: Attentive State Space Bottleneck for efficient Long-range Dependencies in Visual Recognition27 Mar 2025 1 repository listedThe interplay of these components ensures that vGamba leverages the low computational demands of SSMs while maintaining the accuracy of attention mechanisms for modeling long-range dependencies in vision tasks.
-
27 Mar 2025 1 repository listedVideo anomaly detection (VAD) methods are mostly CNN-based or Transformer-based, achieving impressive results, but the focus on detection accuracy often comes at the expense of inference speed.
-
9 Mar 2025 1 repository listedBy utilizing the ConvAttn module, we significantly reduce the reliance on self-attention and its involved memory-bound operations while maintaining the representational capability of transformers.
-
5 Jan 2025 1 repository listedMedical image segmentation is a critical task in medical imaging analysis.
-
17 Dec 2024 1 repository listedTowards this goal, we make the best design choice through extensive experiment settings from data curation to context window extending and utilizing: (1) we analyze data sources and length distributions to construct…
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections