Browse State-of-the-Art › Token Reduction
Token Reduction
37 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 37 papers with code (78 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Jun 2025 1 repository listedSpecifically, we propose a Spatial Token Fusion (STF) method to learn compact vision tokens for short vision token sequence, where spatial-adjacent tokens are fused into one.
-
30 May 2025 1 repository listedIn the second stage, language descriptions are fed into a powerful reasoning LLM to solve complex video-language understanding tasks.
-
29 May 2025 1 repository listedWe introduce grounded video tokenization, a paradigm that organizes tokens based on panoptic sub-object trajectories rather than fixed patches.
-
26 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedHowever, as the interaction between tokens and layers is complicated, this raises a basic question: Is such a simple single-layer criterion sufficient to identify redundancy?
-
23 May 2025 1 repository listedWe highlight its potential to drive new model architectures and learning strategies that improve robustness, increase interpretability, and better align with the objectives of generative modeling.
-
23 May 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedModern reasoning models, such as OpenAI's o1 and DeepSeek-R1, exhibit impressive problem-solving capabilities but suffer from critical inefficiencies: high inference latency, excessive computational resource…
-
22 May 2025 1 repository listedIn this paper, we introduce CrossLMM, decoupling long video sequences from LMMs via a dual cross-attention mechanism, which substantially reduces visual token quantity with minimal performance degradation.
-
21 May 2025 1 repository listed Syntology ran 5 of 9 samples · 4 unverified · 4 pointer-only (licence)Unlike token reduction methods that focus on token-level redundancy, we identify and study the computation-level redundancy on vision tokens to ensure no information loss.
-
11 Apr 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedVisual Language Models require substantial computational resources for inference due to the additional input tokens needed to represent visual information.
-
5 Apr 2025 1 repository listedTo further enhance the performance on fine-grained visual understanding tasks, we introduce WiCo+, which decomposes the visual tokens in later layers of the LLM.
-
26 Mar 2025 1 repository listedParameter-efficient tuning (PET) aims to transfer pre-trained foundation models to downstream tasks by learning a small number of parameters.
-
10 Mar 2025 1 repository listedTo preserve image details while reducing computational complexity, we propose a text-guided token pruning method with Dynamic Image Pyramid (DIP) integration.
-
5 Jan 2025 1 repository listedRecently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their abilities to generate and understand cross-modal data.
-
1 Jan 2025 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)Diffusion models have emerged as a promising approach for generating high-quality, high-dimensional images.
-
1 Jan 2025 1 repository listedParameter-efficient fine-tuning (PEFT) adapts pre-trained models to new tasks by updating only a small subset of parameters, achieving efficiency but still facing significant inference costs driven by input token length.
-
31 Dec 2024 1 repository listedUltra-fine-grained image recognition (UFGIR) is a challenging task that involves classifying images within a macro-category.
-
30 Dec 2024 1 repository listed Syntology ran 2 of 13 samples · 11 unverifiedLeveraging the unique properties of similarity over importance, we introduce FrameFusion, a novel approach that combines similarity-based merging with importance-based pruning for better token reduction in LVLMs.
-
17 Dec 2024 1 repository listedRe-training the token-reduced model enhances the performance of Mamba, by effectively rebuilding the key knowledge.
-
13 Dec 2024 1 repository listedMultimodal large language models have experienced rapid growth, and numerous different models have emerged.
-
13 Dec 2024 1 repository listed Syntology ran 4 of 34 samples · 30 unverified · 4 pointer-only (licence)Our method introduces a lightweight embedding module decoupled from the ViT forward pass to extract dedicated features for token merging, thereby addressing the restriction from using intermediate features.
-
1 Dec 2024 1 repository listedThe adoption of Vision Transformers (ViTs) in resource-constrained applications necessitates improvements in inference throughput.
-
5 Nov 2024 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Our results reveal a surprising trend: for visual reasoning tasks, the inference-optimal behavior in VLMs is achieved by using the largest LLM that fits within the inference budget while minimizing visual token count -…
-
22 Oct 2024 1 repository listed Syntology ran 6 of 11 samples · 5 unverifiedGiven a light-weight LLM, our LongVU also scales effectively into a smaller size with state-of-the-art video understanding performance.
-
16 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)While token reduction techniques offer a straightforward post-training strategy, we find that applying existing methods directly to SSMs leads to substantial performance drops.
-
3 Oct 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedIn this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction.
-
25 Sep 2024 1 repository listed Syntology ran 13 of 13 samples · 0 unverifiedOur research introduces a novel approach for the long context bottleneck to accelerate LLM inference and reduce GPU memory consumption.
-
17 Sep 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverifiedThe rapid advancement of Multimodal Large Language Models (MLLMs) has led to remarkable performances across various domains.
-
18 Jun 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Representation learning on text-attributed graphs (TAGs) is vital for real-world applications, as they combine semantic textual and contextual structural information.
-
14 Jun 2024 1 repository listed Syntology ran 6 of 12 samples · 6 unverified · 12 pointer-only (licence)This work presents Adaptive Local-then-Global Merging (ALGM), a token reduction method for semantic segmentation networks that use plain Vision Transformers.
-
16 Apr 2024 1 repository listed Syntology ran 12 of 15 samples · 3 unverified · 15 pointer-only (licence)Large language models (LLMs) have shown remarkable performance in various natural language processing tasks.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections