Methods › General › Model Compression › Pruning › Papers, page 2
Pruning
Papers archive 2025-07-28
archive papers tagged: 3,874 · with a code link: 1,508 · where Syntology ran a sample: 478 (395 with a run with no instrument failure, 83 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (478 of 3,874 tagged: 395 with a run with no instrument failure, 83 where every run was a failure of Syntology's instrument)
Page 2 of 39: papers 101 to 200 of 3,874, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models 26 May 2025 · 1 repository · arXiv:2505.19959
-
Multimodal Machine Translation with Visual Scene Graph Pruning 26 May 2025 · 0 repositories · arXiv:2505.19507
-
Pangu Light: Weight Re-Initialization for Pruning and Accelerating LLMs 26 May 2025 · 0 repositories · arXiv:2505.20155
-
syftr: Pareto-Optimal Generative AI 26 May 2025 · 1 repository · arXiv:2505.20266
-
Context-Driven Dynamic Pruning for Large Speech Foundation Models 24 May 2025 · 0 repositories · arXiv:2505.18860
-
Is Attention Required for Transformer Inference? Explore Function-preserving Attention Replacement 24 May 2025 · 0 repositories · arXiv:2505.21535
-
μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts 24 May 2025 · 0 repositories · arXiv:2505.18451
-
Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer 24 May 2025 · 0 repositories · arXiv:2505.18713
-
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning 24 May 2025 · 0 repositories · arXiv:2505.18563
-
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT 24 May 2025 · 0 repositories · arXiv:2506.02005
-
ToDRE: Visual Token Pruning via Diversity and Task Awareness for Efficient Large Vision-Language Models 24 May 2025 · 0 repositories · arXiv:2505.18757
-
ELDeR: Getting Efficient LLMs through Data-Driven Regularized Layer-wise Pruning 23 May 2025 · 0 repositories · arXiv:2505.18232
-
Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems 23 May 2025 · 0 repositories · arXiv:2505.18139
-
Task Specific Pruning with LLM-Sieve: How Many Parameters Does Your Task Really Need? 23 May 2025 · 0 repositories · arXiv:2505.18350
-
Align-GRAG: Reasoning-Guided Dual Alignment for Graph Retrieval-Augmented Generation 22 May 2025 · 0 repositories · arXiv:2505.16237
-
Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning 22 May 2025 · 0 repositories · arXiv:2505.16950
-
Fixing Data That Hurts Performance: Cascading LLMs to Relabel Hard Negatives for Robust Information Retrieval 22 May 2025 · 0 repositories · arXiv:2505.16967
-
Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models 22 May 2025 · 1 repository · arXiv:2505.16104
-
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 22 May 2025 · 1 repository · arXiv:2505.16175
-
RAP: Runtime-Adaptive Pruning for LLM Inference 22 May 2025 · 0 repositories · arXiv:2505.17138
-
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning 22 May 2025 · 1 repository · arXiv:2505.16743
-
Clustering and Pruning in Causal Data Fusion 21 May 2025 · 1 repository · arXiv:2505.15215
-
DeFTX: Denoised Sparse Fine-Tuning for Zero-Shot Cross-Lingual Transfer 21 May 2025 · 0 repositories · arXiv:2505.15090
-
DiffProb: Data Pruning for Face Recognition 21 May 2025 · 1 repository · arXiv:2505.15272
-
On the creation of narrow AI: hierarchy and nonlocality of neural network skills 21 May 2025 · 1 repository · arXiv:2505.15811
-
Self-GIVE: Associative Thinking from Limited Structured Knowledge for Enhanced Large Language Model Reasoning 21 May 2025 · 0 repositories · arXiv:2505.15062
-
Adaptive Pruning of Deep Neural Networks for Resource-Aware Embedded Intrusion Detection on the Edge 20 May 2025 · 1 repository · arXiv:2505.14592
-
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning 20 May 2025 · 0 repositories · arXiv:2505.14582
-
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models 20 May 2025 · 0 repositories · arXiv:2505.13975
-
Improved Methods for Model Pruning and Knowledge Distillation 20 May 2025 · 0 repositories · arXiv:2505.14052
-
A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone 19 May 2025 · 1 repository · arXiv:2505.12781
-
A3 : an Analytical Low-Rank Approximation Framework for Attention 19 May 2025 · 0 repositories · arXiv:2505.12942
-
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning 19 May 2025 · 0 repositories · arXiv:2505.12782
-
An Overview of Arithmetic Adaptations for Inference of Convolutional Neural Networks on Re-configurable Hardware 19 May 2025 · 1 repository · arXiv:2505.13575
-
Automatic Complementary Separation Pruning Toward Lightweight CNNs 19 May 2025 · 0 repositories · arXiv:2505.13225
-
Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning 19 May 2025 · 0 repositories · arXiv:2505.13327
-
DimGrow: Memory-Efficient Field-level Embedding Dimension Search 19 May 2025 · 0 repositories · arXiv:2505.12683
-
Efficient training for large-scale optical neural network using an evolutionary strategy and attention pruning 19 May 2025 · 0 repositories · arXiv:2505.12906
-
Exploring Federated Pruning for Large Language Models 19 May 2025 · 1 repository · arXiv:2505.13547
-
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding 19 May 2025 · 0 repositories · arXiv:2505.13254
-
ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving 19 May 2025 · 0 repositories · arXiv:2505.12717
-
Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-Constrained Pruning 18 May 2025 · 0 repositories · arXiv:2505.12281
-
One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models 18 May 2025 · 0 repositories · arXiv:2505.12216
-
PMQ-VE: Progressive Multi-Frame Quantization for Video Enhancement 18 May 2025 · 1 repository · arXiv:2505.12266
-
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference 18 May 2025 · 0 repositories · arXiv:2505.12359
-
SepPrune: Structured Pruning for Efficient Deep Speech Separation 17 May 2025 · 1 repository · arXiv:2505.12079
-
EA-3DGS: Efficient and Adaptive 3D Gaussians with Highly Enhanced Quality for outdoor scenes 16 May 2025 · 1 repository · arXiv:2505.10787
-
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL 16 May 2025 · 2 repositories · arXiv:2505.10832
-
Nosy Layers, Noisy Fixes: Tackling DRAs in Federated Learning Systems using Explainable AI 16 May 2025 · 0 repositories · arXiv:2505.10942
-
Addition is almost all you need: Compressing neural networks with double binary factorization 16 May 2025 · 1 repository · arXiv:2505.11076Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
SubROC: AUC-Based Discovery of Exceptional Subgroup Performance for Binary Classifiers 16 May 2025 · 0 repositories · arXiv:2505.11283
-
Dynamic Base model Shift for Delta Compression 16 May 2025 · 0 repositories · arXiv:2505.11344
-
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning 16 May 2025 · 0 repositories · arXiv:2505.11471
-
BINGO: A Novel Pruning Mechanism to Reduce the Size of Neural Networks 15 May 2025 · 0 repositories · arXiv:2505.09864
-
Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering 15 May 2025 · 0 repositories · arXiv:2505.10118
-
Interim Report on Human-Guided Adaptive Hyperparameter Optimization with Multi-Fidelity Sprints 14 May 2025 · 0 repositories · arXiv:2505.09792
-
The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks 14 May 2025 · 0 repositories · arXiv:2505.09214
-
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression 13 May 2025 · 0 repositories · arXiv:2505.18166
-
Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments 13 May 2025 · 0 repositories · arXiv:2505.08299
-
Low-Complexity Inference in Continual Learning via Compressed Knowledge Transfer 13 May 2025 · 0 repositories · arXiv:2505.08327
-
SPAT: Sensitivity-based Multihead-attention Pruning on Time Series Forecasting Models 13 May 2025 · 0 repositories · arXiv:2505.08768
-
Channel Fingerprint Construction for Massive MIMO: A Deep Conditional Generative Approach 12 May 2025 · 0 repositories · arXiv:2505.07893
-
ICE-Pruning: An Iterative Cost-Efficient Pruning Pipeline for Deep Neural Networks 12 May 2025 · 1 repository · arXiv:2505.07411
-
Semantic Retention and Extreme Compression in LLMs: Can We Have Both? 12 May 2025 · 0 repositories · arXiv:2505.07289
-
Solving Nonlinear PDEs with Sparse Radial Basis Function Networks 12 May 2025 · 0 repositories · arXiv:2505.07765
-
Bi-LSTM based Multi-Agent DRL with Computation-aware Pruning for Agent Twins Migration in Vehicular Embodied AI Networks 9 May 2025 · 0 repositories · arXiv:2505.06378
-
DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing 8 May 2025 · 0 repositories · arXiv:2505.05413
-
Guiding Evolutionary AutoEncoder Training with Activation-Based Pruning Operators 8 May 2025 · 1 repository · arXiv:2505.05138
-
Text2Cypher: Data Pruning using Hard Example Selection 8 May 2025 · 0 repositories · arXiv:2505.05122
-
Towards Initialization-Agnostic Clustering with Iterative Adaptive Resonance Theory 7 May 2025 · 0 repositories · arXiv:2505.04440
-
Mitigating mode collapse in normalizing flows by annealing with an adaptive schedule: Application to parameter estimation 6 May 2025 · 0 repositories · arXiv:2505.03652
-
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods 6 May 2025 · 0 repositories · arXiv:2505.03373
-
End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning 5 May 2025 · 0 repositories · arXiv:2505.13462
-
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques 5 May 2025 · 0 repositories · arXiv:2505.02309
-
ReplaceMe: Network Simplification via Layer Pruning and Linear Transformations 5 May 2025 · 1 repository · arXiv:2505.02819Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Optimization over Trained (and Sparse) Neural Networks: A Surrogate within a Surrogate 4 May 2025 · 0 repositories · arXiv:2505.01985
-
Efficient Shapley Value-based Non-Uniform Pruning of Large Language Models 3 May 2025 · 0 repositories · arXiv:2505.01731
-
A-DARTS: Stable Model Selection for Data Repair in Time Series 1 May 2025 · 1 repository
-
Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Embeddings 1 May 2025 · 0 repositories · arXiv:2505.00560
-
FineScope : Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation 1 May 2025 · 0 repositories · arXiv:2505.00624
-
Redundancy Analysis and Mitigation for Machine Learning-Based Process Monitoring of Additive Manufacturing 30 Apr 2025 · 0 repositories · arXiv:2504.21317
-
Scalable Multi-Task Learning for Particle Collision Event Reconstruction with Heterogeneous Graph Neural Networks 30 Apr 2025 · 0 repositories · arXiv:2504.21844
-
CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks 29 Apr 2025 · 0 repositories · arXiv:2504.21228
-
Efficient LLMs with AMP: Attention Heads and MLP Pruning 29 Apr 2025 · 2 repositories · arXiv:2504.21174
-
Small brains but big challenges: white matter tractography in early life samples 29 Apr 2025 · 0 repositories · arXiv:2504.20554
-
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs 28 Apr 2025 · 0 repositories · arXiv:2504.19659
-
LODAP: On-Device Incremental Learning Via Lightweight Operations and Data Pruning 28 Apr 2025 · 1 repository · arXiv:2504.19638
-
Towards Faster and More Compact Foundation Models for Molecular Property Prediction 28 Apr 2025 · 1 repository · arXiv:2504.19538
-
Uncovering potential effects of spontaneous waves on synaptic development: the visual system as a model 26 Apr 2025 · 0 repositories · arXiv:2504.18991
-
Coding for Computation: Efficient Compression of Neural Networks for Reconfigurable Hardware 24 Apr 2025 · 0 repositories · arXiv:2504.17403
-
BackSlash: Rate Constrained Optimized Training of Large Language Models 23 Apr 2025 · 0 repositories · arXiv:2504.16968
-
Dynamic Superblock Pruning for Fast Learned Sparse Retrieval 23 Apr 2025 · 0 repositories · arXiv:2504.17045
-
DONOD: Robust and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning 21 Apr 2025 · 0 repositories · arXiv:2504.14810
-
Connecting Parameter Magnitudes and Hessian Eigenspaces at Scale using Sketched Methods 20 Apr 2025 · 0 repositories · arXiv:2504.14701
-
Matrix Factorization with Dynamic Multi-view Clustering for Recommender System 20 Apr 2025 · 0 repositories · arXiv:2504.14565
-
NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models 20 Apr 2025 · 1 repository · arXiv:2504.14569Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding 20 Apr 2025 · 0 repositories · arXiv:2504.14692
-
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator 19 Apr 2025 · 0 repositories · arXiv:2504.14365
-
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs 18 Apr 2025 · 0 repositories · arXiv:2504.13471
-
D²MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving 17 Apr 2025 · 0 repositories · arXiv:2504.15299