| Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration added by Syntology |
2026-06 (from id) |
pietrotrope/Dont_Go_Breaking_My_LLM/track_block_activations.py 088898428fc14c3e |
ran
|
no licence file found · pointer only |
| AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization added by Syntology |
2026-06 (from id) |
Superone77/AlphaQ/modelutils.py 42a68e1344b46dfb |
ran
|
no licence file found · pointer only |
| 3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs added by Syntology |
2026-03 (from id) |
mazumder-lab/3BASiL/pipeline.py a27cec0b3dcb1fec |
unverified |
Apache-2.0 (permissive) |
| SFMP: Fine-Grained, Hardware-Friendly and Search-Free Mixed-Precision Quantization for Large Language Models added by Syntology |
2026-02 (from id) |
Nkniexin/SFMP/GPTQ/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Fine-Grained Expert Merging and Bit-packed Inference added by Syntology |
2025-11 (from id) |
Supercomputing-System-AI-Lab/PuzzleMoE/puzzlemoe/utils/merge_experts_function.py 14fab8e91e876c60 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PT$^2$-LLM: Post-Training Ternarization for Large Language Models added by Syntology |
2025-10 (from id) |
XIANGLONGYAN/PT2-LLM/pt2_llm/model_utils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM added by Syntology |
2025-10 (from id) |
log-postech/elsa/lib/utils.py 6644b68819429769 |
ran · our draft was wrong
|
no licence file found · pointer only |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/prune/lib/prune_mlp_unif.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| SAFE: Finding Sparse and Flat Minima to Improve Pruning |
7 Jun 2025 |
LOG-postech/safe-torch/language/lib/prune.py 6644b68819429769 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2506.03781 |
2025-06 (from id) |
OpenGVLab/OmniQuant/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
MIT (permissive) |
| An Empirical Study of Qwen3 Quantization |
4 May 2025 |
efficient-ml/qwen3-quantization/BiLLM/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning |
17 Dec 2024 |
xinykou/nlsr/src/prune_regions/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Squeezed Attention: Accelerating Long Context Length LLM Inference |
14 Nov 2024 |
SqueezeAILab/SqueezedAttention/utils/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization |
8 Nov 2024 |
georgia-tech-synergy-lab/microscopiq-llm-quantization/utils/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models |
12 Oct 2024 |
vengdeng/darex/utils/getwx.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning |
9 Oct 2024 |
abx393/llm-pruning-calibration-data/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| MC-MoE: Mixture Compressor for Mixture-of-Experts LLMs Gains More |
8 Oct 2024 |
Aaronhuang-778/MC-MoE/modelutils.py 42a68e1344b46dfb |
ran
|
no licence file found · pointer only |
| Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization |
27 Sep 2024 |
usamec/double_sparse/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models |
25 Sep 2024 |
microsoft/vptq/vptq/layers/utils.py 6886c53e21133de1 |
ran
|
MIT (permissive) |
| OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition |
20 Sep 2024 |
stephenqz/oats/OATS/pruning_utils.py 14fab8e91e876c60 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Eigen Attention: Attention in Low-Rank Space for KV Cache Compression |
10 Aug 2024 |
utkarshsaxena1/eigenattn/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
no licence file found · pointer only |
| LeanQuant: Accurate Large Language Model Quantization with Loss-Error-Aware Grid |
14 Jul 2024 |
LeanModels/LeanQuant/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs |
24 Jun 2024 |
kiddyboots216/lottery-ticket-adaptation/rlaif/save_mask.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization |
10 Jun 2024 |
gatech-eic/shiftaddllm/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models |
5 Jun 2024 |
pprp/pruner-zero/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
MIT (permissive) |
| Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models |
5 Jun 2024 |
pprp/pruner-zero/lora_ft/evaluate_ppl.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models |
5 Jun 2024 |
pprp/pruner-zero/lib/gradient_computation.py 14fab8e91e876c60 |
ran · our draft was wrong
|
MIT (permissive) |
| DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs |
3 Jun 2024 |
Hsu1023/DuQuant/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
MIT (permissive) |
| MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization |
2 Jun 2024 |
AozhongZhang/MagR/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| Sparse Expansion and Neuronal Disentanglement |
24 May 2024 |
shavit-lab/sparse-expansion/utils/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models |
23 May 2024 |
Aaronhuang-778/SliM-LLM/slim-llm/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models |
23 May 2024 |
Aaronhuang-778/SliM-LLM/slim-llm-plus/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
no licence file found · pointer only |
| A safety realignment framework via subspace-oriented model fusion for large language models |
15 May 2024 |
xinykou/safety_realignment/lm_eval/code_task/modeling.py 3ab3c57755867f9a |
ran
|
no licence file found · pointer only |
| An empirical study of LLaMA3 quantization: from LLMs to MLLMs |
22 Apr 2024 |
macaronlin/llama3-quantization/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
no licence file found · pointer only |
| AffineQuant: Affine Transformation Quantization for Large Language Models |
19 Mar 2024 |
bytedance/affinequant/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
Apache-2.0 (permissive) |
| FrameQuant: Flexible Low-Bit Quantization for Transformers |
10 Mar 2024 |
vsingh-group/framequant/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact |
2 Mar 2024 |
squeezeailab/kvquant/benchmarking/kvquant/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| SparseLLM: Towards Global Pruning for Pre-trained Language Models |
28 Feb 2024 |
BaiTheBest/SparseLLM/pruning_utils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation |
18 Feb 2024 |
linkanonymous/besa/utils/tools.py 14fab8e91e876c60 |
ran · our draft was wrong
|
no licence file found · pointer only |
| OneBit: Towards Extremely Low-bit Large Language Models |
17 Feb 2024 |
xuyuzhuang11/OneBit/evaluation/lm_eval/models_utils.py 904cd61df2fe8fb9 |
ran
|
MIT (permissive) |
| Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers |
14 Feb 2024 |
SamsungLabs/aespa/utils.py 1174e949ac22ea59 |
ran · our draft was wrong
|
licence not identified · pointer only |
| BiLLM: Pushing the Limit of Post-Training Quantization for LLMs |
6 Feb 2024 |
aaronhuang-778/billm/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization |
31 Jan 2024 |
SqueezeAILab/KVQuant/deployment/kvquant/modelutils.py 71eb7727feb23c97 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Fast and Effective Weight Update for Pruned Large Language Models |
1 Jan 2024 |
fmfi-compbio/admm-pruning/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
MIT (permissive) |
| Fluctuation-based Adaptive Structured Pruning for Large Language Models |
19 Dec 2023 |
casia-iva-lab/flap/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models |
10 Dec 2023 |
hahnyuan/asvd4llm/quantization.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| Post-training Quantization for Text-to-Image Diffusion Models with Progressive Calibration and Activation Relaxing |
10 Nov 2023 |
tsa18/PCR/quantization_tools/quantization/quantizer_utils.py 819214853c3d34c7 |
ran
|
no licence file found · pointer only |
| Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models |
8 Nov 2023 |
rocktimjyotidas/gblm-pruner/gradient_computation.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
MIT (permissive) |
| QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models |
25 Oct 2023 |
ist-daslab/qmoe/switch.py 87506509921686ee |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models |
14 Oct 2023 |
talkking/MixGPT/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs |
13 Oct 2023 |
zyxxmu/dsnot/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
no licence file found · pointer only |
| QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models |
12 Oct 2023 |
modeltc/qllm/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
Apache-2.0 (permissive) |
| Composite Backdoor Attacks Against Large Language Models |
11 Oct 2023 |
miraclehh/cba/multimodal/llama2_accessory/eval/modeling.py 3ab3c57755867f9a |
ran
|
no licence file found · pointer only |
| PB-LLM: Partially Binarized Large Language Models |
29 Sep 2023 |
hahnyuan/binaryllm/gptq_pb/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers |
28 Sep 2023 |
kuleshov-group/llmtools/llmtools/utils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
no licence file found · pointer only |
| Sight Beyond Text: Multi-Modal Training Enhances LLMs in Truthfulness and Ethics |
13 Sep 2023 |
ucsc-vlaa/sight-beyond-text/llava/eval/mmlu/mmlu_modeling.py 3ab3c57755867f9a |
ran
|
Apache-2.0 (permissive) |
| OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models |
25 Aug 2023 |
opengvlab/omniquant/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
MIT (permissive) |
| ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation |
28 Jul 2023 |
lianjiatech/belle/models/gptq/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Llama 2: Open Foundation and Fine-Tuned Chat Models |
18 Jul 2023 |
squeezeailab/squeezellm/squeezellm/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| A Simple and Effective Pruning Approach for Large Language Models |
20 Jun 2023 |
locuslab/wanda/lib/prune.py 56ac52f94f2cc872 |
ran · our draft was wrong
|
MIT (permissive) |
| RPTQ: Reorder-based Post-training Quantization for Large Language Models |
3 Apr 2023 |
hahnyuan/rptq4llm/models/models_utils.py 904cd61df2fe8fb9 |
ran
|
MIT (permissive) |
| CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis |
25 Mar 2022 |
openlmlab/moss/models/quantization.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:openreview_d3RFDLBw01 |
|
AI2C-Lab/STLA/utils.py 1174e949ac22ea59 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:openreview_Im05D8gFFn |
|
mazumder-lab/RobOP/RobOP-ALPS/utils_models.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:aaai_28960 |
|
CASIA-IVA-Lab/FLAP/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2025.findings-emnlp.1054 |
|
IST-DASLab/sparsegpt/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2025.findings-acl.394 |
|
TheShineyue/HSR/alignment-attribution-hsr/lib/prune.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
MIT (permissive) |
| arXiv:2025.acl-long.590 |
|
ChnQ/SPIN/src/SPIN.py 2ed16fcdc14ad951 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-naacl.145 |
|
LianjiaTech/BELLE/models/gptq/modelutils.py a9e7f2cdf016b88b |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| arXiv:2024.findings-acl.933 |
|
AutoGPTQ/AutoGPTQ/auto_gptq/modeling/_utils.py ac48f2f323eda99e |
unverified |
MIT (permissive) |
| arXiv:2023.emnlp-main.892 |
|
SamsungLabs/Z-Fold/zfold.py 602607980dfc128c |
unverified |
MIT (permissive) |