| Rethinking Residual Errors in Compensation-based LLM Quantization added by Syntology |
2026-04 (from id) |
list0830/ResComp/fake_quant/data_utils.py 347f7e0a0b0e4f2f |
unverified |
no licence file found · pointer only |
| Garbage Attention in Large Language Models: <BOS> Sink Heads and Sink-aware Pruning added by Syntology |
2026-01 (from id) |
CASIA-LMC-Lab/FLAP/lib/prune.py 2d79e7ef9f571f3d |
unverified |
Apache-2.0 (permissive) |
| DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization added by Syntology |
2025-11 (from id) |
CAS-CLab/DartQuant/NPU_DartQuant/fake_quant/data_utils.py 8715ffbc2b65ebe5 |
unverified |
no licence file found · pointer only |
| QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models added by Syntology |
2025-10 (from id) |
SAI-Lab-NYU/QSVD/fake_quant/data_utils.py 8715ffbc2b65ebe5 |
unverified |
Apache-2.0 (permissive) |
| PT$^2$-LLM: Post-Training Ternarization for Large Language Models added by Syntology |
2025-10 (from id) |
XIANGLONGYAN/PT2-LLM/pt2_llm/data.py bc9148e3beb23d8f |
unverified |
Apache-2.0 (permissive) |
| arXiv:2507.23279 |
2025-07 (from id) |
ZunhaiSu/Super-Experts-Profilling/data_utils.py 24b559c86b2c736b |
unverified |
no licence file found · pointer only |
| arXiv:2507.18553 |
2025-07 (from id) |
IST-DASLab/GPTQ-Babai/quantization/data_utils.py 5a96c1f99eeb15d1 |
unverified |
MIT (permissive) |
| arXiv:2507.01299 |
2025-07 (from id) |
alibaba/EfficientAI/masquant/datautils.py 1beda835d86a2700 |
unverified |
Apache-2.0 (permissive) |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/lib/evaldata.py bbb56ec8bf98c7cf |
unverified |
Apache-2.0 (permissive) |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/prune/lib/calidata_random.py 4eadbf13ea0cc98b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2506.09351 |
2025-06 (from id) |
yuchenblah/DIVE/prune/lib/calidata_select_8.py f3b69fb5e58d1980 |
unverified |
Apache-2.0 (permissive) |
| SAFE: Finding Sparse and Flat Minima to Improve Pruning |
7 Jun 2025 |
LOG-postech/safe-torch/language/lib/prune.py 041beaad3085295f |
unverified |
MIT (permissive) |
| arXiv:2506.03781 |
2025-06 (from id) |
OpenGVLab/OmniQuant/datautils.py 406294899f965c2f |
unverified |
MIT (permissive) |
| DLP: Dynamic Layerwise Pruning in Large Language Models |
27 May 2025 |
ironartisan/dlp/lib/prune.py dd6b0a913778761d |
unverified |
Apache-2.0 (permissive) |
| An Empirical Study of Qwen3 Quantization |
4 May 2025 |
efficient-ml/qwen3-quantization/BiLLM/datautils.py 011909315dfdbc2e |
unverified |
Apache-2.0 (permissive) |
| arXiv:2502.14910 |
2025-02 (from id) |
luffy06/EvoP/src/utils/data_utils.py a9e74891f2269718 |
unverified |
Apache-2.0 (permissive) |
| OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting |
23 Jan 2025 |
brotherhappy/ostquant/utils/data_utils.py 443e801197c1a259 |
unverified |
Apache-2.0 (permissive) |
| RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy |
2 Dec 2024 |
aiha-lab/rilq/rilq_utils/data.py 0cfa46cd7236cebe |
unverified |
Apache-2.0 (permissive) |
| Pushing the Limits of Large Language Model Quantization via the Linearity Theorem |
26 Nov 2024 |
goodevening13/aquakv/aquakv/datautils.py e7d4ca6a631d3399 |
unverified |
Apache-2.0 (permissive) |
| AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment |
15 Nov 2024 |
GATECH-EIC/AmoebaLLM/width_shrink/data.py dee8be97b2cf4692 |
ran
|
MIT (permissive) |
| BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments |
31 Oct 2024 |
xinghaow99/BitStack/bitstack/utils/data_utils.py ded9ae110292b62d |
unverified |
no licence file found · pointer only |
| WAGLE: Strategic Weight Attribution for Effective and Modular Unlearning in Large Language Models |
23 Oct 2024 |
OPTML-Group/WAGLE/src/dataset/dataset.py dee8be97b2cf4692 |
ran
|
MIT (permissive) |
| Quamba: A Post-Training Quantization Recipe for Selective State Space Models |
17 Oct 2024 |
enyac-group/quamba/quamba/data_loaders.py f9f3ddd2d74b2929 |
unverified |
licence not identified · pointer only |
| SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression |
12 Oct 2024 |
mohammad-mozaffari/slim/slim/data.py 83ea71aadeabb49a |
unverified |
MIT (permissive) |
| FlatQuant: Flatness Matters for LLM Quantization |
12 Oct 2024 |
ruikangliu/FlatQuant/flatquant/data_utils.py 4dd3c92f12273ec9 |
unverified |
MIT (permissive) |
| DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models |
12 Oct 2024 |
vengdeng/darex/utils/getwx.py ecc04d6edadf3e8b |
unverified |
no licence file found · pointer only |
| Is C4 Dataset Optimal for Pruning? An Investigation of Calibration Data for LLM Pruning |
9 Oct 2024 |
abx393/llm-pruning-calibration-data/lib/data.py 8be36298f96f0805 |
unverified |
Apache-2.0 (permissive) |
| MC-MoE: Mixture Compressor for Mixture-of-Experts LLMs Gains More |
8 Oct 2024 |
Aaronhuang-778/MC-MoE/datautils.py 011909315dfdbc2e |
unverified |
no licence file found · pointer only |
| Two Sparse Matrices are Better than One: Sparsifying Neural Networks with Double Sparse Factorization |
27 Sep 2024 |
usamec/double_sparse/datautils.py 011909315dfdbc2e |
unverified |
Apache-2.0 (permissive) |
| MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models |
26 Sep 2024 |
NVlabs/MaskLLM/eval_llama_ppl.py dee8be97b2cf4692 |
ran
|
no licence file found · pointer only |
| Eigen Attention: Attention in Low-Rank Space for KV Cache Compression |
10 Aug 2024 |
utkarshsaxena1/eigenattn/datautils.py 46e078681a1cd741 |
unverified |
no licence file found · pointer only |
| EfficientQAT: Efficient Quantization-Aware Training for Large Language Models |
10 Jul 2024 |
opengvlab/efficientqat/datautils_block.py 6d02281de8bfe304 |
unverified |
MIT (permissive) |
| LeanQuant: Accurate Large Language Model Quantization with Loss-Error-Aware Grid |
14 Jul 2024 |
LeanModels/LeanQuant/datautils.py 3bbe14523c98ed57 |
unverified |
no licence file found · pointer only |
| Composable Interventions for Language Models |
9 Jul 2024 |
hartvigsen-group/composable-interventions/sparsellm/lib/data.py dee8be97b2cf4692 |
ran
|
no licence file found · pointer only |
| Composable Interventions for Language Models |
9 Jul 2024 |
hartvigsen-group/composable-interventions/sparsellm/lib/datautils.py ce2dbbe16d039626 |
unverified |
no licence file found · pointer only |
| Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization |
21 Jun 2024 |
log-postech/rethinking-llm-pruning/lib/data.py dee8be97b2cf4692 |
ran
|
no licence file found · pointer only |
| Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox |
15 Jun 2024 |
tsingmaoai/mi-optimize/mi_optimize/datasets/data_loader.py 05cb56deff8da528 |
unverified |
licence not identified · pointer only |
| Examining Post-Training Quantization for Mixture-of-Experts: A Benchmark |
12 Jun 2024 |
unites-lab/moe-quantization/dump_mixtral_routing_distribution.py 219f020120bf3e16 |
unverified |
MIT (permissive) |
| ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization |
10 Jun 2024 |
gatech-eic/shiftaddllm/datautils.py 50f2b672b1461e67 |
unverified |
Apache-2.0 (permissive) |
| Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models |
5 Jun 2024 |
pprp/pruner-zero/lib/data.py 07cb1fc598db64e8 |
unverified |
MIT (permissive) |
| Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models |
5 Jun 2024 |
pprp/pruner-zero/lib/gradient_computation.py 3a97b0120e186c09 |
unverified |
MIT (permissive) |
| DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs |
3 Jun 2024 |
Hsu1023/DuQuant/datautils.py 406294899f965c2f |
unverified |
MIT (permissive) |
| MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization |
2 Jun 2024 |
AozhongZhang/MagR/datautils.py 8f69762e4b3ffac1 |
ran
|
MIT (permissive) |
| Sparse Expansion and Neuronal Disentanglement |
24 May 2024 |
shavit-lab/sparse-expansion/utils/datautils.py 380c081883a25dfd |
unverified |
MIT (permissive) |
| SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models |
23 May 2024 |
Aaronhuang-778/SliM-LLM/slim-llm-plus/datautils.py 85420819f7a0df4a |
unverified |
no licence file found · pointer only |
| SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning |
28 Apr 2024 |
optml-group/soul/src/dataset/dataset.py dee8be97b2cf4692 |
ran
|
MIT (permissive) |
| An empirical study of LLaMA3 quantization: from LLMs to MLLMs |
22 Apr 2024 |
macaronlin/llama3-quantization/datautils.py 406294899f965c2f |
unverified |
no licence file found · pointer only |
| AffineQuant: Affine Transformation Quantization for Large Language Models |
19 Mar 2024 |
bytedance/affinequant/datautils.py 406294899f965c2f |
unverified |
Apache-2.0 (permissive) |
| FrameQuant: Flexible Low-Bit Quantization for Transformers |
10 Mar 2024 |
vsingh-group/framequant/datautils.py ab19b8a72ceae217 |
unverified |
no licence file found · pointer only |
| SparseLLM: Towards Global Pruning for Pre-trained Language Models |
28 Feb 2024 |
BaiTheBest/SparseLLM/datautils.py 011909315dfdbc2e |
unverified |
Apache-2.0 (permissive) |
| BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation |
18 Feb 2024 |
linkanonymous/besa/utils/data.py a8a15ad5400c13dc |
unverified |
no licence file found · pointer only |
| OneBit: Towards Extremely Low-bit Large Language Models |
17 Feb 2024 |
xuyuzhuang11/OneBit/evaluation/lm_eval/datautils.py 27d7c058388f491d |
unverified |
MIT (permissive) |
| NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models |
15 Feb 2024 |
lucius-lsr/nuteprune/eval_ppl.py c31f45d1c5adf4cc |
unverified |
no licence file found · pointer only |
| Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference |
14 Feb 2024 |
hdong920/less/src/data_processing.py 5b616f4c20a2e4bb |
unverified |
no licence file found · pointer only |
| BiLLM: Pushing the Limit of Post-Training Quantization for LLMs |
6 Feb 2024 |
aaronhuang-778/billm/datautils.py 011909315dfdbc2e |
unverified |
MIT (permissive) |
| Fast and Effective Weight Update for Pruned Large Language Models |
1 Jan 2024 |
fmfi-compbio/admm-pruning/lib/data.py dee8be97b2cf4692 |
ran
|
MIT (permissive) |
| The LLM Surgeon |
28 Dec 2023 |
qualcomm-ai-research/llm-surgeon/datautils.py 8f59498c44ab3856 |
unverified |
BSD-3-Clause-Clear · pointer only |
| Fluctuation-based Adaptive Structured Pruning for Large Language Models |
19 Dec 2023 |
casia-iva-lab/flap/lib/data.py 2d79e7ef9f571f3d |
unverified |
Apache-2.0 (permissive) |
| Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models |
8 Nov 2023 |
rocktimjyotidas/gblm-pruner/gradient_computation.py cb81c679c455e959 |
unverified |
MIT (permissive) |
| Atom: Low-bit Quantization for Efficient and Accurate LLM Serving |
29 Oct 2023 |
efeslab/atom/model/datautils.py f9f3ddd2d74b2929 |
unverified |
no licence file found · pointer only |
| One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models |
14 Oct 2023 |
talkking/MixGPT/datautils.py cb4c3b607eb16747 |
unverified |
no licence file found · pointer only |
| QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models |
13 Oct 2023 |
ist-daslab/quik/experiments/datautils.py 79092c10d8a25941 |
unverified |
Apache-2.0 (permissive) |
| Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs |
13 Oct 2023 |
zyxxmu/dsnot/lib/prune.py dee8be97b2cf4692 |
ran
|
no licence file found · pointer only |
| QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models |
12 Oct 2023 |
modeltc/qllm/datautils.py 78a5251d494aa3a6 |
unverified |
Apache-2.0 (permissive) |
| Compressing LLMs: The Truth is Rarely Pure and Never Simple |
2 Oct 2023 |
VITA-Group/llm-kick/GPTQ_experiment/lib/data.py dee8be97b2cf4692 |
ran
|
no licence file found · pointer only |
| PB-LLM: Partially Binarized Large Language Models |
29 Sep 2023 |
hahnyuan/binaryllm/gptq_pb/datautils.py 011909315dfdbc2e |
unverified |
MIT (permissive) |
| OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models |
25 Aug 2023 |
opengvlab/omniquant/datautils.py 406294899f965c2f |
unverified |
MIT (permissive) |
| ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation |
28 Jul 2023 |
lianjiatech/belle/models/gptq/datautils.py 58476252d7ce923b |
unverified |
Apache-2.0 (permissive) |
| Llama 2: Open Foundation and Fine-Tuned Chat Models |
18 Jul 2023 |
squeezeailab/squeezellm/squeezellm/datautils.py efc347eff314b865 |
unverified |
MIT (permissive) |
| A Simple and Effective Pruning Approach for Large Language Models |
20 Jun 2023 |
crystaleye42/eval-safety/lib/prune.py 1afe3719e760f32e |
unverified |
MIT (permissive) |
| RPTQ: Reorder-based Post-training Quantization for Large Language Models |
3 Apr 2023 |
hahnyuan/rptq4llm/datautils.py 8b5943596ef2fc63 |
unverified |
MIT (permissive) |
| arXiv:openreview_d3RFDLBw01 |
|
AI2C-Lab/STLA/data_utils.py ed50f961b7ece8a5 |
unverified |
Apache-2.0 (permissive) |
| arXiv:openreview_Im05D8gFFn |
|
mazumder-lab/RobOP/RobOP-ALPS/datautils.py 0a77e3efb0eb0a48 |
unverified |
MIT (permissive) |
| arXiv:openreview_4iupzej9nT |
|
hikvision-research/STEP/step/lib/data.py 3bfd595806d439a3 |
unverified |
Apache-2.0 (permissive) |
| arXiv:aaai_28960 |
|
CASIA-IVA-Lab/FLAP/lib/data.py 2d79e7ef9f571f3d |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.naacl-long.237 |
|
nbasyl/LLM-FP4/datautils.py 688bbc6963a19c01 |
unverified |
MIT (permissive) |
| arXiv:2025.findings-emnlp.1054 |
|
IST-DASLab/sparsegpt/datautils.py 011909315dfdbc2e |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.acl-long.590 |
|
ChnQ/SPIN/src/eval_ppl.py 75b2f62d7ea8f10e |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.acl-long.498 |
|
OpenGVLab/EfficientQAT/datautils_block.py 6d02281de8bfe304 |
unverified |
MIT (permissive) |
| arXiv:2024.findings-naacl.145 |
|
LianjiaTech/BELLE/models/gptq/datautils.py 58476252d7ce923b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2024.findings-emnlp.579 |
|
RazvanDu/DynamicSlicing/src/slicegpt/data.lib.py dee8be97b2cf4692 |
ran
|
MIT (permissive) |
| arXiv:2023.emnlp-main.892 |
|
SamsungLabs/Z-Fold/datautils.py f011af2b47206903 |
unverified |
MIT (permissive) |