Browse State-of-the-Art › Quantization
Quantization
1,596 papers with code · 10 benchmarks · 18 datasets archive 2025-07-28
Quantization is a promising technique to reduce the computation cost of neural network training, which can replace high-cost floating-point numbers (e.g., float32) with low-cost fixed-point numbers (e.g., int8/int16).
Source: Adaptive Precision Training: Quantify Back Propagation in Neural Networks with Fixed-point Numbers
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 1,596 papers with code (4,925 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Dec 2016 44 repositories listed Syntology ran 3 of 18 samples · 15 unverified · 12 pointer-only (licence)We consider the problem of producing compact architectures for text classification, such that the full model fits in a limited amount of memory.
-
20 Jun 2020 25 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 2 pointer-only (licence)We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler.
-
15 Dec 2017 21 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 4 pointer-only (licence)The rising popularity of intelligent mobile devices and the daunting computational cost of deep learning-based models call for efficient and accurate on-device inference schemes.
-
23 May 2023 20 repositories listed Syntology ran 17 of 26 samples · 9 unverified · 17 pointer-only (licence)Our best model family, which we name Guanaco, outperforms all previous openly released models on the Vicuna benchmark, reaching 99.
-
31 Oct 2022 17 repositories listed Syntology ran 5 of 15 samples · 10 unverified · 1 pointer-only (licence)In this paper, we address this challenge, and propose GPTQ, a new one-shot weight quantization method based on approximate second-order information, that is both highly-accurate and highly-efficient.
-
1 Oct 2015 15 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 1 pointer-only (licence)To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that work together to reduce the storage requirement of neural networks by 35x to…
-
28 Feb 2017 14 repositories listed Syntology ran 8 of 35 samples · 27 unverified · 27 pointer-only (licence)Similarity search finds application in specialized database systems handling complex data such as images or videos, which are typically represented by high-dimensional features and require specific indexing structures.
-
20 Jun 2016 13 repositories listed Syntology ran 1 of 13 samples · 12 unverified · 1 pointer-only (licence)We propose DoReFa-Net, a method to train convolutional neural networks that have low bitwidth weights and activations using low bitwidth parameter gradients.
-
1 Jun 2023 12 repositories listed Syntology ran 12 of 18 samples · 6 unverified · 2 pointer-only (licence)We propose Activation-aware Weight Quantization (AWQ), a hardware-friendly approach for LLM low-bit weight-only quantization.
-
21 Nov 2018 11 repositories listedCompared with conventional methods, our framework is fully automated and can specialize the quantization policy for different neural network architectures and hardware architectures.
-
7 Sep 2016 11 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThis paper considers the problem of approximate nearest neighbor search in the compressed domain.
-
5 Oct 2022 9 repositories listed Syntology ran 5 of 21 samples · 16 unverifiedWe introduce GLM-130B, a bilingual (English and Chinese) pre-trained language model with 130 billion parameters.
-
21 Feb 2019 9 repositories listed Syntology ran 7 of 23 samples · 16 unverified · 6 pointer-only (licence)Deep networks run with low precision operations at inference time offer power and space advantages over high precision alternatives, but need to overcome the challenge of maintaining high accuracy as precision decreases.
-
24 Jun 2020 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThis paper presents XLSR which learns cross-lingual speech representations by pretraining a single model from the raw waveform of speech in multiple languages.
-
26 Apr 2018 8 repositories listedSimilarity search approaches based on graph walks have recently attained outstanding speed-accuracy trade-offs, taking aside the memory requirements.
-
7 Sep 2022 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The YOLO community has prospered overwhelmingly to enrich its use in a multitude of hardware platforms and abundant scenarios.
-
5 Jan 2021 7 repositories listedTransformer based models, like BERT and RoBERTa, have achieved state-of-the-art results in many Natural Language Processing tasks.
-
28 Feb 2020 7 repositories listed Syntology ran 5 of 18 samples · 13 unverifiedLinear relaxation based perturbation analysis (LiRPA) for neural networks, which computes provable linear bounds of output neurons given a certain amount of input perturbation, has become a core component in robustness…
-
7 Oct 2019 7 repositories listedThe homogeneous transformation between a LiDAR and monocular camera is required for sensor fusion tasks, such as SLAM.
-
15 Jul 2024 6 repositories listedThis report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models.
-
12 Mar 2024 6 repositories listed Syntology ran 13 of 28 samples · 15 unverifiedWe introduce Chronos, a simple yet effective framework for pretrained probabilistic time series models.
-
10 Mar 2023 6 repositories listedIn this paper, we present a Quantization-error-aware Variable Rate Framework (QVRF) that utilizes a univariate quantization regulator a to achieve wide-range variable rates within a single model.
-
2 Jan 2023 6 repositories listed Syntology ran 2 of 12 samples · 10 unverified · 9 pointer-only (licence)We show for the first time that large-scale generative pretrained transformer (GPT) family models can be pruned to at least 50% sparsity in one-shot, without any retraining, at minimal loss of accuracy.
-
14 Oct 2021 6 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for…
-
2 Aug 2021 6 repositories listedCompared with previous DR models that use brute-force search, JPQ almost matches the best retrieval performance with 30x compression on index size.
-
7 May 2021 6 repositories listedIn this work, we use ResNet as a case study to systematically investigate the effects of quantization on inference compute cost-quality tradeoff curves.
-
31 Mar 2019 6 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 1 pointer-only (licence)It is easy to train and fast to search.
-
4 Dec 2016 6 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 3 pointer-only (licence)To solve this problem, we propose Trained Ternary Quantization (TTQ), a method that can reduce the precision of weights in neural networks to ternary values.
-
18 Nov 2022 5 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedWe propose SmoothQuant, a training-free, accuracy-preserving, and general-purpose post-training quantization (PTQ) solution to enable 8-bit weight, 8-bit activation (W8A8) quantization for LLMs.
-
27 Sep 2020 5 repositories listedTransformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.
Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections