Papers › Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of...

Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of LPIRC-II)

19 Dec 2018arXiv:1812.07872archive 2025-07-28

Alexander Goncharenko, Andrey Denisov, Sergey Alyamkin, Evgeny Terentev

Neural network quantization procedure is the necessary step for porting of neural networks to mobile devices. Quantization allows accelerating the inference, reducing memory consumption and model size. It can be performed without fine-tuning using calibration procedure (calculation of parameters necessary for quantization), or it is possible to train the network with quantization from scratch. Training with quantization from scratch on the labeled data is rather long and resource-consuming procedure. Quantization of network without fine-tuning leads to accuracy drop because of outliers which appear during the calibration. In this article we suggest to simplify the quantization procedure significantly by introducing the trained scale factors for quantization thresholds. It allows speeding up the process of quantization with fine-tuning up to 8 epochs as well as reducing the requirements to the set of train images. By our knowledge, the proposed method allowed us to get the first public available quantized version of MNAS without significant accuracy reduction - 74.8% vs 75.3% for original full-precision network. Model and code are ready for use and available at: https://github.com/agoncharenko1992/FAT-fast_adjustable_threshold.

PaperPDFCode

Code

NervanaSystems/distiller officialmentioned in papermentioned on GitHubpytorchnot reachable when probed 2026-09-17 — repositories for recent papers often appear after camera-ready report
agoncharenko1992/FAT-fast-adjustable-threshold officialmentioned in papermentioned on GitHubtf report
agoncharenko1992/FAT-fast_adjustable_threshold officialmentioned in papermentioned on GitHubtf report
maheshkaran/Nervana-Distiller mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Quantization

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections