{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/paretoq-scaling-laws-in-extremely-low-bit-llm","title":"ParetoQ: Scaling Laws in Extremely Low-bit LLM Quantization","arxiv_id":"2502.02631","date":"2025-02-04","proceeding":null,"authors":["Zechun Liu","Changsheng Zhao","Hanxian Huang","Sijia Chen","Jing Zhang","Jiawei Zhao","Scott Roy","Lisa Jin","Yunyang Xiong","Yangyang Shi","Lin Xiao","Yuandong Tian","Bilge Soran","Raghuraman Krishnamoorthi","Tijmen Blankevoort","Vikas Chandra"],"abstract":"The optimal bit-width for achieving the best trade-off between quantized model size and accuracy has been a subject of ongoing debate. While some advocate for 4-bit quantization, others propose that 1.58-bit offers superior results. However, the lack of a cohesive framework for different bits has left such conclusions relatively tenuous. We present ParetoQ, the first unified framework that facilitates rigorous comparisons across 1-bit, 1.58-bit, 2-bit, 3-bit, and 4-bit quantization settings. Our findings reveal a notable learning transition between 2 and 3 bits: For 3-bits and above, the fine-tuned models stay close to their original pre-trained distributions, whereas for learning 2-bit networks or below, the representations change drastically. By optimizing training schemes and refining quantization functions, ParetoQ surpasses all previous methods tailored to specific bit widths. Remarkably, our ParetoQ ternary 600M-parameter model even outperforms the previous SoTA ternary 3B-parameter model in accuracy, using only one-fifth of the parameters. Extensive experimentation shows that ternary, 2-bit, and 3-bit quantization maintains comparable performance in the size-accuracy trade-off and generally exceeds 4-bit and binary quantization. Considering hardware constraints, 2-bit quantization offers promising potential for memory reduction and speedup.","url_abs":"https://arxiv.org/abs/2502.02631v1","url_pdf":"https://arxiv.org/pdf/2502.02631v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"paretoq-scaling-laws-in-extremely-low-bit-llm","repo_url":"https://github.com/facebookresearch/LLM-QAT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"quantization","task_name":"Quantization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2502.02631","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2502.02631"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/LLM-QAT","reach":null}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"50a961390acac748","entry":"AsymQuantizer","repo":"facebookresearch/LLM-QAT","repo_kind":"listed","path":"models/utils_quant.py","file_url":"https://github.com/facebookresearch/LLM-QAT/blob/HEAD/models/utils_quant.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"50a961390acac748"}},{"code_sha256_prefix":"1a1205f283561273","entry":"SymQuantizer","repo":"facebookresearch/LLM-QAT","repo_kind":"listed","path":"models/utils_quant.py","file_url":"https://github.com/facebookresearch/LLM-QAT/blob/HEAD/models/utils_quant.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"1a1205f283561273"}},{"code_sha256_prefix":"c6cc2e44490f66e2","entry":"QuantizeLinear","repo":"facebookresearch/LLM-QAT","repo_kind":"listed","path":"models/utils_quant.py","file_url":"https://github.com/facebookresearch/LLM-QAT/blob/HEAD/models/utils_quant.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"c6cc2e44490f66e2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}