{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/teq-trainable-equivalent-transformation-for","title":"TEQ: Trainable Equivalent Transformation for Quantization of LLMs","arxiv_id":"2310.10944","date":"2023-10-17","proceeding":null,"authors":["Wenhua Cheng","Yiyang Cai","Kaokao Lv","Haihao Shen"],"abstract":"As large language models (LLMs) become more prevalent, there is a growing need for new and improved quantization methods that can meet the computationalast layer demands of these modern architectures while maintaining the accuracy. In this paper, we present TEQ, a trainable equivalent transformation that preserves the FP32 precision of the model output while taking advantage of low-precision quantization, especially 3 and 4 bits weight-only quantization. The training process is lightweight, requiring only 1K steps and fewer than 0.1 percent of the original model's trainable parameters. Furthermore, the transformation does not add any computational overhead during inference. Our results are on-par with the state-of-the-art (SOTA) methods on typical LLMs. Our approach can be combined with other methods to achieve even better performance. The code is available at https://github.com/intel/neural-compressor.","url_abs":"https://arxiv.org/abs/2310.10944v1","url_pdf":"https://arxiv.org/pdf/2310.10944v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"teq-trainable-equivalent-transformation-for","repo_url":"https://github.com/intel/neural-compressor","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"quantization","task_name":"Quantization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.10944","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.10944"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/intel/neural-compressor","reach":null}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"084bae2a84aebdd7","entry":"dispatch_model_on_devices","repo":"intel/neural-compressor","repo_kind":"official","path":"examples/pytorch/nlp/huggingface_models/language-modeling/quantization/auto_round/llama3/quantize.py","file_url":"https://github.com/intel/neural-compressor/blob/HEAD/examples/pytorch/nlp/huggingface_models/language-modeling/quantization/auto_round/llama3/quantize.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"084bae2a84aebdd7"}},{"code_sha256_prefix":"7f7d585fdda8fcee","entry":"get_accuracy","repo":"intel/neural-compressor","repo_kind":"official","path":"examples/pytorch/nlp/huggingface_models/language-modeling/quantization/auto_round/llama3/quantize.py","file_url":"https://github.com/intel/neural-compressor/blob/HEAD/examples/pytorch/nlp/huggingface_models/language-modeling/quantization/auto_round/llama3/quantize.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7f7d585fdda8fcee"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}