{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2601-20745","title":"HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs","arxiv_id":"2601.20745","date":"2026-01-28","proceeding":null,"authors":["Guoan Wang","Feiyu Wang","Zongwei Lv","Yikun Zong","Tong Yang"],"abstract":"As large language models (LLMs) continue to scale, deployment is increasingly bottlenecked by the memory wall, motivating a shift toward extremely low-bit quantization. However, most quantization-aware training (QAT) methods apply hard rounding and the straight-through estimator (STE) from the beginning of the training, which prematurely discretizes the optimization landscape and induces persistent gradient mismatch between latent weights and quantized weights, hindering effective optimization of quantized models. To address this, we propose Hestia, a Hessian-guided differentiable QAT framework for extremely low-bit LLMs, which replaces the rigid step function with a temperature-controlled softmax relaxation to maintain gradient flow early in training while progressively hardening quantization. Furthermore, Hestia leverages a tensor-wise Hessian trace metric as a lightweight curvature signal to drive fine-grained temperature annealing, enabling sensitivity-aware discretization across the model. Evaluations on Llama-3.2 show that Hestia consistently outperforms existing ternary QAT baselines, yielding average zero-shot improvements of 5.39% and 4.34% for the 1B and 3B models. These results indicate that Hessian-guided relaxation effectively recovers representational capacity, establishing a more robust training path for 1.58-bit LLMs. The code is available at https://github.com/hestia2026/Hestia.","url_abs":"https://arxiv.org/abs/2601.20745","url_pdf":"https://arxiv.org/pdf/2601.20745","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2601.20745","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2601.20745"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/hestia2026/Hestia","reach":null}],"summary":{"ran":3,"ran_fixture":1,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":6,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"0985d54dc263cdc6","entry":"OptimalTempInitializer","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0985d54dc263cdc6"}},{"code_sha256_prefix":"11f0790ae865ce9d","entry":"ThermoQuantizer","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"11f0790ae865ce9d"}},{"code_sha256_prefix":"39c4c0ef1a432d07","entry":"ThermoScheduler","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"39c4c0ef1a432d07"}},{"code_sha256_prefix":"a3ed4b531711340b","entry":"_reshape_for_grouping","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a3ed4b531711340b"}},{"code_sha256_prefix":"02f9a3a520ff0215","entry":"HestiaLinear","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"02f9a3a520ff0215"}},{"code_sha256_prefix":"20169ba49596b63d","entry":"QuantConfig","repo":"hestia2026/Hestia","repo_kind":"found_in_text","path":"src/hestia/quant_linear.py","file_url":"https://github.com/hestia2026/Hestia/blob/HEAD/src/hestia/quant_linear.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"20169ba49596b63d"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}