{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/minimal-random-code-learning-getting-bits","title":"Minimal Random Code Learning: Getting Bits Back from Compressed Model Parameters","arxiv_id":"1810.00440","date":"2018-09-30","proceeding":"ICLR 2019 5","authors":["Marton Havasi","Robert Peharz","José Miguel Hernández-Lobato"],"abstract":"While deep neural networks are a highly successful model class, their large\nmemory footprint puts considerable strain on energy consumption, communication\nbandwidth, and storage requirements. Consequently, model size reduction has\nbecome an utmost goal in deep learning. A typical approach is to train a set of\ndeterministic weights, while applying certain techniques such as pruning and\nquantization, in order that the empirical weight distribution becomes amenable\nto Shannon-style coding schemes. However, as shown in this paper, relaxing\nweight determinism and using a full variational distribution over weights\nallows for more efficient coding schemes and consequently higher compression\nrates. In particular, following the classical bits-back argument, we encode the\nnetwork weights using a random sample, requiring only a number of bits\ncorresponding to the Kullback-Leibler divergence between the sampled\nvariational distribution and the encoding distribution. By imposing a\nconstraint on the Kullback-Leibler divergence, we are able to explicitly\ncontrol the compression rate, while optimizing the expected loss on the\ntraining set. The employed encoding scheme can be shown to be close to the\noptimal information-theoretical lower bound, with respect to the employed\nvariational family. Our method sets new state-of-the-art in neural network\ncompression, as it strictly dominates previous approaches in a Pareto sense: On\nthe benchmarks LeNet-5/MNIST and VGG-16/CIFAR-10, our approach yields the best\ntest performance for a fixed memory budget, and vice versa, it achieves the\nhighest compression rates for a fixed test performance.","url_abs":"http://arxiv.org/abs/1810.00440v1","url_pdf":"http://arxiv.org/pdf/1810.00440v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"minimal-random-code-learning-getting-bits","repo_url":"https://github.com/cambridge-mlg/miracle","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"minimal-random-code-learning-getting-bits","repo_url":"https://github.com/cambridge-mlg/variational-shannon-coding","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"neural-network-compression","task_name":"Neural Network Compression"},{"task_slug":"quantization","task_name":"Quantization"}],"methods":[{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.00440","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1810.00440"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cambridge-mlg/miracle","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cambridge-mlg/variational-shannon-coding","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"2ab005a11f614a7a","entry":"cache","repo":"cambridge-mlg/miracle","repo_kind":"official","path":"models/cache.py","file_url":"https://github.com/cambridge-mlg/miracle/blob/HEAD/models/cache.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2ab005a11f614a7a"}},{"code_sha256_prefix":"68c8a666f927cfee","entry":"one_hot_encoded","repo":"cambridge-mlg/miracle","repo_kind":"official","path":"models/dataset.py","file_url":"https://github.com/cambridge-mlg/miracle/blob/HEAD/models/dataset.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"68c8a666f927cfee"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}