{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-to-center-binary-deep-boltzmann-machines","title":"How to Center Binary Deep Boltzmann Machines","arxiv_id":"1311.1354","date":"2013-11-06","proceeding":null,"authors":["Jan Melchior","Asja Fischer","Laurenz Wiskott"],"abstract":"This work analyzes centered binary Restricted Boltzmann Machines (RBMs) and\nbinary Deep Boltzmann Machines (DBMs), where centering is done by subtracting\noffset values from visible and hidden variables. We show analytically that (i)\ncentering results in a different but equivalent parameterization for artificial\nneural networks in general, (ii) the expected performance of centered binary\nRBMs/DBMs is invariant under simultaneous flip of data and offsets, for any\noffset value in the range of zero to one, (iii) centering can be reformulated\nas a different update rule for normal binary RBMs/DBMs, and (iv) using the\nenhanced gradient is equivalent to setting the offset values to the average\nover model and data mean. Furthermore, numerical simulations suggest that (i)\noptimal generative performance is achieved by subtracting mean values from\nvisible as well as hidden variables, (ii) centered RBMs/DBMs reach\nsignificantly higher log-likelihood values than normal binary RBMs/DBMs, (iii)\ncentering variants whose offsets depend on the model mean, like the enhanced\ngradient, suffer from severe divergence problems, (iv) learning is stabilized\nif an exponentially moving average over the batch means is used for the offset\nvalues instead of the current batch mean, which also prevents the enhanced\ngradient from diverging, (v) centered RBMs/DBMs reach higher LL values than\nnormal RBMs/DBMs while having a smaller norm of the weight matrix, (vi)\ncentering leads to an update direction that is closer to the natural gradient\nand that the natural gradient is extremly efficient for training RBMs, (vii)\ncentering dispense the need for greedy layer-wise pre-training of DBMs, (viii)\nfurthermore we show that pre-training often even worsen the results\nindependently whether centering is used or not, and (ix) centering is also\nbeneficial for auto encoders.","url_abs":"http://arxiv.org/abs/1311.1354v3","url_pdf":"http://arxiv.org/pdf/1311.1354v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-to-center-binary-deep-boltzmann-machines","repo_url":"https://github.com/MelJan/PyDeep","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}