{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pushing-the-bounds-of-dropout","title":"Pushing the bounds of dropout","arxiv_id":"1805.09208","date":"2018-05-23","proceeding":"ICLR 2019 5","authors":["Gábor Melis","Charles Blundell","Tomáš Kočiský","Karl Moritz Hermann","Chris Dyer","Phil Blunsom"],"abstract":"We show that dropout training is best understood as performing MAP estimation\nconcurrently for a family of conditional models whose objectives are themselves\nlower bounded by the original dropout objective. This discovery allows us to\npick any model from this family after training, which leads to a substantial\nimprovement on regularisation-heavy language modelling. The family includes\nmodels that compute a power mean over the sampled dropout masks, and their less\nstochastic subvariants with tighter and higher lower bounds than the fully\nstochastic dropout objective. We argue that since the deterministic\nsubvariant's bound is equal to its objective, and the highest amongst these\nmodels, the predominant view of it as a good approximation to MC averaging is\nmisleading. Rather, deterministic dropout is the best available approximation\nto the true objective.","url_abs":"http://arxiv.org/abs/1805.09208v2","url_pdf":"http://arxiv.org/pdf/1805.09208v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pushing-the-bounds-of-dropout","repo_url":"https://github.com/deepmind/lamb","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[{"method_slug":"dropout","method_name":"Dropout"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-modelling-on-penn-treebank-word","task":"Language Modelling","dataset":"Penn Treebank (Word Level)","model":"2-layer skip-LSTM + dropout tuning","rank_in_archive_order":24,"of":43,"metrics":{"Params":"24M","Test perplexity":"55.3","Validation perplexity":"57.1"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.09208","atlas_url":"https://app.syntology.ai/?focus=1805.09208","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}