{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/biased-importance-sampling-for-deep-neural","title":"Biased Importance Sampling for Deep Neural Network Training","arxiv_id":"1706.00043","date":"2017-05-31","proceeding":null,"authors":["Angelos Katharopoulos","François Fleuret"],"abstract":"Importance sampling has been successfully used to accelerate stochastic\noptimization in many convex problems. However, the lack of an efficient way to\ncalculate the importance still hinders its application to Deep Learning.\n  In this paper, we show that the loss value can be used as an alternative\nimportance metric, and propose a way to efficiently approximate it for a deep\nmodel, using a small model trained for that purpose in parallel.\n  This method allows in particular to utilize a biased gradient estimate that\nimplicitly optimizes a soft max-loss, and leads to better generalization\nperformance. While such method suffers from a prohibitively high variance of\nthe gradient estimate when using a standard stochastic optimizer, we show that\nwhen it is combined with our sampling mechanism, it results in a reliable\nprocedure.\n  We showcase the generality of our method by testing it on both image\nclassification and language modeling tasks using deep convolutional and\nrecurrent neural networks. In particular, our method results in 30% faster\ntraining of a CNN for CIFAR10 than when using uniform sampling.","url_abs":"http://arxiv.org/abs/1706.00043v2","url_pdf":"http://arxiv.org/pdf/1706.00043v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"biased-importance-sampling-for-deep-neural","repo_url":"https://github.com/idiap/importance-sampling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.00043","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}