{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lookahead-optimizer-k-steps-forward-1-step","title":"Lookahead Optimizer: k steps forward, 1 step back","arxiv_id":"1907.08610","date":"2019-07-19","proceeding":"NeurIPS 2019 12","authors":["Michael R. Zhang","James Lucas","Geoffrey Hinton","Jimmy Ba"],"abstract":"The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam, and (2) accelerated schemes, such as heavy-ball and Nesterov momentum. In this paper, we propose a new optimization algorithm, Lookahead, that is orthogonal to these previous approaches and iteratively updates two sets of weights. Intuitively, the algorithm chooses a search direction by looking ahead at the sequence of fast weights generated by another optimizer. We show that Lookahead improves the learning stability and lowers the variance of its inner optimizer with negligible computation and memory cost. We empirically demonstrate Lookahead can significantly improve the performance of SGD and Adam, even with their default hyperparameter settings on ImageNet, CIFAR-10/100, neural machine translation, and Penn Treebank.","url_abs":"https://arxiv.org/abs/1907.08610v2","url_pdf":"https://arxiv.org/pdf/1907.08610v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/michaelrzhang/lookahead","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/201419/Optimizer-PyTorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/Abhimanyu08/Lookahead_Optimizer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/COMP6248-Reproducability-Challenge/LookaheadOptimizer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/DreamInvoker/LookaheadOptimizer-mx-implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mxnet","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/HamadYA/GhostFaceNets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/SdahlSean/RangerOptimizerTensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/allen108108/Model-Optimizer_Implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/alphadl/lookahead.pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/bojone/keras_lookahead","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/chizhu/BDC2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/kpe/params-flow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/mnikitin/LookaheadOptimizer-mx","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mxnet","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/nachiket273/lookahead_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/nsarang/lookahead_keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/rwightman/pytorch-image-models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/wkcn/LookaheadOptimizer-mx","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"mxnet","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/zhangtj1996/lookahead-sgd-adam-rmsprop-","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"lookahead-optimizer-k-steps-forward-1-step","repo_url":"https://github.com/dseuss/pytorch-lookahead-optimizer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"lookahead","method_name":"Lookahead"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/stochastic-optimization-on-cifar-10-resnet-18","task":"Stochastic Optimization","dataset":"CIFAR-10 ResNet-18 - 200 Epochs","model":"Lookahead","rank_in_archive_order":2,"of":4,"metrics":{"Accuracy":"95.27"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-cifar-10-resnet-18","task":"Stochastic Optimization","dataset":"CIFAR-10 ResNet-18 - 200 Epochs","model":"SGD","rank_in_archive_order":3,"of":4,"metrics":{"Accuracy":"95.23"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-cifar-10-resnet-18","task":"Stochastic Optimization","dataset":"CIFAR-10 ResNet-18 - 200 Epochs","model":"ADAM","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy":"94.84"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-imagenet-resnet-50","task":"Stochastic Optimization","dataset":"ImageNet ResNet-50 - 50 Epochs","model":"Lookahead","rank_in_archive_order":1,"of":2,"metrics":{"Top 1 Accuracy":"75.13%"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-imagenet-resnet-50","task":"Stochastic Optimization","dataset":"ImageNet ResNet-50 - 50 Epochs","model":"SGD","rank_in_archive_order":2,"of":2,"metrics":{"Top 5 Accuracy":"92.15%"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-imagenet-resnet-50-1","task":"Stochastic Optimization","dataset":"ImageNet ResNet-50 - 60 Epochs","model":"Lookahead","rank_in_archive_order":1,"of":2,"metrics":{"Top 1 Accuracy":"75.49%","Top 5 Accuracy":"92.53"},"uses_additional_data":false},{"leaderboard":"/sota/stochastic-optimization-on-imagenet-resnet-50-1","task":"Stochastic Optimization","dataset":"ImageNet ResNet-50 - 60 Epochs","model":"SGD","rank_in_archive_order":2,"of":2,"metrics":{"Top 1 Accuracy":"75.15%","Top 5 Accuracy":"92.56"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1907.08610","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}