{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/overcoming-catastrophic-forgetting-with-hard","title":"Overcoming catastrophic forgetting with hard attention to the task","arxiv_id":"1801.01423","date":"2018-01-04","proceeding":"ICML 2018 7","authors":["Joan Serrà","Dídac Surís","Marius Miron","Alexandros Karatzoglou"],"abstract":"Catastrophic forgetting occurs when a neural network loses the information\nlearned in a previous task after training on subsequent tasks. This problem\nremains a hurdle for artificial intelligence systems with sequential learning\ncapabilities. In this paper, we propose a task-based hard attention mechanism\nthat preserves previous tasks' information without affecting the current task's\nlearning. A hard attention mask is learned concurrently to every task, through\nstochastic gradient descent, and previous masks are exploited to condition such\nlearning. We show that the proposed mechanism is effective for reducing\ncatastrophic forgetting, cutting current rates by 45 to 80%. We also show that\nit is robust to different hyperparameter choices, and that it offers a number\nof monitoring capabilities. The approach features the possibility to control\nboth the stability and compactness of the learned knowledge, which we believe\nmakes it also attractive for online learning or network compression\napplications.","url_abs":"http://arxiv.org/abs/1801.01423v3","url_pdf":"http://arxiv.org/pdf/1801.01423v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"overcoming-catastrophic-forgetting-with-hard","repo_url":"https://github.com/joansj/hat","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"overcoming-catastrophic-forgetting-with-hard","repo_url":"https://github.com/chilung/hat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"continual-learning","task_name":"Continual Learning"},{"task_slug":"hard-attention","task_name":"Hard Attention"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/continual-learning-on-20newsgroup-10-tasks","task":"Continual Learning","dataset":"20Newsgroup (10 tasks)","model":"HAT","rank_in_archive_order":2,"of":6,"metrics":{"F1 - macro":"0.9521"},"uses_additional_data":false},{"leaderboard":"/sota/continual-learning-on-asc-19-tasks","task":"Continual Learning","dataset":"ASC (19 tasks)","model":"HAT","rank_in_archive_order":7,"of":15,"metrics":{"F1 - macro":"0.7816"},"uses_additional_data":false},{"leaderboard":"/sota/continual-learning-on-dsc-10-tasks","task":"Continual Learning","dataset":"DSC (10 tasks)","model":"HAT","rank_in_archive_order":3,"of":6,"metrics":{"F1 - macro":"0.8614"},"uses_additional_data":false},{"leaderboard":"/sota/continual-learning-on-f-celeba-10-tasks","task":"Continual Learning","dataset":"F-CelebA (10 tasks)","model":"HAT","rank_in_archive_order":6,"of":7,"metrics":{"Acc":"0.5673"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.01423","atlas_url":"https://app.syntology.ai/?focus=1801.01423","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}