{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/better-fine-tuning-by-reducing","title":"Better Fine-Tuning by Reducing Representational Collapse","arxiv_id":"2008.03156","date":"2020-08-06","proceeding":"ICLR 2021 1","authors":["Armen Aghajanyan","Akshat Shrivastava","Anchit Gupta","Naman Goyal","Luke Zettlemoyer","Sonal Gupta"],"abstract":"Although widely adopted, existing approaches for fine-tuning pre-trained language models have been shown to be unstable across hyper-parameter settings, motivating recent work on trust region methods. In this paper, we present a simplified and efficient method rooted in trust region theory that replaces previously used adversarial objectives with parametric noise (sampling from either a normal or uniform distribution), thereby discouraging representation change during fine-tuning when possible without hurting performance. We also introduce a new analysis to motivate the use of trust region methods more generally, by studying representational collapse; the degradation of generalizable representations from pre-trained models as they are fine-tuned for a specific end task. Extensive experiments show that our fine-tuning method matches or exceeds the performance of previous trust region methods on a range of understanding and generation tasks (including DailyMail/CNN, Gigaword, Reddit TIFU, and the GLUE benchmark), while also being much faster. We also show that it is less prone to representation collapse; the pre-trained models maintain more generalizable representations every time they are fine-tuned.","url_abs":"https://arxiv.org/abs/2008.03156v1","url_pdf":"https://arxiv.org/pdf/2008.03156v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"better-fine-tuning-by-reducing","repo_url":"https://github.com/pytorch/fairseq/tree/master/examples/rxf","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"better-fine-tuning-by-reducing","repo_url":"https://github.com/cliang1453/camero","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"better-fine-tuning-by-reducing","repo_url":"https://github.com/cosmoquester/2021-dialogue-summary-competition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"abstractive-text-summarization","task_name":"Abstractive Text Summarization"},{"task_slug":"cross-lingual-natural-language-inference","task_name":"Cross-Lingual Natural Language Inference"},{"task_slug":"text-summarization","task_name":"Text Summarization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/abstractive-text-summarization-on-cnn-daily","task":"Abstractive Text Summarization","dataset":"CNN / Daily Mail","model":"BART+R3F","rank_in_archive_order":14,"of":53,"metrics":{"ROUGE-1":"44.38","ROUGE-2":"21.53","ROUGE-L":"41.17"},"uses_additional_data":false},{"leaderboard":"/sota/cross-lingual-natural-language-inference-on","task":"Cross-Lingual Natural Language Inference","dataset":"XNLI Zero-Shot English-to-French","model":"XLM-R R4F","rank_in_archive_order":1,"of":3,"metrics":{"Accuracy":"84.7%"},"uses_additional_data":false},{"leaderboard":"/sota/cross-lingual-natural-language-inference-on-3","task":"Cross-Lingual Natural Language Inference","dataset":"XNLI Zero-Shot English-to-German","model":"XLM-R R4F","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy":"84.2%"},"uses_additional_data":false},{"leaderboard":"/sota/cross-lingual-natural-language-inference-on-1","task":"Cross-Lingual Natural Language Inference","dataset":"XNLI Zero-Shot English-to-Spanish","model":"XLM-R R4F","rank_in_archive_order":1,"of":4,"metrics":{"Accuracy":"85.2%"},"uses_additional_data":false},{"leaderboard":"/sota/text-summarization-on-gigaword","task":"Text Summarization","dataset":"GigaWord","model":"BART-RXF","rank_in_archive_order":4,"of":41,"metrics":{"ROUGE-1":"40.45","ROUGE-2":"20.69","ROUGE-L":"36.56"},"uses_additional_data":false},{"leaderboard":"/sota/text-summarization-on-reddit-tifu","task":"Text Summarization","dataset":"Reddit TIFU","model":"BART+R3F","rank_in_archive_order":2,"of":5,"metrics":{"ROUGE-1":"30.31","ROUGE-2":"10.98","ROUGE-L":"24.74"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2008.03156","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}