{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/opt-iml-scaling-language-model-instruction","title":"OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization","arxiv_id":"2212.12017","date":"2022-12-22","proceeding":null,"authors":["Srinivasan Iyer","Xi Victoria Lin","Ramakanth Pasunuru","Todor Mihaylov","Daniel Simig","Ping Yu","Kurt Shuster","Tianlu Wang","Qing Liu","Punit Singh Koura","Xian Li","Brian O'Horo","Gabriel Pereyra","Jeff Wang","Christopher Dewan","Asli Celikyilmaz","Luke Zettlemoyer","Ves Stoyanov"],"abstract":"Recent work has shown that fine-tuning large pre-trained language models on a collection of tasks described via instructions, a.k.a. instruction-tuning, improves their zero and few-shot generalization to unseen tasks. However, there is a limited understanding of the performance trade-offs of different decisions made during the instruction-tuning process. These decisions include the scale and diversity of the instruction-tuning benchmark, different task sampling strategies, fine-tuning with and without demonstrations, training using specialized datasets for reasoning and dialogue, and finally, the fine-tuning objectives themselves. In this paper, we characterize the effect of instruction-tuning decisions on downstream task performance when scaling both model and benchmark sizes. To this end, we create OPT-IML Bench: a large benchmark for Instruction Meta-Learning (IML) of 2000 NLP tasks consolidated into task categories from 8 existing benchmarks, and prepare an evaluation framework to measure three types of model generalizations: to tasks from fully held-out categories, to held-out tasks from seen categories, and to held-out instances from seen tasks. Through the lens of this framework, we first present insights about instruction-tuning decisions as applied to OPT-30B and further exploit these insights to train OPT-IML 30B and 175B, which are instruction-tuned versions of OPT. OPT-IML demonstrates all three generalization abilities at both scales on four different evaluation benchmarks with diverse tasks and input formats -- PromptSource, FLAN, Super-NaturalInstructions, and UnifiedSKG. Not only does it significantly outperform OPT on all benchmarks but is also highly competitive with existing models fine-tuned on each specific benchmark. We release OPT-IML at both scales, together with the OPT-IML Bench evaluation framework.","url_abs":"https://arxiv.org/abs/2212.12017v3","url_pdf":"https://arxiv.org/pdf/2212.12017v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"opt-iml-scaling-language-model-instruction","repo_url":"https://github.com/tanyuqian/cappy","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[{"method_slug":"opt","method_name":"OPT"},{"method_slug":"opt-iml","method_name":"OPT-IML"}],"datasets_introduced":[],"methods_introduced":[{"slug":"opt-iml","name":"OPT-IML","full_name":"OPT-IML"}],"results":[{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT-IML 175B","rank_in_archive_order":26,"of":90,"metrics":{"Accuracy":"84.8%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT-IML 30B","rank_in_archive_order":31,"of":90,"metrics":{"Accuracy":"83.8%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT-IML 1.3B","rank_in_archive_order":63,"of":90,"metrics":{"Accuracy":"66.8%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT 175B","rank_in_archive_order":72,"of":90,"metrics":{"Accuracy":"60.3%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT 30B","rank_in_archive_order":76,"of":90,"metrics":{"Accuracy":"58.1%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"OPT 1.3B","rank_in_archive_order":85,"of":90,"metrics":{"Accuracy":"54.2%"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT-IML 175B","rank_in_archive_order":42,"of":65,"metrics":{"Accuracy":"71.4"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT-IML 30B","rank_in_archive_order":45,"of":65,"metrics":{"Accuracy":"66.9"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT 30B (0-shot)","rank_in_archive_order":49,"of":65,"metrics":{"Accuracy":"64"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT-IML 1.3B (0-shot)","rank_in_archive_order":53,"of":65,"metrics":{"Accuracy":"61.5"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT 1.3B (zero-shot)","rank_in_archive_order":56,"of":65,"metrics":{"Accuracy":"60.5"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-boolq","task":"Question Answering","dataset":"BoolQ","model":"OPT 175B","rank_in_archive_order":58,"of":65,"metrics":{"Accuracy":"60.1"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2212.12017","atlas_url":"https://app.syntology.ai/?focus=2212.12017","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}