{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/textbooks-are-all-you-need-ii-phi-1-5","title":"Textbooks Are All You Need II: phi-1.5 technical report","arxiv_id":"2309.05463","date":"2023-09-11","proceeding":null,"authors":["Yuanzhi Li","Sébastien Bubeck","Ronen Eldan","Allie Del Giorno","Suriya Gunasekar","Yin Tat Lee"],"abstract":"We continue the investigation into the power of smaller Transformer-based language models as initiated by \\textbf{TinyStories} -- a 10 million parameter model that can produce coherent English -- and the follow-up work on \\textbf{phi-1}, a 1.3 billion parameter model with Python coding performance close to the state-of-the-art. The latter work proposed to use existing Large Language Models (LLMs) to generate ``textbook quality\" data as a way to enhance the learning process compared to traditional web data. We follow the ``Textbooks Are All You Need\" approach, focusing this time on common sense reasoning in natural language, and create a new 1.3 billion parameter model named \\textbf{phi-1.5}, with performance on natural language tasks comparable to models 5x larger, and surpassing most non-frontier LLMs on more complex reasoning tasks such as grade-school mathematics and basic coding. More generally, \\textbf{phi-1.5} exhibits many of the traits of much larger LLMs, both good -- such as the ability to ``think step by step\" or perform some rudimentary in-context learning -- and bad, including hallucinations and the potential for toxic and biased generations -- encouragingly though, we are seeing improvement on that front thanks to the absence of web data. We open-source \\textbf{phi-1.5} to promote further research on these urgent topics.","url_abs":"https://arxiv.org/abs/2309.05463v1","url_pdf":"https://arxiv.org/pdf/2309.05463v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"textbooks-are-all-you-need-ii-phi-1-5","repo_url":"https://github.com/knowlab/bi-weekly-paper-presentation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"code-generation","task_name":"Code Generation"},{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"},{"task_slug":"multi-task-language-understanding","task_name":"Multi-task Language Understanding"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/code-generation-on-mbpp","task":"Code Generation","dataset":"MBPP","model":"phi-1.5-web 1.3B","rank_in_archive_order":79,"of":99,"metrics":{"Accuracy":"43.5"},"uses_additional_data":false},{"leaderboard":"/sota/common-sense-reasoning-on-arc-challenge","task":"Common Sense Reasoning","dataset":"ARC (Challenge)","model":"phi-1.5-web 1.3B (zero-shot)","rank_in_archive_order":42,"of":54,"metrics":{"Accuracy":"44.9"},"uses_additional_data":false},{"leaderboard":"/sota/common-sense-reasoning-on-arc-easy","task":"Common Sense Reasoning","dataset":"ARC (Easy)","model":"phi-1.5-web 1.3B (0-shot)","rank_in_archive_order":21,"of":47,"metrics":{"Accuracy":"76.1"},"uses_additional_data":false},{"leaderboard":"/sota/common-sense-reasoning-on-winogrande","task":"Common Sense Reasoning","dataset":"WinoGrande","model":"phi-1.5-web 1.3B (zero-shot)","rank_in_archive_order":31,"of":77,"metrics":{"Accuracy":"74.0"},"uses_additional_data":false},{"leaderboard":"/sota/multi-task-language-understanding-on-mmlu","task":"Multi-task Language Understanding","dataset":"MML","model":"phi-1.5-web 1.3B","rank_in_archive_order":36,"of":44,"metrics":{"Average (%)":"37.9"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-piqa","task":"Question Answering","dataset":"PIQA","model":"phi-1.5-web (1.3B)","rank_in_archive_order":42,"of":67,"metrics":{"Accuracy":"77"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-social-iqa","task":"Question Answering","dataset":"SIQA","model":"phi-1.5-web 1.3B (zero-shot)","rank_in_archive_order":16,"of":24,"metrics":{"Accuracy":"53.0"},"uses_additional_data":false},{"leaderboard":"/sota/question-answering-on-social-iqa","task":"Question Answering","dataset":"SIQA","model":"phi-1.5 1.3B (zero-shot)","rank_in_archive_order":17,"of":24,"metrics":{"Accuracy":"52.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2309.05463","atlas_url":"https://app.syntology.ai/?focus=2309.05463","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}