{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scaling-granite-code-models-to-128k-context","title":"Scaling Granite Code Models to 128K Context","arxiv_id":"2407.13739","date":"2024-07-18","proceeding":null,"authors":["Matt Stallone","Vaibhav Saxena","Leonid Karlinsky","Bridget McGinn","Tim Bula","Mayank Mishra","Adriana Meza Soria","Gaoyuan Zhang","Aditya Prasad","Yikang Shen","Saptha Surendran","Shanmukha Guttula","Hima Patel","Parameswaran Selvam","Xuan-Hong Dang","Yan Koyfman","Atin Sood","Rogerio Feris","Nirmit Desai","David D. Cox","Ruchir Puri","Rameswar Panda"],"abstract":"This paper introduces long-context Granite code models that support effective context windows of up to 128K tokens. Our solution for scaling context length of Granite 3B/8B code models from 2K/4K to 128K consists of a light-weight continual pretraining by gradually increasing its RoPE base frequency with repository-level file packing and length-upsampled long-context data. Additionally, we also release instruction-tuned models with long-context support which are derived by further finetuning the long context base models on a mix of permissively licensed short and long-context instruction-response pairs. While comparing to the original short-context Granite code models, our long-context models achieve significant improvements on long-context tasks without any noticeable performance degradation on regular code completion benchmarks (e.g., HumanEval). We release all our long-context Granite code models under an Apache 2.0 license for both research and commercial use.","url_abs":"https://arxiv.org/abs/2407.13739v1","url_pdf":"https://arxiv.org/pdf/2407.13739v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scaling-granite-code-models-to-128k-context","repo_url":"https://github.com/ibm/data-prep-kit","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"2k","task_name":"2k"},{"task_slug":"4k","task_name":"4k"},{"task_slug":"code-completion","task_name":"Code Completion"},{"task_slug":"continual-pretraining","task_name":"Continual Pretraining"},{"task_slug":"humaneval","task_name":"HumanEval"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.13739","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}