{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/xgen-7b-technical-report","title":"XGen-7B Technical Report","arxiv_id":"2309.03450","date":"2023-09-07","proceeding":null,"authors":["Erik Nijkamp","Tian Xie","Hiroaki Hayashi","Bo Pang","Congying Xia","Chen Xing","Jesse Vig","Semih Yavuz","Philippe Laban","Ben Krause","Senthil Purushwalkam","Tong Niu","Wojciech Kryściński","Lidiya Murakhovs'ka","Prafulla Kumar Choubey","Alex Fabbri","Ye Liu","Rui Meng","Lifu Tu","Meghana Bhat","Chien-Sheng Wu","Silvio Savarese","Yingbo Zhou","Shafiq Joty","Caiming Xiong"],"abstract":"Large Language Models (LLMs) have become ubiquitous across various domains, transforming the way we interact with information and conduct research. However, most high-performing LLMs remain confined behind proprietary walls, hindering scientific progress. Most open-source LLMs, on the other hand, are limited in their ability to support longer sequence lengths, which is a key requirement for many tasks that require inference over an input context. To address this, we have trained XGen, a series of 7B parameter models on up to 8K sequence length for up to 1.5T tokens. We have also finetuned the XGen models on public-domain instructional data, creating their instruction-tuned counterparts (XGen-Inst). We open-source our models for both research advancements and commercial applications. Our evaluation on standard benchmarks shows that XGen models achieve comparable or better results when compared with state-of-the-art open-source LLMs. Our targeted evaluation on long sequence modeling tasks shows the benefits of our 8K-sequence models over 2K-sequence open-source LLMs.","url_abs":"https://arxiv.org/abs/2309.03450v1","url_pdf":"https://arxiv.org/pdf/2309.03450v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"xgen-7b-technical-report","repo_url":"https://github.com/salesforce/xgen","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"2k","task_name":"2k"},{"task_slug":null,"task_name":"8k"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.03450","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}