{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/skywork-a-more-open-bilingual-foundation","title":"Skywork: A More Open Bilingual Foundation Model","arxiv_id":"2310.19341","date":"2023-10-30","proceeding":null,"authors":["Tianwen Wei","Liang Zhao","Lichang Zhang","Bo Zhu","Lijie Wang","Haihua Yang","Biye Li","Cheng Cheng","Weiwei Lü","Rui Hu","Chenxia Li","Liu Yang","Xilin Luo","Xuejie Wu","Lunan Liu","Wenjun Cheng","Peng Cheng","Jianhao Zhang","XiaoYu Zhang","Lei Lin","Xiaokun Wang","Yutuan Ma","Chuanhai Dong","Yanqi Sun","Yifu Chen","Yongyi Peng","Xiaojuan Liang","Shuicheng Yan","Han Fang","Yahui Zhou"],"abstract":"In this technical report, we present Skywork-13B, a family of large language models (LLMs) trained on a corpus of over 3.2 trillion tokens drawn from both English and Chinese texts. This bilingual foundation model is the most extensively trained and openly published LLMs of comparable size to date. We introduce a two-stage training methodology using a segmented corpus, targeting general purpose training and then domain-specific enhancement training, respectively. We show that our model not only excels on popular benchmarks, but also achieves \\emph{state of the art} performance in Chinese language modeling on diverse domains. Furthermore, we propose a novel leakage detection method, demonstrating that test data contamination is a pressing issue warranting further investigation by the LLM community. To spur future research, we release Skywork-13B along with checkpoints obtained during intermediate stages of the training process. We are also releasing part of our SkyPile corpus, a collection of over 150 billion tokens of web text, which is the largest high quality open Chinese pre-training corpus to date. We hope Skywork-13B and our open corpus will serve as a valuable open-source resource to democratize access to high-quality LLMs.","url_abs":"https://arxiv.org/abs/2310.19341v1","url_pdf":"https://arxiv.org/pdf/2310.19341v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"skywork-a-more-open-bilingual-foundation","repo_url":"https://github.com/skyworkai/skywork","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.19341","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.19341"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/skyworkai/skywork","reach":null}],"summary":{"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"e4adafd3040ed495","entry":"decode","repo":"skyworkai/skywork","repo_kind":"official","path":"eval/eval_gsm8k.py","file_url":"https://github.com/skyworkai/skywork/blob/HEAD/eval/eval_gsm8k.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e4adafd3040ed495"}},{"code_sha256_prefix":"073b5576f7ab301e","entry":"get_match_str","repo":"skyworkai/skywork","repo_kind":"official","path":"eval/eval_gsm8k.py","file_url":"https://github.com/skyworkai/skywork/blob/HEAD/eval/eval_gsm8k.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"073b5576f7ab301e"}},{"code_sha256_prefix":"48ae121813e44a58","entry":"doc_to_text","repo":"skyworkai/skywork","repo_kind":"official","path":"eval/eval_gsm8k.py","file_url":"https://github.com/skyworkai/skywork/blob/HEAD/eval/eval_gsm8k.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"48ae121813e44a58"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}