{"url":"/method/cpm-2","slug":"cpm-2","name":"CPM-2","full_name":"CPM-2","full_name_withheld":false,"description_markdown":"**CPM-2** is a 11 billion parameters pre-trained language model based on a standard Transformer architecture consisting of a bidirectional encoder and a unidirectional decoder. The model is pre-trained on WuDaoCorpus which contains 2.3TB cleaned Chinese data as well as 300GB cleaned English data. The pre-training process of CPM-2 can be divided into three stages: Chinese pre-training, bilingual pre-training, and MoE pre-training. Multi-stage training with knowledge inheritance can significantly reduce the computation cost.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2106.10715v3","title":"CPM-2: Large-scale Cost-effective Pre-trained Language Models","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/cpm-2-large-scale-cost-effective-pre-trained","title":"CPM-2: Large-scale Cost-effective Pre-trained Language Models","date":"2021-06-20","arxiv_id":"2106.10715","n_code_links":2,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/decoder","name":"Decoder","papers":1},{"task":null,"name":"GPU","papers":1}],"tasks_shown":2,"n_tasks":2,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cpm-2"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}