Methods › Natural Language Processing › Language Models › PanGu-α
PanGu-α
Introduced by Wei Zeng et al. in PanGu-α: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
PanGu-α is an autoregressive language model (ALM) with up to 200 billion parameters pretrained on a large corpus of text, mostly in Chinese language. The architecture of PanGu-α is based on Transformer, which has been extensively used as the backbone of a variety of pretrained language models such as BERT and GPT. Different from them, there's an additional query layer developed on top of Transformer layers which aims to explicitly induce the expected output.
Papers archive 2025-07-28
1 shown of 1, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
PanGu-α: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation 26 Apr 2021 · 5 repositories · arXiv:2104.12369
Tasks archive 2025-07-28
20 shown of 21 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections