{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unilmv2-pseudo-masked-language-models-for","title":"UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training","arxiv_id":"2002.12804","date":"2020-02-28","proceeding":null,"authors":["Hangbo Bao","Li Dong","Furu Wei","Wenhui Wang","Nan Yang","Xiaodong Liu","Yu Wang","Songhao Piao","Jianfeng Gao","Ming Zhou","Hsiao-Wuen Hon"],"abstract":"We propose to pre-train a unified language model for both autoencoding and partially autoregressive language modeling tasks using a novel training procedure, referred to as a pseudo-masked language model (PMLM). Given an input text with masked tokens, we rely on conventional masks to learn inter-relations between corrupted tokens and context via autoencoding, and pseudo masks to learn intra-relations between masked spans via partially autoregressive modeling. With well-designed position embeddings and self-attention masks, the context encodings are reused to avoid redundant computation. Moreover, conventional masks used for autoencoding provide global masking information, so that all the position embeddings are accessible in partially autoregressive language modeling. In addition, the two tasks pre-train a unified language model as a bidirectional encoder and a sequence-to-sequence decoder, respectively. Our experiments show that the unified language models pre-trained using PMLM achieve new state-of-the-art results on a wide range of natural language understanding and generation tasks across several widely used benchmarks.","url_abs":"https://arxiv.org/abs/2002.12804v1","url_pdf":"https://arxiv.org/pdf/2002.12804v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unilmv2-pseudo-masked-language-models-for","repo_url":"https://github.com/microsoft/unilm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"unilmv2-pseudo-masked-language-models-for","repo_url":"https://github.com/facebookresearch/data2vec_vision","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"unilmv2-pseudo-masked-language-models-for","repo_url":"https://github.com/microsoft/dialoglm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"abstractive-text-summarization","task_name":"Abstractive Text Summarization"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":null,"task_name":"Position"},{"task_slug":"question-generation","task_name":"Question Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/abstractive-text-summarization-on-cnn-daily","task":"Abstractive Text Summarization","dataset":"CNN / Daily Mail","model":"UniLMv2","rank_in_archive_order":24,"of":53,"metrics":{"ROUGE-1":"43.16","ROUGE-2":"20.42","ROUGE-L":"40.14"},"uses_additional_data":true},{"leaderboard":"/sota/question-generation-on-squad11","task":"Question Generation","dataset":"SQuAD1.1","model":"UniLMv2","rank_in_archive_order":4,"of":13,"metrics":{"BLEU-4":"24.43"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2002.12804","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2002.12804"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/dialoglm","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/data2vec_vision","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/unilm","reach":null}],"summary":{"ran_fixture":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8b9e23cfcb2e7648","entry":"repeat_kv","repo":"microsoft/unilm","repo_kind":"official","path":"Diff-Transformer/multihead_attention.py","file_url":"https://github.com/microsoft/unilm/blob/HEAD/Diff-Transformer/multihead_attention.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8b9e23cfcb2e7648"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}