{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/varmae-pre-training-of-variational-masked","title":"VarMAE: Pre-training of Variational Masked Autoencoder for Domain-adaptive Language Understanding","arxiv_id":"2211.00430","date":"2022-11-01","proceeding":null,"authors":["Dou Hu","Xiaolong Hou","Xiyang Du","Mengyuan Zhou","Lianxin Jiang","Yang Mo","Xiaofeng Shi"],"abstract":"Pre-trained language models have achieved promising performance on general benchmarks, but underperform when migrated to a specific domain. Recent works perform pre-training from scratch or continual pre-training on domain corpora. However, in many specific domains, the limited corpus can hardly support obtaining precise representations. To address this issue, we propose a novel Transformer-based language model named VarMAE for domain-adaptive language understanding. Under the masked autoencoding objective, we design a context uncertainty learning module to encode the token's context into a smooth latent distribution. The module can produce diverse and well-formed contextual representations. Experiments on science- and finance-domain NLU tasks demonstrate that VarMAE can be efficiently adapted to new domains with limited resources.","url_abs":"https://arxiv.org/abs/2211.00430v1","url_pdf":"https://arxiv.org/pdf/2211.00430v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"citation-intent-classification","task_name":"Citation Intent Classification"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[{"slug":"iee","name":"IEE","full_name":""},{"slug":"mtc","name":"MTC","full_name":""},{"slug":"oir","name":"OIR","full_name":""},{"slug":"psm","name":"PSM","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/citation-intent-classification-on-acl-arc","task":"Citation Intent Classification","dataset":"ACL-ARC","model":"VarMAE","rank_in_archive_order":8,"of":8,"metrics":{"Macro-F1":"Not reported","Micro-F1":"76.50"},"uses_additional_data":false},{"leaderboard":"/sota/participant-intervention-comparison-outcome","task":"Participant Intervention Comparison Outcome Extraction","dataset":"EBM-NLP","model":"VarMAE","rank_in_archive_order":1,"of":5,"metrics":{"F1":"76.01"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2211.00430","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}