Browse State-of-the-Art › Masked Language Modeling
Masked Language Modeling
254 papers with code · 0 benchmarks · 7 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 254 papers with code (475 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Mar 2020 19 repositories listed Syntology ran 26 of 40 samples · 14 unverified · 10 pointer-only (licence)Then, instead of training a model that predicts the original identities of the corrupted tokens, we train a discriminative model that predicts whether each token in the corrupted input was replaced by a generator sample…
-
20 Aug 2019 9 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 3 pointer-only (licence)In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder.
-
20 Apr 2020 7 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 6 pointer-only (licence)Since BERT neglects dependency among predicted tokens, XLNet introduces permuted language modeling (PLM) for pre-training to address this problem.
-
25 Oct 2019 7 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedThis generalization ability has been attributed to the use of a shared subword vocabulary and joint training across multiple languages giving rise to deep multilingual abstractions.
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
3 Jul 2020 6 repositories listedWhile BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings…
-
10 Feb 2020 6 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedLanguage model pre-training has been shown to capture a surprising amount of world knowledge, crucial for NLP tasks such as question answering.
-
21 Dec 2020 5 repositories listedTransformer is the backbone of modern NLP models.
-
18 Apr 2022 4 repositories listedIn this paper, we propose \textbf{LayoutLMv3} to pre-train multimodal Transformers for Document AI with unified text and image masking.
-
7 Aug 2021 4 repositories listedIn particular, when compared to published models such as conformer-based wav2vec~2.
-
5 Mar 2020 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)We introduce "talking-heads attention" - a variation on multi-head attention which includes linearprojections across the attention-heads dimension, immediately before and after the softmax operation.
-
16 Jun 2022 3 repositories listed Syntology ran 14 of 34 samples · 20 unverified · 1 pointer-only (licence)Manual annotation of question and answers for videos, however, is tedious and prohibits scalability.
-
1 May 2020 3 repositories listed Syntology ran 5 of 13 samples · 8 unverified · 8 pointer-only (licence)We present HERO, a novel framework for large-scale video+language omni-representation learning.
-
17 Mar 2020 3 repositories listedThis allows the coupling of the MLM as pre-training with a downstream anomaly detection task.
-
11 Jun 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)While diffusion models excel at generating high-quality images, prior work reports a significant performance gap between diffusion and autoregressive (AR) methods in language modeling.
-
19 May 2023 2 repositories listed Syntology ran 6 of 16 samples · 10 unverifiedIn this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective.
-
24 Nov 2022 2 repositories listed Syntology ran 6 of 10 samples · 4 unverifiedMedical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information.
-
17 Oct 2022 2 repositories listed Syntology ran 2 of 18 samples · 16 unverifiedPretraining a language model (LM) on text has been shown to help various downstream NLP tasks.
-
11 Oct 2022 2 repositories listed Syntology ran 6 of 17 samples · 11 unverifiedThis paper proposes the Mixture of Attention Heads (MoA), a new architecture that combines multi-head attention with the MoE mechanism.
-
22 Aug 2022 2 repositories listedA big convergence of language, vision, and multimodal pretraining is emerging.
-
21 Jul 2022 2 repositories listedWe find that our proposed pre-training methods help in modeling the data at a patient and population level and improve performance in different fine-tuning tasks on all datasets.
-
20 Jun 2022 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedAccurate prediction of Drug-Target Affinity (DTA) is of vital importance in early-stage drug discovery, facilitating the identification of drugs that can effectively interact with specific targets and regulate their…
-
11 May 2022 2 repositories listedHowever, the CNN feature maps still maintain the spatial relationship and we utilize this property to design self-supervised learning approaches to train the encoder of object detection transformers in pretraining and…
-
29 Apr 2022 2 repositories listed Syntology ran 5 of 7 samples · 2 unverifiedIn this paper, we specifically adapt vision-language joint learning for scene text detection, a task that intrinsically involves cross-modal interaction between the two modalities: vision and language, since text is the…
-
21 Feb 2022 2 repositories listedWe revisit the design choices in Transformers, and propose methods to address their weaknesses in handling long sequences.
-
29 Dec 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Then it diversifies the experts and continues to train the MoE with a novel Dense-to-Sparse gate (DTS-Gate).
-
15 Nov 2021 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present a self-supervised framework iBOT that can perform masked prediction with an online tokenizer.
-
14 Oct 2021 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Both these masks can then be composed with the pretrained model.
-
4 Aug 2021 2 repositories listedTuning pre-trained language models (PLMs) with task-specific prompts has been a promising approach for text classification.
-
3 Jun 2021 2 repositories listed Syntology ran 5 of 10 samples · 5 unverified · 5 pointer-only (licence)Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length.
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections