Browse State-of-the-Art › Language Modeling
Language Modeling
5,620 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 5,620 papers with code (14,182 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
4 Nov 2015 161 repositories listedIn our experiments, we find that long short term memory recurrent networks after being pretrained with the two approaches are more stable and generalize better.
-
28 May 2020 67 repositories listed Syntology ran 15 of 65 samples · 50 unverified · 4 pointer-only (licence)By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do.
-
26 Jul 2019 67 repositories listed Syntology ran 22 of 48 samples · 26 unverified · 23 pointer-only (licence)Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging.
-
18 Jan 2018 66 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 3 pointer-only (licence)Inductive transfer learning has greatly impacted computer vision, but existing approaches in NLP still require task-specific modifications and training from scratch.
-
24 Jun 2018 59 repositories listed Syntology ran 66 of 156 samples · 90 unverified · 48 pointer-only (licence)This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner.
-
15 Feb 2018 46 repositories listed Syntology ran 23 of 58 samples · 35 unverified · 25 pointer-only (licence)We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.
-
7 Aug 2017 45 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)Recurrent neural networks (RNNs), such as long short-term memory networks (LSTMs), serve as a fundamental building block for many sequence learning tasks, including machine translation, language modeling, and question…
-
31 Mar 2015 44 repositories listed Syntology ran 2 of 15 samples · 13 unverified · 5 pointer-only (licence)For the former our approach is competitive with Memory Networks, but with less supervision.
-
5 Aug 2015 40 repositories listed Syntology ran 12 of 51 samples · 39 unverified · 9 pointer-only (licence)Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly.
-
2 Oct 2019 37 repositories listed Syntology ran 19 of 27 samples · 8 unverifiedAs Transfer Learning from large-scale pre-trained models becomes more prevalent in Natural Language Processing (NLP), operating these large models in on-the-edge and/or under constrained computational training or…
-
9 Jan 2019 37 repositories listed Syntology ran 63 of 143 samples · 80 unverified · 43 pointer-only (licence)Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.
-
1 Dec 2023 35 repositories listed Syntology ran 18 of 62 samples · 44 unverified · 28 pointer-only (licence)Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.
-
5 Nov 2019 35 repositories listed Syntology ran 25 of 59 samples · 34 unverified · 52 pointer-only (licence)We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and…
-
18 Apr 2019 30 repositories listed Syntology ran 1 of 18 samples · 17 unverifiedOn LibriSpeech, we achieve 6.
-
29 May 2023 29 repositories listed Syntology ran 6 of 31 samples · 25 unverified · 2 pointer-only (licence)Existing methods for gaining such steerability collect human labels of the relative quality of model generations and fine-tune the unsupervised LM to align with these preferences, often with reinforcement learning from…
-
19 Jun 2019 27 repositories listed Syntology ran 10 of 24 samples · 14 unverified · 3 pointer-only (licence)With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling.
-
13 Jun 2016 26 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 6 pointer-only (licence)Our algorithm improves one-shot accuracy on ImageNet from 87.
-
16 May 2020 25 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 2 pointer-only (licence)Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).
-
31 Dec 2020 22 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedRecent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models.
-
10 Apr 2020 22 repositories listed Syntology ran 15 of 35 samples · 20 unverified · 5 pointer-only (licence)To address this limitation, we introduce the Longformer with an attention mechanism that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer.
-
8 Aug 2019 21 repositories listed Syntology ran 4 of 13 samples · 9 unverifiedThe learning rate warmup heuristic achieves remarkable success in stabilizing training, accelerating convergence and improving generalization for adaptive stochastic optimization algorithms like RMSprop and Adam.
-
14 Feb 2019 21 repositories listedNatural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
-
8 Sep 2014 21 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units.
-
2 Jun 2021 20 repositories listed Syntology ran 17 of 26 samples · 9 unverified · 6 pointer-only (licence)In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling.
-
28 Jan 2022 19 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning.
-
23 Mar 2020 19 repositories listed Syntology ran 26 of 40 samples · 14 unverified · 10 pointer-only (licence)Then, instead of training a model that predicts the original identities of the corrupted tokens, we train a discriminative model that predicts whether each token in the corrupted input was replaced by a generator sample…
-
16 Feb 2018 18 repositories listed Syntology ran 2 of 21 samples · 19 unverified · 1 pointer-only (licence)This non-linear probabilistic model enables us to go beyond the limited modeling capacity of linear factor models which still largely dominate collaborative filtering research.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
22 Apr 2019 17 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 11 pointer-only (licence)Despite considerable advancements with deep neural language models, the enigma of neural text degeneration persists when these models are tested as text generators.
-
22 Jan 2019 17 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 1 pointer-only (licence)On unsupervised machine translation, we obtain 34.
Syntology lines on 28 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections