Browse State-of-the-Art › Language Modelling

Language Modelling

7,012 papers with code · 55 benchmarks · 166 datasets archive 2025-07-28

MedicalMiscellaneousNatural Language Processing

A language model is a model of natural language. Language models are useful for a variety of tasks, including speech recognition, machine translation, natural language generation (generating more human-like text), optical character recognition, route optimization, handwriting recognition, grammar induction, and information retrieval.

Large language models (LLMs), currently their most advanced form, are predominantly based on transformers trained on larger datasets (frequently using words scraped from the public internet). They have superseded recurrent neural network-based models, which had previously superseded the purely statistical models, such as word n-gram language model.

Source: Wikipedia

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

55 leaderboard tables shown for this task, 55 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 55 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
WikiText-103 (89 rows) RETRO (7.5B) Improving language models by retrieving from trillions of tokens code Syntology ran 16 of 23 samples · 7 unverified Compare
Penn Treebank (Word Level) (43 rows) GPT-3 (Zero-Shot) Language Models are Few-Shot Learners code Syntology ran 15 of 65 samples · 50 unverified Compare
enwik8 (42 rows) GPT-2 (48 layers, h=1600) Language Models are Unsupervised Multitask Learners code — Compare
The Pile (39 rows) Test-Time Fine-Tuning with SIFT + Llama-3.2 (3B) Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs code Syntology ran 2 of 2 samples · 0 unverified Compare
WikiText-2 (38 rows) SparseGPT (175B, 50% Sparsity) SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot code Syntology ran 2 of 12 samples · 10 unverified Compare
LAMBADA (37 rows) PaLM-540B (Few-Shot) PaLM: Scaling Language Modeling with Pathways code Syntology ran 30 of 37 samples · 7 unverified Compare
One Billion Word (27 rows) MDLM (AR baseline) Simple and Effective Masked Diffusion Language Models code Syntology ran 1 of 1 samples · 0 unverified Compare
Text8 (24 rows) GPT-2 Language Models are Unsupervised Multitask Learners code — Compare
Penn Treebank (Character Level) (20 rows) Mogrifier LSTM + dynamic eval Mogrifier LSTM code — Compare
Hutter Prize (18 rows) Transformer-XL + RMS dynamic eval Dynamic Evaluation of Transformer Language Models code — Compare
OpenWebText (12 rows) MDLM-Prime Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking — — Compare
SALMon (10 rows) Spirit-LM (Expr.) Spirit LM: Interleaved Spoken and Written Language Model code Syntology ran 1 of 1 samples · 0 unverified Compare
C4 (9 rows) Primer Primer: Searching for Efficient Transformers for Language Modeling code Syntology ran 3 of 3 samples · 0 unverified Compare
BIG-bench-lite (3 rows) GLM-130B (3-shot) GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
Wiki-40B (3 rows) FLASH-Quad-8k Transformer Quality in Linear Time code — Compare
CLUE (AFQMC) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (C3) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (CMNLI) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (CMRC2018) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (DRCD) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (OCNLI_50K) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
CLUE (WSC1.1) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
FewCLUE (BUSTM) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
FewCLUE (CHID-FC) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
FewCLUE (CLUEWSC-FC) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
FewCLUE (EPRSTMT) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
FewCLUE (OCNLI-FC) (2 rows) GLM-130B GLM-130B: An Open Bilingual Pre-trained Model code Syntology ran 5 of 21 samples · 16 unverified Compare
VietMed (2 rows) Hybrid 4-gram VietMed-Train + ExtraText VietMed: A Dataset and Benchmark for Automatic Speech Recognition... code — Compare
100 sleep nights of 8 caregivers (1 row) Gpt3 007: Democratically Finding The Cause of Packet Drops code — Compare
2000 HUB5 English (1 row) MMLU Spirit LM: Interleaved Spoken and Written Language Model code Syntology ran 1 of 1 samples · 0 unverified Compare
A1 (1 row) Fdyf A Neural Algorithm of Artistic Style code Syntology ran 42 of 110 samples · 68 unverified Compare
Arxiv HEP-TH citation graph (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Bookcorpus2 (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Books3 (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Curation Corpus (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
DM Mathematics (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
enwik8 dev (1 row) Transformer-LS (small) Long-Short Transformer: Efficient Transformers for Language and Vision code Syntology ran 4 of 4 samples · 0 unverified Compare
enwiki8 (1 row) PAR Transformer 24B Pay Attention when Required code — Compare
FreeLaw (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
GitHub (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Gutenberg PG-19 (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
HackerNews (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
language-modeling-recommendation (1 row) GPT2 Zero-Shot Recommendation as Language Modeling code — Compare
NIH ExPorter (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
OpenSubtitles (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
OpenWebtext2 (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
PhilPapers (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Pile CC (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
PTB Diagnostic ECG Database (1 row) I-DARTS Improved Differentiable Architecture Search for Language Modeling... code — Compare
PubMed Cognitive Control Abstracts (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
PubMed Central (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
StackExchange (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
Text8 dev (1 row) Transformer-LS (small) Long-Short Transformer: Efficient Transformers for Language and Vision code Syntology ran 4 of 4 samples · 0 unverified Compare
Ubuntu IRC (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare
USPTO Backgrounds (1 row) Gopher Scaling Language Models: Methods, Analysis & Insights from Training Gopher code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

166 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 166 until expanded.

IMDb Movie ReviewsPubmedWikiText-2Penn TreebankC4WikiText-103The PileBIG-benchBookCorpusLAMBADAPubMedQALAMAOpenSubtitlesOpenWebTextWiCClothoAISHELL-1Libri-LightXQuADHate SpeechLRAPAWS-XATOMICELI5OLIDPG-19Billion Word BenchmarkBLUEBLiMPSIQAWritingPromptsT-RExCC100CLUEE2EKP20kTweetEvalMetaQARealNewsSemantic ScholarCMRCOSCARTED-LIUMSciDocsC3JerichoemrQACMRC 2018DARTMLSUMEmotionLinesOCNLIRWCPeerReadSMM4HChIDTatoebaAdvGLUEArxiv HEP-TH citation graph2000 HUB5 EnglishWorldtreeWiki-40BMassiveTextMuseDataKPTimesSentiCapKELMNatural StoriesDo-Not-AnswerPTB Diagnostic ECG DatabaseText8CLOTHSenseval-2Taskmaster-1CASIA-HWDBBeerAdvocateCMU DoGHumicroeditBabyLMCoDrawDakshinaIndoNLU BenchmarkOpen-PlatypusOVAD benchmarkANTIQUEPersonalDialogFewCLUEHutter PrizeHousekeepCC-StoriesART DatasetComQAIndoSumSALMonPEYMATUT Sound Events 2017UKPWinogender SchemasWMT 2018 NewsDefinite Pronoun Resolution DatasetGINCTUT-SED Synthetic 2016arXiv-10CLUECorpus2020CoarseWSD-20MovieFIBWikiCREMWNLaMProcaWaCCoached Conversational Preference ElicitationLiputan6Tencent ML-ImagesCCPE-MCommitChronicleDatabricks Dolly 15kHLA-Chatirc-disentanglementRONECSLINGVietMedChrEnHPLT v2IndicNLP CorpusKMIRSLNETWikiConvertWikiText-TL-39CircaComparative Question CompletionCzech restaurant informationNQuADPubMed Cognitive Control AbstractsRTCS-TESTSARTSentimental LIARSMC Text CorpusSpadesStack ExchangeViText2SQLAlexa Point of ViewblbooksBooks3Curation CorpusDpgMedia2019Fake News Filipino DatasetFreeLawGlot500-cGlotSparseGlotStoryBookHALvestHiXSTestInstructOpenWikiKitelanguage-modeling-recommendationLenta Short SentencesLipogram-eNTextPentachromatic Cultural Palette DatasetPerformance Improving Code Edits (PIE)PhilPapersPolyNewsSGXSTestSVLDUSPTO BackgroundsVerified Smart Contracts

Subtasks archive 2025-07-28

6 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 7,012 papers with code (17,610 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 26 Aug 2015 284 repositories listed Syntology ran 42 of 110 samples · 68 unverified · 39 pointer-only (licence)
    In fine art, especially painting, humans have mastered the skill to create unique visual experiences through composing a complex interplay between the content and style of an image.
  • 4 Nov 2015 161 repositories listed
    In our experiments, we find that long short term memory recurrent networks after being pretrained with the two approaches are more stable and generalize better.
  • 17 Jun 2021 74 repositories listed Syntology ran 34 of 84 samples · 50 unverified · 28 pointer-only (licence)
    We propose Low-Rank Adaptation, or LoRA, which freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of…
  • 28 May 2020 67 repositories listed Syntology ran 15 of 65 samples · 50 unverified · 4 pointer-only (licence)
    By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do.
  • 26 Jul 2019 67 repositories listed Syntology ran 22 of 48 samples · 26 unverified · 23 pointer-only (licence)
    Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging.
  • 18 Jan 2018 66 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 3 pointer-only (licence)
    Inductive transfer learning has greatly impacted computer vision, but existing approaches in NLP still require task-specific modifications and training from scratch.
  • 24 Jun 2018 59 repositories listed Syntology ran 66 of 156 samples · 90 unverified · 48 pointer-only (licence)
    This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner.
  • 4 Aug 2013 59 repositories listed Syntology ran 7 of 37 samples · 30 unverified · 4 pointer-only (licence)
    This paper shows how Long Short-term Memory recurrent neural networks can be used to generate complex sequences with long-range structure, simply by predicting one data point at a time.
  • 15 Feb 2018 46 repositories listed Syntology ran 23 of 58 samples · 35 unverified · 25 pointer-only (licence)
    We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.
  • 7 Aug 2017 45 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)
    Recurrent neural networks (RNNs), such as long short-term memory networks (LSTMs), serve as a fundamental building block for many sequence learning tasks, including machine translation, language modeling, and question…
  • 31 Mar 2015 44 repositories listed Syntology ran 2 of 15 samples · 13 unverified · 5 pointer-only (licence)
    For the former our approach is competitive with Memory Networks, but with less supervision.
  • 23 Aug 2019 40 repositories listed Syntology ran 6 of 32 samples · 26 unverified
    Recent developments in natural language representations have been accompanied by large and expensive models that leverage vast amounts of general-domain text through self-supervised pre-training.
  • 5 Aug 2015 40 repositories listed Syntology ran 12 of 51 samples · 39 unverified · 9 pointer-only (licence)
    Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly.
  • 2 Oct 2019 37 repositories listed Syntology ran 19 of 27 samples · 8 unverified
    As Transfer Learning from large-scale pre-trained models becomes more prevalent in Natural Language Processing (NLP), operating these large models in on-the-edge and/or under constrained computational training or…
  • 9 Jan 2019 37 repositories listed Syntology ran 63 of 143 samples · 80 unverified · 43 pointer-only (licence)
    Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.
  • 1 Dec 2023 35 repositories listed Syntology ran 18 of 62 samples · 44 unverified · 28 pointer-only (licence)
    Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.
  • 5 Nov 2019 35 repositories listed Syntology ran 25 of 59 samples · 34 unverified · 52 pointer-only (licence)
    We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and…
  • 4 Mar 2018 35 repositories listed Syntology ran 2 of 10 samples · 8 unverified · 1 pointer-only (licence)
    Our results indicate that a simple convolutional architecture outperforms canonical recurrent networks such as LSTMs across a diverse range of tasks and datasets, while demonstrating longer effective memory.
  • 18 Apr 2019 30 repositories listed Syntology ran 1 of 18 samples · 17 unverified
    On LibriSpeech, we achieve 6.
  • 29 May 2023 29 repositories listed Syntology ran 6 of 31 samples · 25 unverified · 2 pointer-only (licence)
    Existing methods for gaining such steerability collect human labels of the relative quality of model generations and fine-tune the unsupervised LM to align with these preferences, often with reinforcement learning from…
  • 9 Feb 2018 28 repositories listed Syntology ran 18 of 38 samples · 20 unverified · 25 pointer-only (licence)
    The controller is trained with policy gradient to select a subgraph that maximizes the expected reward on the validation set.
  • 19 Jun 2019 27 repositories listed Syntology ran 10 of 24 samples · 14 unverified · 3 pointer-only (licence)
    With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling.
  • 13 Jun 2016 26 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 6 pointer-only (licence)
    Our algorithm improves one-shot accuracy on ImageNet from 87.
  • 16 May 2020 25 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 2 pointer-only (licence)
    Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs).
  • 31 Dec 2020 22 repositories listed Syntology ran 1 of 1 samples · 0 unverified
    Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models.
  • 10 Apr 2020 22 repositories listed Syntology ran 15 of 35 samples · 20 unverified · 5 pointer-only (licence)
    To address this limitation, we introduce the Longformer with an attention mechanism that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer.
  • 8 Aug 2019 21 repositories listed Syntology ran 4 of 13 samples · 9 unverified
    The learning rate warmup heuristic achieves remarkable success in stabilizing training, accelerating convergence and improving generalization for adaptive stochastic optimization algorithms like RMSprop and Adam.
  • 14 Feb 2019 21 repositories listed
    Natural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
  • 8 Sep 2014 21 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)
    We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units.
  • 23 May 2023 20 repositories listed Syntology ran 17 of 26 samples · 9 unverified · 17 pointer-only (licence)
    Our best model family, which we name Guanaco, outperforms all previous openly released models on the Vicuna benchmark, reaching 99.

Syntology lines on 28 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections