Browse State-of-the-Art › Text Generation
Text Generation
2,047 papers with code · 20 benchmarks · 162 datasets archive 2025-07-28
Text Generation is the task of generating text with the goal of appearing indistinguishable to human-written text. This task is more formally known as "natural language generation" in the literature.
Text generation can be addressed with Markov processes or deep generative models like LSTMs. Recently, some of the most advanced methods for text generation include BART, GPT and other GAN-based approaches. Text generation systems are evaluated either through human ratings or automatic evaluation metrics like METEOR, ROUGE, and BLEU.
Further readings:
( Image credit: Adversarial Ranking for Language Generation )
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
155 leaderboard tables shown for this task, 20 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 155 until expanded.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DART (7 rows) | T5B Baseline | FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text... | code | — | Compare |
| COCO Captions (5 rows) | LeakGAN | Long Text Generation via Adversarial Training with Leaked Information | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| EMNLP2017 WMT (5 rows) | LeakGAN | Long Text Generation via Adversarial Training with Leaked Information | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| ReDial (5 rows) | UniCRS | Towards Unified Conversational Recommender Systems via... | code | — | Compare |
| CommonGen (4 rows) | UniLM | CommonGen: A Constrained Text Generation Challenge for Generative... | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
| ROCStories (4 rows) | Beam search + A*esque (beam) | NeuroLogic A*esque Decoding: Constrained Text Generation with... | code | — | Compare |
| Chinese Poems (3 rows) | RankGAN | Adversarial Ranking for Language Generation | code | — | Compare |
| Czech restaurant information (3 rows) | TGen++ | The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics | — | — | Compare |
| OpenWebText (3 rows) | GPT2-Hermite | Polynomial, trigonometric, and tropical activations | code | — | Compare |
| SciQ (3 rows) | LLaMA-65B+CFG (zero-shot) | Stay on topic with Classifier-Free Guidance | — | — | Compare |
| Yahoo Questions (3 rows) | Aggressive VAE | Lagging Inference Networks and Posterior Collapse in Variational... | code | Syntology ran 1 of 10 samples · 9 unverified | Compare |
| ADGEN (1 row) | BART (TextBox 2.0) | TextBox 2.0: A Text Generation Library with Pre-trained Language Models | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| CMU-SE (1 row) | STWGAN-GP | Generating Text through Adversarial Training using Skip-Thought Vectors | code | — | Compare |
| CNN/Daily Mail (1 row) | PALM | PALM: Pre-training an Autoencoding&Autoregressive Language Model... | code | — | Compare |
| CSL (1 row) | BART (TextBox 2.0) | TextBox 2.0: A Text Generation Library with Pre-trained Language Models | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| DailyDialog (1 row) | AEM+Attention | An Auto-Encoder Matching Model for Learning Utterance-Level... | code | — | Compare |
| HarmfulQA (1 row) | GPT-4 | Red-Teaming Large Language Models using Chain of Utterances for... | code | — | Compare |
| LCSTS (1 row) | BART (TextBox 2.0) | TextBox 2.0: A Text Generation Library with Pre-trained Language Models | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| LDC2016E25 (1 row) | Graph2Seq | A Graph-to-Sequence Model for AMR-to-Text Generation | code | — | Compare |
| One Billion Word (1 row) | WGANGP + DGflow | Refining Deep Generative Models via Discriminator Gradient Flow | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
| AI2 Reasoning Challenge (25-Shot) (0 rows) | no rows in the archive | — | — | ||
| AI2 Reasoning Challenge TR v0.2 (0 rows) | no rows in the archive | — | — | ||
| AlpacaEval (0 rows) | no rows in the archive | — | — | ||
| AlpacaEval 2.0 (GPT-4-1106-Preview) (0 rows) | no rows in the archive | — | — | ||
| Arabic Poetry Dataset (6th - 21st century) (0 rows) | no rows in the archive | — | — | ||
| arc_challenge (0 rows) | no rows in the archive | — | — | ||
| ARC Challenge (25-Shot) (0 rows) | no rows in the archive | — | — | ||
| ARC-Challenge (PT) (0 rows) | no rows in the archive | — | — | ||
| arc_easy (0 rows) | no rows in the archive | — | — | ||
| Assin2 RTE (0 rows) | no rows in the archive | — | — | ||
| Assin2 STS (0 rows) | no rows in the archive | — | — | ||
| axb (0 rows) | no rows in the archive | — | — | ||
| axg (0 rows) | no rows in the archive | — | — | ||
| BBH (3-Shot) (0 rows) | no rows in the archive | — | — | ||
| Big Bench Hard (3-Shot) (0 rows) | no rows in the archive | — | — | ||
| BLUEX (No Images) (0 rows) | no rows in the archive | — | — | ||
| BoolQ (0 rows) | no rows in the archive | — | — | ||
| CALAME-PT (0 rows) | no rows in the archive | — | — | ||
| cb (0 rows) | no rows in the archive | — | — | ||
| Censorship (0-shot) (0 rows) | no rows in the archive | — | — | ||
| CoLA (0 rows) | no rows in the archive | — | — | ||
| Creativity (0-shot) (0 rows) | no rows in the archive | — | — | ||
| CrimeStats (0 rows) | no rows in the archive | — | — | ||
| crows_pairs_english (0 rows) | no rows in the archive | — | — | ||
| crows_pairs_french (0 rows) | no rows in the archive | — | — | ||
| DiaBLa (0 rows) | no rows in the archive | — | — | ||
| Drop (3-Shot) (0 rows) | no rows in the archive | — | — | ||
| ENEM Challenge (No Images) (0 rows) | no rows in the archive | — | — | ||
| FaQuAD NLI (0 rows) | no rows in the archive | — | — | ||
| GPQA (0-shot) (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_afr (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_asm (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_cat (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ceb (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ces (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ckb (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_cym (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_dan (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_deu (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ell (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_eng (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_est (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_fas (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_fin (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_fra (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ful (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_glg (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_guj (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_hau (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_heb (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_mlt (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_mon (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_msa (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_mya (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_nld (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_nob (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_npi (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_nso (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_nya (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_oci (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_orm (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ory (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_pan (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_pol (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_por (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ron (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_rus (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_slk (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_slv (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_som (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tam (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tel (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tgk (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tgl (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tha (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_tur (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_ukr (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_umb (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_urd (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_uzb (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_vie (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_wol (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_xho (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_yor (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_zho_simpl (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_zho_trad (0 rows) | no rows in the archive | — | — | ||
| gsarti/flores_101_zul (0 rows) | no rows in the archive | — | — | ||
| GSM8k (5-shot) (0 rows) | no rows in the archive | — | — | ||
| GSM8k TR (0 rows) | no rows in the archive | — | — | ||
| GSM8k TR v0.2 (0 rows) | no rows in the archive | — | — | ||
| HateBR Binary (0 rows) | no rows in the archive | — | — | ||
| HeadQA (0 rows) | no rows in the archive | — | — | ||
| HellaSwag (0 rows) | no rows in the archive | — | — | ||
| HellaSwag (10-Shot) (0 rows) | no rows in the archive | — | — | ||
| HellaSwag (PT) (0 rows) | no rows in the archive | — | — | ||
| HellaSwag TR (0 rows) | no rows in the archive | — | — | ||
| Humanness (0-shot) (0 rows) | no rows in the archive | — | — | ||
| IFEval (0-Shot) (0 rows) | no rows in the archive | — | — | ||
| IndicGLUE (0 rows) | no rows in the archive | — | — | ||
| Internet (0 rows) | no rows in the archive | — | — | ||
| LAMBADA-PT (0 rows) | no rows in the archive | — | — | ||
| LogiQA (0 rows) | no rows in the archive | — | — | ||
| Math HARD (4-Shot) (0 rows) | no rows in the archive | — | — | ||
| MATH Lvl 5 (4-Shot) (0 rows) | no rows in the archive | — | — | ||
| MMLU (5-Shot) (0 rows) | no rows in the archive | — | — | ||
| MMLU-PRO (5-shot) (0 rows) | no rows in the archive | — | — | ||
| MMLU TR (0 rows) | no rows in the archive | — | — | ||
| MMLU TR v0.2 (0 rows) | no rows in the archive | — | — | ||
| MNLI (0 rows) | no rows in the archive | — | — | ||
| MT-Bench (0 rows) | no rows in the archive | — | — | ||
| MuSR (0-shot) (0 rows) | no rows in the archive | — | — | ||
| Nexa Scientific Tokens (0 rows) | no rows in the archive | — | — | ||
| OAB Exams (0 rows) | no rows in the archive | — | — | ||
| Open Australian Legal QA (0 rows) | no rows in the archive | — | — | ||
| Open-Mindedness (0-shot) (0 rows) | no rows in the archive | — | — | ||
| OpenBookQA (0 rows) | no rows in the archive | — | — | ||
| PIQA (0 rows) | no rows in the archive | — | — | ||
| PolContro (0 rows) | no rows in the archive | — | — | ||
| PT Hate Speech Binary (0 rows) | no rows in the archive | — | — | ||
| Stories/Jokes (0 rows) | no rows in the archive | — | — | ||
| Talking (0-shot) (0 rows) | no rows in the archive | — | — | ||
| TriviaQA (0 rows) | no rows in the archive | — | — | ||
| TruthfulQA (0-shot) (0 rows) | no rows in the archive | — | — | ||
| TruthfulQA (PT) (0 rows) | no rows in the archive | — | — | ||
| TruthfulQA TR v0.2 (0 rows) | no rows in the archive | — | — | ||
| tweetSentBR (0 rows) | no rows in the archive | — | — | ||
| Unruly (0 rows) | no rows in the archive | — | — | ||
| W/10 (0 rows) | no rows in the archive | — | — | ||
| WiC (0 rows) | no rows in the archive | — | — | ||
| WikiText-103 (0 rows) | no rows in the archive | — | — | ||
| WinoGrande (0 rows) | no rows in the archive | — | — | ||
| Winogrande (5-shot) (0 rows) | no rows in the archive | — | — | ||
| Winogrande TR (0 rows) | no rows in the archive | — | — | ||
| Winogrande TR v0.2 (0 rows) | no rows in the archive | — | — | ||
| World Knowledge (0-shot) (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
162 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 162 until expanded.
Subtasks archive 2025-07-28
25 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 2,047 papers with code (5,335 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
17 Nov 2014 74 repositories listed Syntology ran 13 of 34 samples · 21 unverified · 6 pointer-only (licence)Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions.
-
4 Aug 2013 59 repositories listed Syntology ran 7 of 37 samples · 30 unverified · 4 pointer-only (licence)This paper shows how Long Short-term Memory recurrent neural networks can be used to generate complex sequences with long-range structure, simply by predicting one data point at a time.
-
29 Oct 2019 47 repositories listed Syntology ran 22 of 53 samples · 31 unverified · 7 pointer-only (licence)We evaluate a number of noising approaches, finding the best performance by both randomly shuffling the order of the original sentences and using a novel in-filling scheme, where spans of text are replaced with a single…
-
7 Oct 2016 25 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 14 pointer-only (licence)We observe that our method consistently outperforms BS and previously proposed techniques for diverse decoding from neural sequence models.
-
18 Sep 2016 23 repositories listed Syntology ran 10 of 21 samples · 11 unverified · 13 pointer-only (licence)As a new way of training generative models, Generative Adversarial Nets (GAN) that uses a discriminative model to guide the training of the generative model has enjoyed considerable success in generating real-valued…
-
14 Feb 2019 21 repositories listedNatural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
-
21 Apr 2019 20 repositories listed Syntology ran 44 of 76 samples · 32 unverified · 33 pointer-only (licence)We propose BERTScore, an automatic evaluation metric for text generation.
-
22 May 2020 18 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
1 Jan 2021 13 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedFine-tuning is the de facto way to leverage large pretrained language models to perform downstream tasks.
-
9 Oct 2019 9 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedTransformer architectures have facilitated building higher-capacity models and pretraining has made it possible to effectively utilize this capacity for a wide variety of tasks.
-
8 May 2019 9 repositories listedThis paper presents a new Unified pre-trained Language Model (UniLM) that can be fine-tuned for both natural language understanding and generation tasks.
-
20 Mar 2024 8 repositories listed Syntology ran 3 of 17 samples · 14 unverifiedEfficient fine-tuning is vital for adapting large language models (LLMs) to downstream tasks.
-
21 Sep 2021 8 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedText recognition is a long-standing research problem for document digitalization.
-
11 Sep 2019 8 repositories listed Syntology ran 13 of 14 samples · 1 unverifiedLarge-scale language models show promising text generation capabilities, but users cannot easily control particular aspects of the generated text.
-
15 May 2023 7 repositories listedOur proposed framework provides access: (i) for verifying whether automatic metrics are faithful to human preference, regardless of their correlation level to human; and (ii) for inspecting the strengths and limitations…
-
4 Dec 2019 7 repositories listed Syntology ran 4 of 22 samples · 18 unverifiedLarge transformer-based language models (LMs) trained on huge text corpora have shown unparalleled generation capabilities.
-
7 May 2019 7 repositories listedPre-training and fine-tuning, e.
-
19 Dec 2024 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.
-
7 Jul 2021 6 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedHere, we introduce Discrete Denoising Diffusion Probabilistic Models (D3PMs), diffusion-like generative models for discrete data that generalize the multinomial diffusion model of Hoogeboom et al.
-
12 Aug 2019 6 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 4 pointer-only (licence)Neural text generation is a key tool in natural language applications, but it is well known there are major problems at its core.
-
10 Jun 2019 6 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedThe rapid improvement of language models has raised the specter of abuse of text generation systems.
-
23 May 2019 6 repositories listedGenerative Adversarial Networks (GANs) enjoy great success at image generation, but have proven difficult to train in the domain of natural language.
-
1 Apr 2019 6 repositories listed Syntology ran 1 of 11 samples · 10 unverified · 5 pointer-only (licence)fairseq is an open-source sequence modeling toolkit that allows researchers and developers to train custom models for translation, summarization, language modeling, and other text generation tasks.
-
24 Sep 2017 6 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Automatically generating coherent and semantically meaningful text has many applications in machine translation, dialogue systems, image captioning, etc.
-
8 May 2017 6 repositories listedGenerative Adversarial Nets (GANs) represent an important milestone for effective generative models, which has inspired numerous variants seemingly different from each other.
-
6 Apr 2017 6 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We consider the problem of parsing natural language descriptions into source code written in a general-purpose programming language like Python.
-
27 Feb 2017 6 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe introduce a method for training GANs with discrete data that uses the estimated difference measure from the discriminator to compute importance weights for generated samples, thus providing a policy gradient for…
-
9 Jun 2016 6 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedIn this work, we introduce a model and beam-search training scheme, based on the work of Daume III and Marcu (2005), that extends seq2seq to learn global sequence scores.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections