Browse State-of-the-Art › Text Generation

Text Generation

2,047 papers with code · 20 benchmarks · 162 datasets archive 2025-07-28

AdversarialComputer CodeNatural Language ProcessingSpeech

Text Generation is the task of generating text with the goal of appearing indistinguishable to human-written text. This task is more formally known as "natural language generation" in the literature.

Text generation can be addressed with Markov processes or deep generative models like LSTMs. Recently, some of the most advanced methods for text generation include BART, GPT and other GAN-based approaches. Text generation systems are evaluated either through human ratings or automatic evaluation metrics like METEOR, ROUGE, and BLEU.

Further readings:

( Image credit: Adversarial Ranking for Language Generation )

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

155 leaderboard tables shown for this task, 20 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 155 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
DART (7 rows) T5B Baseline FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text... code — Compare
COCO Captions (5 rows) LeakGAN Long Text Generation via Adversarial Training with Leaked Information code Syntology ran 2 of 2 samples · 0 unverified Compare
EMNLP2017 WMT (5 rows) LeakGAN Long Text Generation via Adversarial Training with Leaked Information code Syntology ran 2 of 2 samples · 0 unverified Compare
ReDial (5 rows) UniCRS Towards Unified Conversational Recommender Systems via... code — Compare
CommonGen (4 rows) UniLM CommonGen: A Constrained Text Generation Challenge for Generative... code Syntology ran 0 of 6 samples · 6 unverified Compare
ROCStories (4 rows) Beam search + A*esque (beam) NeuroLogic A*esque Decoding: Constrained Text Generation with... code — Compare
Chinese Poems (3 rows) RankGAN Adversarial Ranking for Language Generation code — Compare
Czech restaurant information (3 rows) TGen++ The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics — — Compare
OpenWebText (3 rows) GPT2-Hermite Polynomial, trigonometric, and tropical activations code — Compare
SciQ (3 rows) LLaMA-65B+CFG (zero-shot) Stay on topic with Classifier-Free Guidance — — Compare
Yahoo Questions (3 rows) Aggressive VAE Lagging Inference Networks and Posterior Collapse in Variational... code Syntology ran 1 of 10 samples · 9 unverified Compare
ADGEN (1 row) BART (TextBox 2.0) TextBox 2.0: A Text Generation Library with Pre-trained Language Models code Syntology ran 1 of 1 samples · 0 unverified Compare
CMU-SE (1 row) STWGAN-GP Generating Text through Adversarial Training using Skip-Thought Vectors code — Compare
CNN/Daily Mail (1 row) PALM PALM: Pre-training an Autoencoding&Autoregressive Language Model... code — Compare
CSL (1 row) BART (TextBox 2.0) TextBox 2.0: A Text Generation Library with Pre-trained Language Models code Syntology ran 1 of 1 samples · 0 unverified Compare
DailyDialog (1 row) AEM+Attention An Auto-Encoder Matching Model for Learning Utterance-Level... code — Compare
HarmfulQA (1 row) GPT-4 Red-Teaming Large Language Models using Chain of Utterances for... code — Compare
LCSTS (1 row) BART (TextBox 2.0) TextBox 2.0: A Text Generation Library with Pre-trained Language Models code Syntology ran 1 of 1 samples · 0 unverified Compare
LDC2016E25 (1 row) Graph2Seq A Graph-to-Sequence Model for AMR-to-Text Generation code — Compare
One Billion Word (1 row) WGANGP + DGflow Refining Deep Generative Models via Discriminator Gradient Flow code Syntology ran 0 of 1 samples · 1 unverified Compare
AI2 Reasoning Challenge (25-Shot) (0 rows) no rows in the archive — —
AI2 Reasoning Challenge TR v0.2 (0 rows) no rows in the archive — —
AlpacaEval (0 rows) no rows in the archive — —
AlpacaEval 2.0 (GPT-4-1106-Preview) (0 rows) no rows in the archive — —
Arabic Poetry Dataset (6th - 21st century) (0 rows) no rows in the archive — —
arc_challenge (0 rows) no rows in the archive — —
ARC Challenge (25-Shot) (0 rows) no rows in the archive — —
ARC-Challenge (PT) (0 rows) no rows in the archive — —
arc_easy (0 rows) no rows in the archive — —
Assin2 RTE (0 rows) no rows in the archive — —
Assin2 STS (0 rows) no rows in the archive — —
axb (0 rows) no rows in the archive — —
axg (0 rows) no rows in the archive — —
BBH (3-Shot) (0 rows) no rows in the archive — —
Big Bench Hard (3-Shot) (0 rows) no rows in the archive — —
BLUEX (No Images) (0 rows) no rows in the archive — —
BoolQ (0 rows) no rows in the archive — —
CALAME-PT (0 rows) no rows in the archive — —
cb (0 rows) no rows in the archive — —
Censorship (0-shot) (0 rows) no rows in the archive — —
CoLA (0 rows) no rows in the archive — —
Creativity (0-shot) (0 rows) no rows in the archive — —
CrimeStats (0 rows) no rows in the archive — —
crows_pairs_english (0 rows) no rows in the archive — —
crows_pairs_french (0 rows) no rows in the archive — —
DiaBLa (0 rows) no rows in the archive — —
Drop (3-Shot) (0 rows) no rows in the archive — —
ENEM Challenge (No Images) (0 rows) no rows in the archive — —
FaQuAD NLI (0 rows) no rows in the archive — —
GPQA (0-shot) (0 rows) no rows in the archive — —
gsarti/flores_101_afr (0 rows) no rows in the archive — —
gsarti/flores_101_asm (0 rows) no rows in the archive — —
gsarti/flores_101_cat (0 rows) no rows in the archive — —
gsarti/flores_101_ceb (0 rows) no rows in the archive — —
gsarti/flores_101_ces (0 rows) no rows in the archive — —
gsarti/flores_101_ckb (0 rows) no rows in the archive — —
gsarti/flores_101_cym (0 rows) no rows in the archive — —
gsarti/flores_101_dan (0 rows) no rows in the archive — —
gsarti/flores_101_deu (0 rows) no rows in the archive — —
gsarti/flores_101_ell (0 rows) no rows in the archive — —
gsarti/flores_101_eng (0 rows) no rows in the archive — —
gsarti/flores_101_est (0 rows) no rows in the archive — —
gsarti/flores_101_fas (0 rows) no rows in the archive — —
gsarti/flores_101_fin (0 rows) no rows in the archive — —
gsarti/flores_101_fra (0 rows) no rows in the archive — —
gsarti/flores_101_ful (0 rows) no rows in the archive — —
gsarti/flores_101_glg (0 rows) no rows in the archive — —
gsarti/flores_101_guj (0 rows) no rows in the archive — —
gsarti/flores_101_hau (0 rows) no rows in the archive — —
gsarti/flores_101_heb (0 rows) no rows in the archive — —
gsarti/flores_101_mlt (0 rows) no rows in the archive — —
gsarti/flores_101_mon (0 rows) no rows in the archive — —
gsarti/flores_101_msa (0 rows) no rows in the archive — —
gsarti/flores_101_mya (0 rows) no rows in the archive — —
gsarti/flores_101_nld (0 rows) no rows in the archive — —
gsarti/flores_101_nob (0 rows) no rows in the archive — —
gsarti/flores_101_npi (0 rows) no rows in the archive — —
gsarti/flores_101_nso (0 rows) no rows in the archive — —
gsarti/flores_101_nya (0 rows) no rows in the archive — —
gsarti/flores_101_oci (0 rows) no rows in the archive — —
gsarti/flores_101_orm (0 rows) no rows in the archive — —
gsarti/flores_101_ory (0 rows) no rows in the archive — —
gsarti/flores_101_pan (0 rows) no rows in the archive — —
gsarti/flores_101_pol (0 rows) no rows in the archive — —
gsarti/flores_101_por (0 rows) no rows in the archive — —
gsarti/flores_101_ron (0 rows) no rows in the archive — —
gsarti/flores_101_rus (0 rows) no rows in the archive — —
gsarti/flores_101_slk (0 rows) no rows in the archive — —
gsarti/flores_101_slv (0 rows) no rows in the archive — —
gsarti/flores_101_som (0 rows) no rows in the archive — —
gsarti/flores_101_tam (0 rows) no rows in the archive — —
gsarti/flores_101_tel (0 rows) no rows in the archive — —
gsarti/flores_101_tgk (0 rows) no rows in the archive — —
gsarti/flores_101_tgl (0 rows) no rows in the archive — —
gsarti/flores_101_tha (0 rows) no rows in the archive — —
gsarti/flores_101_tur (0 rows) no rows in the archive — —
gsarti/flores_101_ukr (0 rows) no rows in the archive — —
gsarti/flores_101_umb (0 rows) no rows in the archive — —
gsarti/flores_101_urd (0 rows) no rows in the archive — —
gsarti/flores_101_uzb (0 rows) no rows in the archive — —
gsarti/flores_101_vie (0 rows) no rows in the archive — —
gsarti/flores_101_wol (0 rows) no rows in the archive — —
gsarti/flores_101_xho (0 rows) no rows in the archive — —
gsarti/flores_101_yor (0 rows) no rows in the archive — —
gsarti/flores_101_zho_simpl (0 rows) no rows in the archive — —
gsarti/flores_101_zho_trad (0 rows) no rows in the archive — —
gsarti/flores_101_zul (0 rows) no rows in the archive — —
GSM8k (5-shot) (0 rows) no rows in the archive — —
GSM8k TR (0 rows) no rows in the archive — —
GSM8k TR v0.2 (0 rows) no rows in the archive — —
HateBR Binary (0 rows) no rows in the archive — —
HeadQA (0 rows) no rows in the archive — —
HellaSwag (0 rows) no rows in the archive — —
HellaSwag (10-Shot) (0 rows) no rows in the archive — —
HellaSwag (PT) (0 rows) no rows in the archive — —
HellaSwag TR (0 rows) no rows in the archive — —
Humanness (0-shot) (0 rows) no rows in the archive — —
IFEval (0-Shot) (0 rows) no rows in the archive — —
IndicGLUE (0 rows) no rows in the archive — —
Internet (0 rows) no rows in the archive — —
LAMBADA-PT (0 rows) no rows in the archive — —
LogiQA (0 rows) no rows in the archive — —
Math HARD (4-Shot) (0 rows) no rows in the archive — —
MATH Lvl 5 (4-Shot) (0 rows) no rows in the archive — —
MMLU (5-Shot) (0 rows) no rows in the archive — —
MMLU-PRO (5-shot) (0 rows) no rows in the archive — —
MMLU TR (0 rows) no rows in the archive — —
MMLU TR v0.2 (0 rows) no rows in the archive — —
MNLI (0 rows) no rows in the archive — —
MT-Bench (0 rows) no rows in the archive — —
MuSR (0-shot) (0 rows) no rows in the archive — —
Nexa Scientific Tokens (0 rows) no rows in the archive — —
OAB Exams (0 rows) no rows in the archive — —
Open Australian Legal QA (0 rows) no rows in the archive — —
Open-Mindedness (0-shot) (0 rows) no rows in the archive — —
OpenBookQA (0 rows) no rows in the archive — —
PIQA (0 rows) no rows in the archive — —
PolContro (0 rows) no rows in the archive — —
PT Hate Speech Binary (0 rows) no rows in the archive — —
Stories/Jokes (0 rows) no rows in the archive — —
Talking (0-shot) (0 rows) no rows in the archive — —
TriviaQA (0 rows) no rows in the archive — —
TruthfulQA (0-shot) (0 rows) no rows in the archive — —
TruthfulQA (PT) (0 rows) no rows in the archive — —
TruthfulQA TR v0.2 (0 rows) no rows in the archive — —
tweetSentBR (0 rows) no rows in the archive — —
Unruly (0 rows) no rows in the archive — —
W/10 (0 rows) no rows in the archive — —
WiC (0 rows) no rows in the archive — —
WikiText-103 (0 rows) no rows in the archive — —
WinoGrande (0 rows) no rows in the archive — —
Winogrande (5-shot) (0 rows) no rows in the archive — —
Winogrande TR (0 rows) no rows in the archive — —
Winogrande TR v0.2 (0 rows) no rows in the archive — —
World Knowledge (0-shot) (0 rows) no rows in the archive — —

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

162 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 162 until expanded.

GLUEMMLGSM8KMultiNLIWikiText-2HellaSwagTriviaQAPIQACoLAWinoGrandeBoolQOpenBookQATruthfulQAWikiText-103MT-BenchCNN/Daily MailDailyDialogDROPBookCorpusOpenWebTextWiCCOCO CaptionsSciQELI5ROCStoriesBillion Word BenchmarkAlpacaEvalLogiQAWritingPromptsThe StackCommonGenReDialE2EFLoRes-101RealNewsVQGSentence CompressionTinyStoriesLCSTSMuTualOpenDialKGMTNTDARTCCMatrixCSLImage Paragraph CaptioningCONANKPTimesKELMMoral StoriesHeadQAAMR BankCOCO-CNWIQARoboCupIndicGLUECelebV-TextOpusparcusAmazonQATimeTravelPersonalDialogROSCOEOPUSCC-StoriesChart2TextChatHaruhiWikiAtomicEditsHarmfulQALogic2TextLongFormWikiTableTVNHSGEarXiv-10CELLSCLUECorpus2020Microsoft Research Social Media Conversation CorpusWebLINXMLB DatasetPoMoRiSAWOZRuCoLASpatialVOC2KSTACKEXTaiga CorpustweetSentBRDiaBLaGoalWikipedia GenerationCANNOTDR.BENCHOAB ExamsTextBox 2.0Visual Writing PromptsWikiDocEditsYTD-18MCLSECodeSyntaxConciseCzech restaurant informationDIALOCONANDTGBExHVVKaleidoscopeLive Comment DatasetNegotiation Dialogues DatasetQTunaTaoDescribeThreatGram 101 - Extreme Telegram DataAAVE/SAE Paired DatasetAlpaca Data GalicianAppealCaseAskParentsAvicenna: Deductive Commonsense ReasoningCOLLIE-v1Colorsdiaforge-utc-r-0725DocBank-TBDpgMedia2019Educational Grade School Math (EGSM)ENT-DESCExpository ProseFood.com Recipes and InteractionsFraud_Case_VerdictsHALvestHAVOCHiXSTestIce Hockey News DatasetILSP Greek Evaluation SuiteImage Caption Quality DatasetLenta Short SentencesLipogram-eLitMind DictionaryLLMafiaLMSYS-USPMachine_Mindset_MBTI_datasetMatToolsMTTNneedadviceOnlySports DatasetOpenDebateEvidenceOQGendOQRanDPentachromatic Cultural Palette DatasetPolyNewsPTVDReglamento_Aeronautico_Colombiano_2024Rotowire-ModifiedSAD-InstructSG-NLGSGXSTestShort Stories, Adjudicator Scores and Written ReflectionsTamil AlpacaTamil Alpaca OrcaTempWikiBioTexygen PlatformThe Mafia DatasetUHGEvalDatasetUKIL-DB-ENWebBrain-RawWikiWeb2MXWikiRefASSIN2

Subtasks archive 2025-07-28

25 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 2,047 papers with code (5,335 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections