Datasets › WebNLG

WebNLG

Introduced by Claire Gardent et al. in Creating Training Corpora for NLG Micro-Planners1 Jan 2017 archive 2025-07-28

The WebNLG corpus comprises of sets of triplets describing facts (entities and relations between them) and the corresponding facts in form of natural language text. The corpus contains sets with up to 7 triplets each along with one or more reference texts for each set. The test set is split into two parts: seen, containing inputs created for entities and relations belonging to DBpedia categories that were seen in the training data, and unseen, containing inputs extracted for entities and relations belonging to 5 unseen categories.

Initially, the dataset was used for the WebNLG natural language generation challenge which consists of mapping the sets of triplets to text, including referring expression generation, aggregation, lexicalization, surface realization, and sentence segmentation. The corpus is also used for a reverse task of triplets extraction.

Versioning history of the dataset can be found here.

Source: Step-by-Step: Separating Planning from Realization in Neural Data-to-Text Generation Image Source: https://paperswithcode.com/paper/creating-training-corpora-for-nlg-micro/

It's also available here: https://huggingface.co/datasets/web_nlg Note: "The v3 release (release_v3.0_en, release_v3.0_ru) for the WebNLG2020 challenge also supports a semantic parsing task."

Benchmarks archive 2025-07-28

All 17 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Data-to-Text Generation WebNLG Control Prefixes (A1, T5-large) BLEU 67.32 Control Prefixes for Parameter-Efficient Text Generation Yale-LILY/dart +1 20 Compare
Relation Extraction WebNLG UniRel F1 94.7 UniRel: Unified Representation and Interaction for Joint... wtangdev/unirel 14 Compare
KG-to-Text Generation WebNLG 2.0 (Unconstrained) GAP - Me,r+γ BLEU 66.2 GAP: A Graph-aware Language Model Framework for... acolas1/GAP_COLING2022 13 Compare
Joint Entity and Relation Extraction WebNLG 3.0 ReGen (Ours) T2G.CE F1 72.3 ReGen: Reinforcement Learning for Text and Knowledge... IBM/regen 10 Compare
KG-to-Text Generation WebNLG 2.0 (Constrained) FactT5B BLEU 67.08 FactSpotter: Evaluating the Factual Faithfulness of... guihuzhang/FactSpotter 9 Compare
Data-to-Text Generation WebNLG Full Control Prefixes (A1, A2, T5-large) BLEU 62.27 Control Prefixes for Parameter-Efficient Text Generation Yale-LILY/dart +1 8 Compare
Joint Entity and Relation Extraction WebNLG SPN F1 93.4 Joint Entity and Relation Extraction with Set Prediction Networks DianboWork/SPN4RE 3 Compare
Data-to-Text Generation WebNLG en mBART METEOR 0.462 The GEM Benchmark: Natural Language Generation, its... — 2 Compare
KG-to-Text Generation WebNLG (All) T5_large BLEU 59.70 Investigating Pretrained Language Models for... bjascob/amrlib +2 2 Compare
KG-to-Text Generation WebNLG (Seen) T5_large BLEU 64.71 Investigating Pretrained Language Models for... bjascob/amrlib +2 2 Compare
KG-to-Text Generation WebNLG (Unseen) T5_large BLEU 53.67 Investigating Pretrained Language Models for... bjascob/amrlib +2 2 Compare
Table-to-Text Generation WebNLG (All) HTLM (fine-tuning) BLEU 55.6 HTLM: Hyper-Text Pre-Training and Prompting of Language Models — 2 Compare
Table-to-Text Generation WebNLG (Seen) HTLM (fine-tuning) BLEU 65.4 HTLM: Hyper-Text Pre-Training and Prompting of Language Models — 2 Compare
Table-to-Text Generation WebNLG (Unseen) HTLM (fine-tuning) BLEU 48.4 HTLM: Hyper-Text Pre-Training and Prompting of Language Models — 2 Compare
Graph-to-Sequence WebNLG CGE-LW BLEU 63.69 Modeling Global and Local Node Contexts for Text... UKPLab/kg2text 1 Compare
Unsupervised KG-to-Text Generation WebNLG v2.1 GT-BT (sampled noise) BLEU 37.7 An Unsupervised Joint System for Text Generation from... mnschmit/unsupervised-graph-text-conversion 1 Compare
Unsupervised semantic parsing WebNLG v2.1 GT-BT (sampled noise) F1 39.1 An Unsupervised Joint System for Text Generation from... mnschmit/unsupervised-graph-text-conversion 1 Compare

Papers archive 2025-07-28

30 shown of 39 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 149. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy 0 2 25 Jan 2024 not harvested
FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text Generation 1 8 25 Oct 2023 not harvested
TextBox 2.0: A Text Generation Library with Pre-trained Language Models 1 1 26 Dec 2022 ran 1 of 1 samples (0 unverified)
Knowledge Graph Generation From Text 1 6 18 Nov 2022 not harvested
UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction 1 1 16 Nov 2022 not harvested
GAP: A Graph-aware Language Model Framework for Knowledge Graph-to-Text Generation 1 6 13 Apr 2022 not harvested
TDEER: An Efficient Translating Decoding Schema for Joint Extraction of Entities and Relations 1 2 1 Nov 2021 not harvested
Control Prefixes for Parameter-Efficient Text Generation 2 4 15 Oct 2021 not harvested
ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models 1 4 27 Aug 2021 ran 0 of 1 samples (1 unverified)
A Partition Filter Network for Joint Entity and Relation Extraction 1 1 27 Aug 2021 ran 5 of 7 samples (2 unverified)
HTLM: Hyper-Text Pre-Training and Prompting of Language Models 0 8 14 Jul 2021 not harvested
JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs 1 8 19 Jun 2021 not harvested
Stage-wise Fine-tuning for Graph-to-Text Generation 1 2 17 May 2021 not harvested
Representation Iterative Fusion based on Heterogeneous Graph Neural Network for Joint Entity and Relation Extraction 1 1 8 May 2021 not harvested
Structural Information Preserving for Graph-to-Text Generation 1 1 12 Feb 2021 not harvested
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics 0 2 2 Feb 2021 not harvested
Joint Entity and Relation Extraction with Set Prediction Networks 1 2 3 Nov 2020 not harvested
TPLinker: Single-stage Joint Extraction of Entities and Relations Through Token Pair Linking 1 1 26 Oct 2020 not harvested
KGPT: Knowledge-Grounded Pre-Training for Data-to-Text Generation 1 2 5 Oct 2020 not harvested
OpenUE: An Open Toolkit of Universal Extraction from Text 1 1 1 Oct 2020 not harvested
Contrastive Triple Extraction with Generative Transformer 0 1 14 Sep 2020 not harvested
Investigating Pretrained Language Models for Graph-to-Text Generation 3 8 16 Jul 2020 ran 2 of 12 samples (10 unverified)
A Relation-Specific Attention Network for Joint Entity and Relation Extraction 1 1 1 Jul 2020 not harvested
Modeling Graph Structure via Relative Position for Text Generation from Knowledge Graphs 0 1 16 Jun 2020 not harvested
Text-to-Text Pre-Training for Data-to-Text Tasks 2 2 21 May 2020 not harvested
Recurrent Interaction Network for Jointly Extracting Entities and Classifying Relations 0 1 1 May 2020 not harvested
Have Your Text and Use It Too! End-to-End Neural Data-to-Text Generation with Semantic Fidelity 1 1 8 Apr 2020 not harvested
Modeling Global and Local Node Contexts for Text Generation from Knowledge Graphs 1 2 29 Jan 2020 not harvested
CopyMTL: Copy Mechanism for Joint Extraction of Entities and Relations with Multi-Task Learning 2 1 24 Nov 2019 not harvested
Joint Extraction of Entities and Relations Based on a Novel Decomposition Strategy 1 1 10 Sep 2019 not harvested

The full list of 39 is in the JSON twin.

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • WebNLG 3.0
  • WebNLG 2.0 (Unconstrained)
  • WebNLG 2.0 (Constrained)
  • WebNLG (Unseen)
  • WebNLG (Seen)
  • WebNLG (All)
  • WebNLG (Constrained)
  • WebNLG(C)
  • WebNLG(U)
  • WebNLG en
  • WebNLG v2.1
  • WebNLG Full
  • WebNLG

13 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections