Papers › Challenges in Data-to-Document Generation

Challenges in Data-to-Document Generation

25 Jul 2017EMNLP 2017 9arXiv:1707.08052archive 2025-07-28

Sam Wiseman, Stuart M. Shieber, Alexander M. Rush

Recent neural models have shown significant progress on the problem of generating short descriptive texts conditioned on a small number of database records. In this work, we suggest a slightly more difficult data-to-text generation task, and investigate how effective current approaches are on this task. In particular, we introduce a new, large-scale corpus of data records paired with descriptive documents, propose a series of extractive evaluation methods for analyzing performance, and obtain baseline results using current neural generation methods. Experiments show that these models produce fluent text, but fail to convincingly approximate human-generated documents. Moreover, even templated baselines exceed the performance of these neural models on some metrics, though copy- and reconstruction-based extensions lead to noticeable improvements.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

harvardnlp/boxscore-data officialmentioned in papermentioned on GitHub report
harvardnlp/data2text officialmentioned in papermentioned on GitHub report
KaijuML/rotowire-rg-metric mentioned on GitHubpytorch report
ratishsp/data2text-1 mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Data-to-Text GenerationDescriptiveText Generation

Datasets

Introduced by this paper, per the archive.

RotoWire

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Data-to-Text Generation RotoWire Encoder-decoder + conditional copy BLEU 14.19 #6 of 6 Archive leaderboard report
Data-to-Text Generation RotoWire (Content Ordering) Encoder-decoder + conditional copy BLEU 14.49 #5 of 5 Archive leaderboard report
Data-to-Text Generation RotoWire (Content Ordering) Encoder-decoder + conditional copy DLD 8.68% #5 of 5 Archive leaderboard report
Data-to-Text Generation RotoWire (Relation Generation) Encoder-decoder + conditional copy Precision 74.80% #6 of 6 Archive leaderboard report
Data-to-Text Generation RotoWire (Relation Generation) Encoder-decoder + conditional copy count 23.72 #6 of 6 Archive leaderboard report
Data-to-Text Generation Rotowire (Content Selection) Encoder-decoder + conditional copy Precision 29.49% #5 of 5 Archive leaderboard report
Data-to-Text Generation Rotowire (Content Selection) Encoder-decoder + conditional copy Recall 36.18% #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections