Browse State-of-the-Art › Story Generation
Story Generation
110 papers with code · 5 benchmarks · 8 datasets archive 2025-07-28
Story generation is the task of automatically generating a coherent narrative, often from a set of premises or a brief summary.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Fandom dev (1 row) | (NN) Oracle plot + summary + oracle char. desc. | TVStoryGen: A Dataset for Generating Stories with Character Descriptions | code | — | Compare |
| Fandom test (1 row) | (NN) Oracle plot + summary + oracle char. desc. | TVStoryGen: A Dataset for Generating Stories with Character Descriptions | code | — | Compare |
| TVMegaSite dev (1 row) | (NN) Oracle plot + summary + oracle char. desc. | TVStoryGen: A Dataset for Generating Stories with Character Descriptions | code | — | Compare |
| TVMegaSite test (1 row) | (NN) Oracle plot + summary + oracle char. desc. | TVStoryGen: A Dataset for Generating Stories with Character Descriptions | code | — | Compare |
| WritingPrompts (1 row) | BART (TextBox 2.0) | TextBox 2.0: A Text Generation Library with Pre-trained Language Models | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 110 papers with code (235 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 May 2018 7 repositories listedWe explore story generation: creative systems that can build coherent and fluent passages of text about a topic.
-
1 Feb 2022 3 repositories listed Syntology ran 6 of 25 samples · 19 unverified · 15 pointer-only (licence)Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, locally typical sampling offers competitive performance (in both abstractive summarization and story generation) in terms of…
-
14 Sep 2021 3 repositories listedRecent language models can generate interesting and grammatically correct text in story generation but often lack plot development and long-term coherence.
-
22 Jun 2024 2 repositories listedChildren with Autism Spectrum Disorder (ASD) often misunderstand social situations and struggle to participate in daily routines.
-
22 May 2023 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedIn practical scenarios, however, the detector faces texts from various domains or LLMs without knowing their sources.
-
4 Jan 2021 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)In this paper, we advocate to revive latent variable modeling, essentially the power of representation learning, in the era of Transformers to enhance controllability without hurting state-of-the-art generation…
-
2 May 2020 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedIt is well known that the standard likelihood training and approximate decoding objectives in neural text generation models lead to less human-like responses for open-ended tasks such as language modeling and story…
-
30 Apr 2020 2 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedWe propose the task of outline-conditioned story generation: given an outline as a set of phrases that describe key characters and events to appear in a story, the task is to generate a coherent narrative that is…
-
19 Jun 2025 1 repository listedLong story generation remains a challenge for existing large language models (LLMs), primarily due to two main factors: (1) discourse coherence, which requires plot consistency, logical coherence, and completeness in…
-
11 Jun 2025 1 repository listedText-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a…
-
5 Jun 2025 1 repository listed Syntology ran 5 of 16 samples · 11 unverifiedWe show that language models are capable of counterfactual reasoning in this controlled setup and provide insights that counterfactual reasoning for a broad class of functions can be reduced to a transformation on…
-
20 May 2025 1 repository listedRobustly evaluating the long-form storytelling capabilities of Large Language Models (LLMs) remains a significant challenge, as existing benchmarks often lack the necessary scale, diversity, or objective measures.
-
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation15 May 2025 1 repository listedVisual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations.
-
28 Mar 2025 1 repository listed Syntology ran 1 of 15 samples · 14 unverifiedGenerating high-quality stories spanning thousands of tokens requires competency across a variety of skills, from tracking plot and character arcs to keeping a consistent and engaging style.
-
18 Feb 2025 1 repository listedHuman evaluation highlights the high quality of our Author Writing Sheet and provides valuable insights into the personalized story generation task.
-
26 Jan 2025 1 repository listedThe stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms.
-
23 Jan 2025 1 repository listedDrawing inspiration from the inherent context consistency, we propose a novel training-free method for consistent text-to-image (T2I) generation, termed "One-Prompt-One-Story" (1Prompt1Story).
-
18 Dec 2024 1 repository listedLong-form story generation task aims to produce coherent and sufficiently lengthy text, essential for applications such as novel writingand interactive storytelling.
-
4 Dec 2024 1 repository listedAlthough several fine-tuning and prompting techniques have been suggested to tackle the issue, they are often tailored to specific tasks or come with a substantial increase in computational cost and latency.
-
11 Nov 2024 1 repository listedWhile a large body of work inspects language models for biases concerning gender, race, occupation and religion, biases of geographical nature are relatively less explored.
-
7 Nov 2024 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedRecent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs.
-
4 Nov 2024 1 repository listedAutomated metrics show that LLMs generate stylistically complex stories, but tend to fall short in terms of novelty, surprise and diversity when compared to average human writers.
-
3 Oct 2024 1 repository listedGenerating a long story of several thousand words with narrative coherence using Large Language Models (LLMs) has been a challenging task.
-
3 Sep 2024 1 repository listedIn this position paper, we argue that the speed of advancement of this technology requires us, as computer and data scientists, to mobilize and develop a values-based auditing framework containing a community…
-
15 Aug 2024 1 repository listedWe highlight the limitations of current LLMs in fiction writing and advocate for future research to test and create story worlds for LLMs to reside in.
-
9 Aug 2024 1 repository listedData-driven storytelling is a powerful method for conveying insights by combining narrative techniques with visualizations and text.
-
11 Jul 2024 1 repository listed Syntology ran 14 of 17 samples · 3 unverified · 17 pointer-only (licence)We further propose multimodal attention sink mechanism to enable the generation of stories with up to 25 sequences (only 10 for training) in a highly efficient autoregressive manner.
-
10 Jul 2024 1 repository listedFor our GPT-41 data, we introduce crafted prompts that allow us to generate data well-suited to the Arabic context in both Modern Standard Arabic (MSA) and two Arabic dialects (Egyptian and Moroccan).
-
Human-AI Collaborative Taxonomy Construction: A Case Study in Profession-Specific Writing Assistants26 Jun 2024 1 repository listedLarge Language Models (LLMs) have assisted humans in several writing tasks, including text revision and story generation.
-
18 Jun 2024 1 repository listedDespite this, there have not been efforts to explore such collaborative writing.
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections