{"url":"/dataset/tf1-en-3m","name":"TF1-EN-3M","full_name":"klusai/ds-tf1-en-3m","description_markdown":"# TF1-EN-3M: Three Million Synthetic Moral Fables for Open Language Models\r\n\r\n**TF1-EN-3M** is a large-scale synthetic dataset of **3,000,000 English-language moral fables**, generated by instruction-tuned language models with no more than 8 billion parameters. The stories are aimed at child-friendly educational and moral reasoning applications and follow a consistent six-part narrative scaffold: `character → trait → setting → conflict → resolution → moral`.\r\n\r\n## Dataset Characteristics\r\n\r\n- **Size**: 3 million stories (~1B tokens)\r\n- **Format**: JSON Lines, with detailed metadata including prompt elements, model configuration, generation time, token counts, and costs\r\n- **Generation Models**: Evaluated across 10 open-weight LLMs; final dataset generated using `LLaMA-3.1-8B-Instruct` for optimal quality and cost balance\r\n- **Story Structure**: Each story ends with an explicit moral and follows a template-driven structure\r\n- **Target Audience**: Designed primarily for children aged 4–7 (age group B), with simple vocabulary and accessible narratives\r\n\r\n## Motivation\r\n\r\nNatural language processing lacks large, structured corpora of fables that combine creative storytelling with explicit moral lessons. Existing human-authored datasets like Aesop's Fables are limited in scale and diversity. TF1-EN-3M bridges this gap by:\r\n\r\n- Demonstrating that mid-sized open models can reliably generate coherent, instructive stories\r\n- Enabling research into value alignment, narrative intelligence, and low-resource model fine-tuning\r\n- Offering a reproducible, cost-efficient alternative to proprietary LLM pipelines\r\n\r\n## Summary of Content\r\n\r\nEach entry in the dataset includes:\r\n\r\n- A structured prompt with narrative elements\r\n- The generated fable text\r\n- Metadata (model name, inference time, token usage, cost, etc.)\r\n- Quality assessments (via LLM-based scoring for grammar, creativity, moral clarity, and structure adherence)\r\n\r\n## Use Cases\r\n\r\nTF1-EN-3M is suitable for a wide range of tasks and applications:\r\n\r\n- **Training** small or medium-sized LLMs for story generation or moral reasoning\r\n- **Benchmarking** models on tasks like moral inference, story-to-moral mapping, or story quality evaluation\r\n- **Educational AI tools**, such as interactive storytelling tutors or automated moral education platforms\r\n- **Creative NLP** research, including literary analysis and narrative generation\r\n- **Multilingual extension** by swapping out prompt elements for other languages\r\n\r\n## Citation\r\n\r\nIf you use TF1-EN-3M, please cite:\r\n\r\n```bibtex\r\n@misc{nadas2025tf1en3m,\r\n  title={TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models},\r\n  author={Mihai Nădaș and Laura Dioșan and Andreea Tomescu and Andrei Pișcoran},\r\n  year={2025},\r\n  eprint={2504.20605},\r\n  archivePrefix={arXiv},\r\n  primaryClass={cs.CL}\r\n}\r\n```\r\n\r\n## Dataset Access\r\n\r\n- **Hugging Face**: [klusai/ds-tf1-en-3m](https://huggingface.co/datasets/klusai/ds-tf1-en-3m)\r\n- **Generation & Evaluation Code**: [TinyFabulist GitHub](https://github.com/klusai/tinyfabulist)","description_withheld":null,"homepage":"https://huggingface.co/datasets/klusai/ds-tf1-en-3m","introduced_date":"2025-04-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/tf1-en-3m-three-million-synthetic-moral","title":"TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models","first_author":"Mihai Nadas","url":null},"license":{"name":"MIT","url":"https://choosealicense.com/licenses/mit/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Synthetic Data Generation","url":"/task/synthetic-data-generation","datasets_with_task":"/datasets/task/synthetic-data-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["TF1-EN-3M"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}