{"url":"/dataset/mttn","name":"MTTN","full_name":null,"description_markdown":"**MTTN** is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc. MTTN consists of over 2.4M sentences that are divided over 5 stages creating a combination amounting to over 12M pairs, along with a vocab size of consisting more than 300 thousands unique words that creates an abundance of variations.\r\n\r\nSource: [MTTN: Multi-Pair Text to Text Narratives for Prompt Generation](https://arxiv.org/pdf/2301.10172v1.pdf)","description_withheld":null,"homepage":"https://github.com/mttn2023/mttn","introduced_date":"2023-01-21","introduced_date_note":null,"introduced_by":{"paper":"/paper/mttn-multi-pair-text-to-text-narratives-for","title":"MTTN: Multi-Pair Text to Text Narratives for Prompt Generation","first_author":"Archan Ghosh","url":null},"license":{"name":"CC0-1.0 license","url":"https://github.com/mttn2023/mttn/blob/main/LICENSE.md"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Generation","url":"/task/text-generation","datasets_with_task":"/datasets/task/text-generation"},{"name":"Text2text Generation","url":"/task/text2text-generation-1","datasets_with_task":"/datasets/task/text2text-generation-1"}],"languages":[],"variants":["MTTN"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}