{"url":"/dataset/gyafc","name":"GYAFC","full_name":"Grammarly’s Yahoo Answers Formality Corpus","description_markdown":"Grammarly’s Yahoo Answers Formality Corpus (GYAFC) is the largest dataset for any style containing a total of 110K informal / formal sentence pairs.\r\n\r\nYahoo Answers is a question answering forum, contains a large number of informal sentences and allows redistribution of data. The authors used the Yahoo Answers L6 corpus to create the GYAFC dataset of informal and formal sentence pairs. In order to ensure a uniform distribution of data, they removed sentences that are questions, contain URLs, and are shorter than 5 words or longer than 25. After these preprocessing steps, 40 million sentences remain. \r\n\r\nThe Yahoo Answers corpus consists of several different domains like Business, Entertainment & Music, Travel, Food, etc. Pavlick and Tetreault formality classifier (PT16) shows that the formality level varies significantly\r\nacross different genres. In order to control for this variation, the authors work with two specific domains that contain the most informal sentences and show results on training and testing within those categories. The authors use the formality classifier from PT16 to identify informal sentences and train this classifier on the Answers genre of the PT16 corpus\r\nwhich consists of nearly 5,000 randomly selected sentences from Yahoo Answers manually annotated on a scale of -3 (very informal) to 3 (very formal). They find that the domains of Entertainment & Music and Family & Relationships contain the most informal sentences and create the GYAFC dataset using these domains.\r\n\r\nSource: [Dear Sir or Madam, May I Introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer](https://arxiv.org/pdf/1803.06535v2.pdf)","description_withheld":null,"homepage":"https://github.com/raosudha89/GYAFC-corpus","introduced_date":"2018-03-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/dear-sir-or-madam-may-i-introduce-the-gyafc","title":"Dear Sir or Madam, May I introduce the GYAFC Dataset: Corpus, Benchmarks and Metrics for Formality Style Transfer","first_author":"Sudha Rao","url":null},"license":{"name":"Custom (research-only)","url":"https://github.com/raosudha89/GYAFC-corpus"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Style Transfer","url":"/task/style-transfer","datasets_with_task":"/datasets/task/style-transfer"},{"name":"Unsupervised Text Style Transfer","url":"/task/unsupervised-text-style-transfer","datasets_with_task":"/datasets/task/unsupervised-text-style-transfer"},{"name":"Formality Style Transfer","url":"/task/formality-style-transfer","datasets_with_task":"/datasets/task/formality-style-transfer"}],"languages":[],"variants":["GYAFC"],"data_loaders":[{"repo":"https://github.com/raosudha89/GYAFC-corpus","url":"https://github.com/raosudha89/GYAFC-corpus","frameworks":[]}],"num_papers_in_archive":103,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/formality-style-transfer-on-gyafc","task":"Formality Style Transfer","dataset_variant":"GYAFC","rows":1,"metrics":["BLEU"],"first_row_in_archive_order":{"model":"Consistency Training","paper":"/paper/semi-supervised-formality-style-transfer-with","metrics":{"BLEU":"81.37"},"code_links":[{"title":"aolius/semi-fst","url":"https://github.com/aolius/semi-fst"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/style-transfer-on-gyafc","task":"Style Transfer","dataset_variant":"GYAFC","rows":1,"metrics":["Accuracy","BLEU-4","Harmonic mean"],"first_row_in_archive_order":{"model":"BART (TextBox 2.0)","paper":"/paper/textbox-2-0-a-text-generation-library-with","metrics":{"Accuracy":"94.37","BLEU-4":"76.93","Harmonic mean":"84.74"},"code_links":[{"title":"RUCAIBox/TextBox","url":"https://github.com/RUCAIBox/TextBox"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/textbox-2-0-a-text-generation-library-with","title":"TextBox 2.0: A Text Generation Library with Pre-trained Language Models","date":"2022-12-26","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/semi-supervised-formality-style-transfer-with","title":"Semi-Supervised Formality Style Transfer with Consistency Training","date":"2022-03-25","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}