{"url":"/dataset/fairtranslate-fr","name":"FairTranslate_fr","full_name":null,"description_markdown":"The FairTranslate Dataset includes **2,418 sentence pairs**, each centered around an occupation, designed to assess gender expression and translation in English-French contexts. Each English sentence appears in **three gender variants** (male, female, inclusive), allowing for direct counterfactual comparisons. This structure supports fairness evaluations and helps analyze how models handle **grammatical gender, inclusive forms**, and **coreference resolution** in translation.\r\n\r\nEach example in the dataset is annotated with rich metadata:\r\n\r\n- **English**: Sentence involving an occupation, designed to test both explicit and subtle cues of gender.\r\n- **French**: Ground-truth translation faithfully aligned with the intended gender variant.\r\n- **Gender**: Target gender for translation: `male`, `female`, or `inclusive`.\r\n- **Ambiguity**: Level of gender ambiguity in the English source sentence:\r\n  - `ambiguous`: No explicit gender cues.\r\n  - `unambiguous`: Clear pronouns or cues (e.g., “he”, “she”, “they”).\r\n  - `long unambiguous`: Gender resolvable from distant context, testing long-range coreference.\r\n- **Stereotype**: Whether the occupation is `male-stereotyped`, `female-stereotyped`, or `gender-balanced`, based on real-world statistics from [Statbel](https://statbel.fgov.be/).\r\n- **Occupation**: A list of the three gendered French forms for each occupation, e.g., `[\"infirmier\", \"infirmière\", \"infirmier.ière\"]`.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Fannyjrd/FairTranslate_fr","introduced_date":"2025-04-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/fairtranslate-an-english-french-dataset-for","title":"FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity","first_author":"Fanny Jourdan","url":null},"license":{"name":"mit","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Translation","url":"/task/translation","datasets_with_task":"/datasets/task/translation"}],"languages":[{"name":"English","url":"/datasets/language/english"},{"name":"French","url":"/datasets/language/french"}],"variants":["FairTranslate_fr"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}