{"url":"/dataset/namextend","name":"NAMEXTEND","full_name":null,"description_markdown":"This dataset extends [NAMEXACT](https://huggingface.co/datasets/aieng-lab/namexact) by including words that can be used as names, but may not exclusively be used as names in every context.\r\n\r\n\r\n## Dataset Details\r\n\r\n### Dataset Description\r\n\r\n<!-- Provide a longer summary of what this dataset is. -->\r\nUnlike [NAMEXACT](https://huggingface.co/datasets/aieng-lab/namexact), this datasets contains *words* that are mostly used as *names*, but may also be used in other contexts, such as\r\n\r\n- *Christian* (believer in Christianity)\r\n- *Drew* (simple past of the verb to draw)\r\n- *Florence* (an Italian city)\r\n- *Henry* (the SI unit of inductance)\r\n- *Mercedes* (a car brand)\r\n\r\nIn addition, names with ambiguous gender are included - once for each `gender`. For instance, `Skyler` is included as female (`F`) name with a `probability` of 37.3%, and as male (`M`) name with a `probability` of 62.7%.\r\n\r\n\r\n### Dataset Sources [optional]\r\n\r\n<!-- Provide the basic links for the dataset. -->\r\n\r\n- **Repository:** [github.com/aieng-lab/gradiend](https://github.com/aieng-lab/gradiend)\r\n\r\n- **Original Dataset:** [Gender by Name](https://archive.ics.uci.edu/dataset/591/gender+by+name)\r\n\r\n\r\n## Dataset Structure\r\n\r\n- `name`: the name\r\n- `gender`: the gender of the name (`M` for male and `F` for female)\r\n- `count`: the count value of this name (raw value from the original dataset)\r\n- `probability`: the probability of this name (raw value from original dataset; not normalized to this dataset!)\r\n- `gender_agreement`: a value describing the certainty that this name has an unambiguous gender computed as the maximum probability of that name across both genders, e.g., $max(37.7%, 62.7%)=62.7%$ for *Skyler*. For names with a unique `gender` in this dataset, this value is 1.0\r\n- `primary_gender`: is equal to `gender` for names with a unique gender in this dataset, and equals otherwise the gender of that name with higher probability\r\n- `genders`: label `B` if *both* genders are contained for this name in this dataset, otherwise equal to `gender`\r\n- `prob_F`: the probability of that name being used as a female name (i.e., 0.0 or 1.0 if `genders` != `B`)\r\n- `prob_M`: the probability of that name being used as a male name\r\n\r\n## Dataset Creation\r\n\r\n\r\n### Source Data\r\n\r\n<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->\r\nThe data is created by filtering [Gender by Name](https://archive.ics.uci.edu/dataset/591/gender+by+name).\r\n\r\n\r\n#### Data Collection and Processing\r\n\r\n<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->\r\n\r\nThe original data is filtered to contain only names with a `count` of at least 100 to remove very rare names. This threshold reduces the total number of names by $72%, from 133910 to 37425.\r\n\r\n\r\n\r\n## Bias, Risks, and Limitations\r\n\r\n<!-- This section is meant to convey both technical and sociotechnical limitations. -->\r\n\r\nThe original dataset provides counts of names (with their gender) for male and female babies from open-source government authorities in the US (1880-2019), UK (2011-2018), Canada (2011-2018), and Australia (1944-2019) in these periods","description_withheld":null,"homepage":"https://huggingface.co/datasets/aieng-lab/namextend","introduced_date":"2025-02-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/gradiend-monosemantic-feature-learning-within","title":"GRADIEND: Monosemantic Feature Learning within Neural Networks Applied to Gender Debiasing of Transformer Models","first_author":"Jonathan Drechsel","url":null},"license":{"name":"cc-by.4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["NAMEXTEND"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}