{"url":"/dataset/chem-finese","name":"Chem-FINESE","full_name":null,"description_markdown":"The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET.\r\nFor each dataset, we randomly sample a subset based on the frequency of each type class. Specifically, given a dataset, we first set the number of maximum entity mentions $k$ for the most frequent entity type in the dataset. We then randomly sample other types and ensure that the distribution of each type remains the same as in the original dataset. We choose the values $6, 9, 12, 15, 18$ as the potential maximum entity mentions for $k$. The ChemNER+ and CHEMET few-shot datasets contain 52 and 28 types respectively.","description_withheld":null,"homepage":"https://github.com/EagleW/Chem-FINESE/tree/main/data","introduced_date":"2024-01-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/chem-finese-validating-fine-grained-few-shot-1","title":"Chem-FINESE: Validating Fine-Grained Few-shot Entity Extraction through Text Reconstruction","first_author":"Qingyun Wang","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Named Entity Recognition (NER)","url":"/task/named-entity-recognition-ner","datasets_with_task":"/datasets/task/named-entity-recognition-ner"},{"name":"Few-shot NER","url":"/task/few-shot-ner","datasets_with_task":"/datasets/task/few-shot-ner"},{"name":"Chemical Entity Recognition","url":"/task/chemical-entity-recognition","datasets_with_task":"/datasets/task/chemical-entity-recognition"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Chem-FINESE"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}