{"url":"/dataset/nlp-taxonomy-classification-test-data","name":"NLP Taxonomy Classification Data","full_name":null,"description_markdown":"The dataset consists of titles and abstracts from NLP-related papers. Each paper is annotated with multiple fields of study from an NLP taxonomy. The training dataset contains 178,521 weakly annotated samples. The test dataset consists of 828 manually annotated samples from the EMNLP22 conference. The manually labeled test dataset might not contain all possible classes since it consists of EMNLP22 papers only, and some rarer classes haven’t been published there. Therefore, we advise creating an additional test or validation set from the train data that includes all the possible classes.","description_withheld":null,"homepage":"https://huggingface.co/datasets/TimSchopf/nlp_taxonomy_data","introduced_date":"2023-07-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/exploring-the-landscape-of-natural-language","title":"Exploring the Landscape of Natural Language Processing Research","first_author":"Tim Schopf","url":null},"license":{"name":"MIT","url":"https://huggingface.co/datasets/TimSchopf/nlp_taxonomy_data"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Classification","url":"/task/classification-1","datasets_with_task":"/datasets/task/classification-1"},{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"},{"name":"Multi-Label Classification","url":"/task/multi-label-classification","datasets_with_task":"/datasets/task/multi-label-classification"},{"name":"Multilabel Text Classification","url":"/task/multilabel-text-classification","datasets_with_task":"/datasets/task/multilabel-text-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["NLP Taxonomy Classification Data"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}