{"url":"/dataset/nkjp-ner","name":"NKJP-NER","full_name":null,"description_markdown":"The **NKJP-NER dataset** is based on a human-annotated part of the **National Corpus of Polish (NKJP)**. In this dataset, sentences containing **named entities** of exactly one type have been extracted. The primary task associated with this dataset is to **predict the type of the named entity**. The dataset provides examples split into three categories:\r\n\r\n1. **Train**: Contains **15,794** examples.\r\n2. **Validation**: Contains **1,941** examples.\r\n3. **Test**: Contains **2,058** examples.\r\n\r\nThe named entity types in this dataset include:\r\n- **geogName**\r\n- **noEntity**\r\n- **orgName**\r\n- **persName**\r\n- **placeName**\r\n- **time**\r\n\r\nThe NKJP-NER dataset is a valuable resource for natural language processing tasks related to named entity recognition in the Polish language. It is available under the **GNU GPL v.3** license and is currently at version **1.1.0**¹²³.\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) nkjp-ner | TensorFlow Datasets. https://www.tensorflow.org/datasets/community_catalog/huggingface/nkjp-ner.\r\n(2) +86 Ner Datasets - NLP Database - Metatext. https://metatext.io/datasets-list/ner-task.\r\n(3) The Best Polish Language Datasets of 2022 | Twine. https://www.twine.net/blog/polish-language-datasets/.\r\n(4) nkjp-ner · Datasets at Hugging Face. https://huggingface.co/datasets/nkjp-ner/viewer/default/test.\r\n(5) undefined. https://klejbenchmark.com/static/data/klej_nkjp-ner.zip.","description_withheld":null,"homepage":"https://www.tensorflow.org/datasets/community_catalog/huggingface/nkjp-ner","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NKJP-NER"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}