Datasets › WikiANN
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset. It consists of Wikipedia articles that have been annotated with LOC (location), PER (person), and ORG (organization) tags in the IOB2 format¹². This dataset serves as a valuable resource for training and evaluating named entity recognition models across various languages.
For instance, it includes information about notable individuals, places, and organizations mentioned in Wikipedia articles. Researchers and practitioners can use WikiANN to develop and improve natural language processing systems that identify and classify named entities in text.
(1) wikiann · Datasets at Hugging Face. https://huggingface.co/datasets/wikiann. (2) wikiann | TensorFlow Datasets. https://tensorflow.google.cn/datasets/catalog/wikiann. (3) wikiann · Datasets at Hugging Face. https://huggingface.co/datasets/wikiann/viewer/en. (4) WikiAnn Dataset | Papers With Code. https://paperswithcode.com/dataset/wikiann-1.
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Cross-Lingual NER | WikiAnn NER | ByT5 XXL F1 67.7 | ByT5: Towards a token-free future with pre-trained... | huggingface/transformers +4 | 1 | Compare |
| UIE | WikiANN | KnowCoder-7b-IE F1 score 87.0 | KnowCoder: Coding Structured Knowledge into LLMs for... | ICT-GoKnow/KnowCoder | 1 | Compare |
| Token Classification | Wikiann | no rows | — | — | 0 | Compare |
| Token Classification | wikiann sk | no rows | — | — | 0 | Compare |
Papers archive 2025-07-28
2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 72. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction | 1 | 1 | 12 Mar 2024 | not harvested |
| ByT5: Towards a token-free future with pre-trained byte-to-byte models | 5 | 1 | 28 May 2021 | ran 0 of 6 samples (6 unverified) |
Dataset loaders archive 2025-07-28
5 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
Variants archive 2025-07-28
- Wikiann
- WikiANN
- WikiAnn NER
- wikiann sk
- HiNER Collapsed
- HiNER Original
- WikiAnn (Maltese)
- WikiAnn (en, ko, es, pt)
8 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections