Datasets › WikiANN

WikiANN (PAN-X)

Introduced by Xiaoman Pan et al. in Cross-lingual Name Tagging and Linking for 282 Languages14 Sep 2023 archive 2025-07-28

WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset. It consists of Wikipedia articles that have been annotated with LOC (location), PER (person), and ORG (organization) tags in the IOB2 format¹². This dataset serves as a valuable resource for training and evaluating named entity recognition models across various languages.

For instance, it includes information about notable individuals, places, and organizations mentioned in Wikipedia articles. Researchers and practitioners can use WikiANN to develop and improve natural language processing systems that identify and classify named entities in text.

(1) wikiann · Datasets at Hugging Face. https://huggingface.co/datasets/wikiann. (2) wikiann | TensorFlow Datasets. https://tensorflow.google.cn/datasets/catalog/wikiann. (3) wikiann · Datasets at Hugging Face. https://huggingface.co/datasets/wikiann/viewer/en. (4) WikiAnn Dataset | Papers With Code. https://paperswithcode.com/dataset/wikiann-1.

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Cross-Lingual NER WikiAnn NER ByT5 XXL F1 67.7 ByT5: Towards a token-free future with pre-trained... huggingface/transformers +4 1 Compare
UIE WikiANN KnowCoder-7b-IE F1 score 87.0 KnowCoder: Coding Structured Knowledge into LLMs for... ICT-GoKnow/KnowCoder 1 Compare
Token Classification Wikiann no rows — — 0 Compare
Token Classification wikiann sk no rows — — 0 Compare

Papers archive 2025-07-28

2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 72. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction 1 1 12 Mar 2024 not harvested
ByT5: Towards a token-free future with pre-trained byte-to-byte models 5 1 28 May 2021 ran 0 of 6 samples (6 unverified)

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

EnglishFrenchSpanishGermanItalianChineseBengaliMultilingualJapaneseRussianPortugueseAfrikaansAlbanianAmharicArabicArmenianBambaraBasqueBelarusianBretonBulgarianCatalanChurch SlavicCroatianCzechDanishDutchErzyaEstonianFaroeseFinnishGalicianGothicHebrewHindiHungarianIcelandicIndonesianIrishKazakhKomi-PermyakKoreanLatinLatvianLithuanianLivviMalteseManxMarathiMokshaNorthern SamiNorwegianPersianPolishRomanianRussia BuriatSanskritScottish GaelicSerbianSlovakSlovenianSwedishTagalogTamilTeluguThaiTurkishUighurUkrainianUpper SorbianUrduVietnameseWelshWolofYorubaAfarAbkhazianAchineseAdygheAkanTosk AlbanianOld English (ca. 450-1100)Official Aramaic (700-300 BCE)AragoneseEgyptian ArabicAssameseAsturianAvaricAymaraSouth AzerbaijaniAzerbaijaniBashkirBavarianCentral BikolBislamaBanjarTibetanBosnianBishnupriyaBugineseMin Dong ChineseCebuanoChamorroChechenChoctawCherokeeChuvashCheyenneCentral KurdishCornishCorsicanCreeCrimean TatarKashubianDimli (individual language)DhivehiLower SorbianDzongkhaModern Greek (1453-)EsperantoEweExtremaduranFijianArpitanNorthern FrisianWestern FrisianFulahFriulianGagauzGan ChineseGilakiGoan KonkaniGuaraniGujaratiHakka ChineseHaitianHausaHawaiianSerbo-CroatianHereroFiji HindiHiri MotuIgboIdoSichuan YiInuktitutInterlingueIlokoInterlingua (International Auxiliary Language Association)InupiaqJamaican Creole EnglishJavaneseLojbanKara-KalpakKabyleKalaallisutKannadaKashmiriGeorgianKanuriKabardianCentral KhmerKikuyuKinyarwandaKirghizKomiKongoKarachay-BalkarKölschKuanyamaKurdishLadinoLaoLakLezghianLigurianLimburganLingalaLombardNorthern LuriLatgalianLuxembourgishGandaMarshalleseMaithiliMalayalamEastern MariMinangkabauMacedonianMalagasyMongolianMaoriWestern MariMalay (macrolanguage)CreekMirandeseBurmeseMazanderaniNeapolitanNauruNavajoNdongaLow GermanNepali (macrolanguage)NewariNorwegian NynorskNovialNaromPediNyanjaOccitan (post 1500)Oriya (macrolanguage)OromoOssetianPangasinanPampangaPapiamentoPicardPennsylvania GermanPfaelzischPitcairn-NorfolkPaliPiemonteseWestern PanjabiPonticPushtoQuechuaVlax RomaniRomanshRusynRundiSangoYakutSicilianScotsSinhalaSamoanShonaSindhiSomaliSouthern SothoSardinianSranan TongoSwatiSaterfriesischSundaneseSwahili (macrolanguage)SilesianTahitianTatarTuluTetumTajikTigrinyaTonga (Tonga Islands)Tok PisinTswanaTsongaTurkmenTumbukaTwiTuvinianUdmurtUzbekVenetianVendaVepsVlaamsVolapükWaray (Philippines)WalloonWu ChineseKalmykXhosaMingrelianYiddishZeeuwsZhuangZulu

Variants archive 2025-07-28

  • Wikiann
  • WikiANN
  • WikiAnn NER
  • wikiann sk
  • HiNER Collapsed
  • HiNER Original
  • WikiAnn (Maltese)
  • WikiAnn (en, ko, es, pt)

8 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections