{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pioner-datasets-and-baselines-for-armenian","title":"pioNER: Datasets and Baselines for Armenian Named Entity Recognition","arxiv_id":"1810.08699","date":"2018-10-19","proceeding":null,"authors":["Tsolak Ghukasyan","Garnik Davtyan","Karen Avetisyan","Ivan Andrianov"],"abstract":"In this work, we tackle the problem of Armenian named entity recognition,\nproviding silver- and gold-standard datasets as well as establishing baseline\nresults on popular models. We present a 163000-token named entity corpus\nautomatically generated and annotated from Wikipedia, and another 53400-token\ncorpus of news sentences with manual annotation of people, organization and\nlocation named entities. The corpora were used to train and evaluate several\npopular named entity recognition models. Alongside the datasets, we release\n50-, 100-, 200-, 300-dimensional GloVe word embeddings trained on a collection\nof Armenian texts from Wikipedia, news, blogs, and encyclopedia.","url_abs":"http://arxiv.org/abs/1810.08699v1","url_pdf":"http://arxiv.org/pdf/1810.08699v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pioner-datasets-and-baselines-for-armenian","repo_url":"https://github.com/ispras-texterra/pioner","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"pioner-datasets-and-baselines-for-armenian","repo_url":"https://github.com/vt257/allnews-am","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[{"method_slug":"glove","method_name":"GloVe"}],"datasets_introduced":[{"slug":"pioner","name":"pioNER","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.08699","atlas_url":"https://app.syntology.ai/?focus=1810.08699","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}