{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/peyma-a-tagged-corpus-for-persian-named","title":"PEYMA: A Tagged Corpus for Persian Named Entities","arxiv_id":"1801.09936","date":"2018-01-30","proceeding":null,"authors":["Mahsa Sadat Shahshahani","Mahdi Mohseni","Azadeh Shakery","Heshaam Faili"],"abstract":"The goal in the NER task is to classify proper nouns of a text into classes\nsuch as person, location, and organization. This is an important preprocessing\nstep in many NLP tasks such as question-answering and summarization. Although\nmany research studies have been conducted in this area in English and the\nstate-of-the-art NER systems have reached performances of higher than 90\npercent in terms of F1 measure, there are very few research studies for this\ntask in Persian. One of the main important causes of this may be the lack of a\nstandard Persian NER dataset to train and test NER systems. In this research we\ncreate a standard, big-enough tagged Persian NER dataset which will be\ndistributed for free for research purposes. In order to construct such a\nstandard dataset, we studied standard NER datasets which are constructed for\nEnglish researches and found out that almost all of these datasets are\nconstructed using news texts. So we collected documents from ten news websites.\nLater, in order to provide annotators with some guidelines to tag these\ndocuments, after studying guidelines used for constructing CoNLL and MUC\nstandard English datasets, we set our own guidelines considering the Persian\nlinguistic rules.","url_abs":"http://arxiv.org/abs/1801.09936v1","url_pdf":"http://arxiv.org/pdf/1801.09936v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"peyma-a-tagged-corpus-for-persian-named","repo_url":"https://github.com/Pirata-Codex/Tag-Persian-Entities-Using-Bert","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"cg","task_name":"NER"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[{"slug":"peyma","name":"PEYMA","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.09936","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}