{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-multilingual-named-entity-recognition","title":"Robust Multilingual Named Entity Recognition with Shallow Semi-Supervised Features","arxiv_id":"1701.09123","date":"2017-01-31","proceeding":null,"authors":["Rodrigo Agerri","German Rigau"],"abstract":"We present a multilingual Named Entity Recognition approach based on a robust\nand general set of features across languages and datasets. Our system combines\nshallow local information with clustering semi-supervised features induced on\nlarge amounts of unlabeled text. Understanding via empirical experimentation\nhow to effectively combine various types of clustering features allows us to\nseamlessly export our system to other datasets and languages. The result is a\nsimple but highly competitive system which obtains state of the art results\nacross five languages and twelve datasets. The results are reported on standard\nshared task evaluation data such as CoNLL for English, Spanish and Dutch.\nFurthermore, and despite the lack of linguistically motivated features, we also\nreport best results for languages such as Basque and German. In addition, we\ndemonstrate that our method also obtains very competitive results even when the\namount of supervised data is cut by half, alleviating the dependency on\nmanually annotated data. Finally, the results show that our emphasis on\nclustering features is crucial to develop robust out-of-domain models. The\nsystem and models are freely available to facilitate its use and guarantee the\nreproducibility of results.","url_abs":"http://arxiv.org/abs/1701.09123v1","url_pdf":"http://arxiv.org/pdf/1701.09123v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-multilingual-named-entity-recognition","repo_url":"https://github.com/ixa-ehu/ixa-pipe-nerc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"multilingual-named-entity-recognition","task_name":"Multilingual Named Entity Recognition"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/named-entity-recognition-ner-on-conll-2003","task":"Named Entity Recognition (NER)","dataset":"CoNLL 2003 (English)","model":"IXA pipes","rank_in_archive_order":63,"of":73,"metrics":{"F1":"91.36"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.09123","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}