{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bengal-an-automatic-benchmark-generator-for","title":"BENGAL: An Automatic Benchmark Generator for Entity Recognition and Linking","arxiv_id":"1710.08691","date":"2017-10-24","proceeding":"WS 2018 11","authors":["Axel-Cyrille Ngonga Ngomo","Michael Röder","Diego Moussallem","Ricardo Usbeck","René Speck"],"abstract":"The manual creation of gold standards for named entity recognition and entity\nlinking is time- and resource-intensive. Moreover, recent works show that such\ngold standards contain a large proportion of mistakes in addition to being\ndifficult to maintain. We hence present BENGAL, a novel automatic generation of\nsuch gold standards as a complement to manually created benchmarks. The main\nadvantage of our benchmarks is that they can be readily generated at any time.\nThey are also cost-effective while being guaranteed to be free of annotation\nerrors. We compare the performance of 11 tools on benchmarks in English\ngenerated by BENGAL and on 16benchmarks created manually. We show that our\napproach can be ported easily across languages by presenting results achieved\nby 4 tools on both Brazilian Portuguese and Spanish. Overall, our results\nsuggest that our automatic benchmark generation approach can create varied\nbenchmarks that have characteristics similar to those of existing benchmarks.\nOur approach is open-source. Our experimental results are available at\nhttp://faturl.com/bengalexpinlg and the code at\nhttps://github.com/dice-group/BENGAL.","url_abs":"http://arxiv.org/abs/1710.08691v3","url_pdf":"http://arxiv.org/pdf/1710.08691v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bengal-an-automatic-benchmark-generator-for","repo_url":"https://github.com/dice-group/BENGAL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"entity-linking","task_name":"Entity Linking"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}