{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lener-br-a-dataset-for-named-entity","title":"LeNER-Br: a Dataset for Named Entity Recognition in Brazilian Legal Text","arxiv_id":null,"date":"2018-09-24","proceeding":"International Conference on the Computational Processing of Portuguese (PROPOR) 2018 9","authors":["Pedro H. Luz de Araujo","Teófilo E. de Campos","Renato R. R. de Oliveira","Matheus Stauffer","Samuel Couto","Paulo Bermejo"],"abstract":"Named entity recognition systems have the untapped potential to extract information from legal documents, which can improve\r\ninformation retrieval and decision-making processes. In this paper, a dataset for named entity recognition in Brazilian legal documents is presented. Unlike other Portuguese language datasets, this dataset is composed entirely of legal documents. In addition to tags for persons, locations, time entities and organizations, the dataset contains specific tags for law and legal cases entities. To establish a set of baseline results, we first performed experiments on another Portuguese dataset: Paramopama. This evaluation demonstrate that LSTM-CRF gives results that are significantly better than those previously reported. We then retrained LSTM-CRF, on our dataset and obtained F 1 scores of 97.04% and 88.82% for Legislation and Legal case entities, respectively.\r\nThese results show the viability of the proposed dataset for legal applications.","url_abs":"https://cic.unb.br/~teodecampos/LeNER-Br/","url_pdf":"https://cic.unb.br/~teodecampos/LeNER-Br/luz_etal_propor2018.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lener-br-a-dataset-for-named-entity","repo_url":"https://github.com/peluz/lener-br","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[{"slug":"lener-br","name":"LeNER-Br","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/named-entity-recognition-on-lener-br","task":"Named Entity Recognition (NER)","dataset":"LeNER-Br","model":"LSTM-CRF","rank_in_archive_order":1,"of":1,"metrics":{"Micro F1 (Exact Span)":"0.8661","Micro F1 (Tokens)":"0.9253"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}