{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/identification-of-tasks-datasets-evaluation","title":"Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction","arxiv_id":"1906.09317","date":"2019-06-21","proceeding":"ACL 2019 7","authors":["Yufang Hou","Charles Jochim","Martin Gleize","Francesca Bonin","Debasis Ganguly"],"abstract":"While the fast-paced inception of novel tasks and new datasets helps foster active research in a community towards interesting directions, keeping track of the abundance of research activity in different areas on different datasets is likely to become increasingly difficult. The community could greatly benefit from an automatic system able to summarize scientific results, e.g., in the form of a leaderboard. In this paper we build two datasets and develop a framework (TDMS-IE) aimed at automatically extracting task, dataset, metric and score from NLP papers, towards the automatic construction of leaderboards. Experiments show that our model outperforms several baselines by a large margin. Our model is a first step towards automatic leaderboard construction, e.g., in the NLP domain.","url_abs":"https://arxiv.org/abs/1906.09317v1","url_pdf":"https://arxiv.org/pdf/1906.09317v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"identification-of-tasks-datasets-evaluation","repo_url":"https://github.com/IBM/science-result-extractor","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"scientific-results-extraction","task_name":"Scientific Results Extraction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scientific-results-extraction-on-nlp-tdms-exp","task":"Scientific Results Extraction","dataset":"NLP-TDMS (Exp, arXiv only)","model":"TDMS-IE","rank_in_archive_order":2,"of":2,"metrics":{"Macro F1":"8.8","Macro Precision":"9.5","Macro Recall":"8.6","Micro F1":"7.5","Micro Precision":"6.8","Micro Recall":"8.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1906.09317","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1906.09317"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IBM/science-result-extractor","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1b46c9ce24fc6db3","entry":"file_based_input_fn_builder","repo":"IBM/science-result-extractor","repo_kind":"official","path":"bert_tdms/run_classifier_sci.py","file_url":"https://github.com/IBM/science-result-extractor/blob/HEAD/bert_tdms/run_classifier_sci.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1b46c9ce24fc6db3"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}