{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/replicability-analysis-for-natural-language","title":"Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets","arxiv_id":"1709.09500","date":"2017-09-27","proceeding":"TACL 2017 1","authors":["Rotem Dror","Gili Baumer","Marina Bogomolov","Roi Reichart"],"abstract":"With the ever-growing amounts of textual data from a large variety of\nlanguages, domains, and genres, it has become standard to evaluate NLP\nalgorithms on multiple datasets in order to ensure consistent performance\nacross heterogeneous setups. However, such multiple comparisons pose\nsignificant challenges to traditional statistical analysis methods in NLP and\ncan lead to erroneous conclusions. In this paper, we propose a Replicability\nAnalysis framework for a statistically sound analysis of multiple comparisons\nbetween algorithms for NLP tasks. We discuss the theoretical advantages of this\nframework over the current, statistically unjustified, practice in the NLP\nliterature, and demonstrate its empirical value across four applications:\nmulti-domain dependency parsing, multilingual POS tagging, cross-domain\nsentiment classification and word similarity prediction.","url_abs":"http://arxiv.org/abs/1709.09500v1","url_pdf":"http://arxiv.org/pdf/1709.09500v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"replicability-analysis-for-natural-language","repo_url":"https://github.com/rtmdrr/replicability-analysis-NLP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"dependency-parsing","task_name":"Dependency Parsing"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"pos","task_name":"POS"},{"task_slug":"pos-tagging","task_name":"POS Tagging"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"},{"task_slug":"sentiment-classification","task_name":"Sentiment Classification"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.09500","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}