{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/thresholding-classifiers-to-maximize-f1-score","title":"Thresholding Classifiers to Maximize F1 Score","arxiv_id":"1402.1892","date":"2014-02-08","proceeding":null,"authors":["Zachary Chase Lipton","Charles Elkan","Balakrishnan Narayanaswamy"],"abstract":"This paper provides new insight into maximizing F1 scores in the context of\nbinary classification and also in the context of multilabel classification. The\nharmonic mean of precision and recall, F1 score is widely used to measure the\nsuccess of a binary classifier when one class is rare. Micro average, macro\naverage, and per instance average F1 scores are used in multilabel\nclassification. For any classifier that produces a real-valued output, we\nderive the relationship between the best achievable F1 score and the\ndecision-making threshold that achieves this optimum. As a special case, if the\nclassifier outputs are well-calibrated conditional probabilities, then the\noptimal threshold is half the optimal F1 score. As another special case, if the\nclassifier is completely uninformative, then the optimal behavior is to\nclassify all examples as positive. Since the actual prevalence of positive\nexamples typically is low, this behavior can be considered undesirable. As a\ncase study, we discuss the results, which can be surprising, of applying this\nprocedure when predicting 26,853 labels for Medline documents.","url_abs":"http://arxiv.org/abs/1402.1892v2","url_pdf":"http://arxiv.org/pdf/1402.1892v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"thresholding-classifiers-to-maximize-f1-score","repo_url":"https://github.com/Raghuvar/pneumonia_detection_using_X-rays","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1402.1892","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}