{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-no-regret-generalization-of-hierarchical","title":"A no-regret generalization of hierarchical softmax to extreme multi-label classification","arxiv_id":"1810.11671","date":"2018-10-27","proceeding":"NeurIPS 2018 12","authors":["Marek Wydmuch","Kalina Jasinska","Mikhail Kuznetsov","Róbert Busa-Fekete","Krzysztof Dembczyński"],"abstract":"Extreme multi-label classification (XMLC) is a problem of tagging an instance\nwith a small subset of relevant labels chosen from an extremely large pool of\npossible labels. Large label spaces can be efficiently handled by organizing\nlabels as a tree, like in the hierarchical softmax (HSM) approach commonly used\nfor multi-class problems. In this paper, we investigate probabilistic label\ntrees (PLTs) that have been recently devised for tackling XMLC problems. We\nshow that PLTs are a no-regret multi-label generalization of HSM when\nprecision@k is used as a model evaluation metric. Critically, we prove that\npick-one-label heuristic - a reduction technique from multi-label to\nmulti-class that is routinely used along with HSM - is not consistent in\ngeneral. We also show that our implementation of PLTs, referred to as\nextremeText (XT), obtains significantly better results than HSM with the\npick-one-label heuristic and XML-CNN, a deep network specifically designed for\nXMLC problems. Moreover, XT is competitive to many state-of-the-art approaches\nin terms of statistical performance, model size and prediction time which makes\nit amenable to deploy in an online system.","url_abs":"http://arxiv.org/abs/1810.11671v1","url_pdf":"http://arxiv.org/pdf/1810.11671v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-no-regret-generalization-of-hierarchical","repo_url":"https://github.com/mwydmuch/extremeText","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"extreme-multi-label-classification","task_name":"Extreme Multi-Label Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"}],"methods":[{"method_slug":"hierarchical-softmax","method_name":"Hierarchical Softmax"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.11671","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}