{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-the-vector-space-model-sparse","title":"Revisiting the Vector Space Model: Sparse Weighted Nearest-Neighbor Method for Extreme Multi-Label Classification","arxiv_id":"1802.03938","date":"2018-02-12","proceeding":null,"authors":["Tatsuhiro Aoshima","Kei Kobayashi","Mihoko Minami"],"abstract":"Machine learning has played an important role in information retrieval (IR)\nin recent times. In search engines, for example, query keywords are accepted\nand documents are returned in order of relevance to the given query; this can\nbe cast as a multi-label ranking problem in machine learning. Generally, the\nnumber of candidate documents is extremely large (from several thousand to\nseveral million); thus, the classifier must handle many labels. This problem is\nreferred to as extreme multi-label classification (XMLC). In this paper, we\npropose a novel approach to XMLC termed the Sparse Weighted Nearest-Neighbor\nMethod. This technique can be derived as a fast implementation of\nstate-of-the-art (SOTA) one-versus-rest linear classifiers for very sparse\ndatasets. In addition, we show that the classifier can be written as a sparse\ngeneralization of a representer theorem with a linear kernel. Furthermore, our\nmethod can be viewed as the vector space model used in IR. Finally, we show\nthat the Sparse Weighted Nearest-Neighbor Method can process data points in\nreal time on XMLC datasets with equivalent performance to SOTA models, with a\nsingle thread and smaller storage footprint. In particular, our method exhibits\nsuperior performance to the SOTA models on a dataset with 3 million labels.","url_abs":"http://arxiv.org/abs/1802.03938v1","url_pdf":"http://arxiv.org/pdf/1802.03938v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"revisiting-the-vector-space-model-sparse","repo_url":"https://github.com/hiro4bbh/sticker","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"extreme-multi-label-classification","task_name":"Extreme Multi-Label Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}