{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-weight-for-text-classification","title":"Learning to Weight for Text Classification","arxiv_id":"1903.12090","date":"2019-03-28","proceeding":null,"authors":["Alejandro Moreo Fernández","Andrea Esuli","Fabrizio Sebastiani"],"abstract":"In information retrieval (IR) and related tasks, term weighting approaches\ntypically consider the frequency of the term in the document and in the\ncollection in order to compute a score reflecting the importance of the term\nfor the document. In tasks characterized by the presence of training data (such\nas text classification) it seems logical that the term weighting function\nshould take into account the distribution (as estimated from training data) of\nthe term across the classes of interest. Although `supervised term weighting'\napproaches that use this intuition have been described before, they have failed\nto show consistent improvements. In this article we analyse the possible\nreasons for this failure, and call consolidated assumptions into question.\nFollowing this criticism we propose a novel supervised term weighting approach\nthat, instead of relying on any predefined formula, learns a term weighting\nfunction optimised on the training set of interest; we dub this approach\n\\emph{Learning to Weight} (LTW). The experiments that we run on several\nwell-known benchmarks, and using different learning methods, show that our\nmethod outperforms previous term weighting approaches in text classification.","url_abs":"http://arxiv.org/abs/1903.12090v1","url_pdf":"http://arxiv.org/pdf/1903.12090v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-weight-for-text-classification","repo_url":"https://github.com/AlexMoreo/learning-to-weight","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}