{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ranking-distillation-learning-compact-ranking","title":"Ranking Distillation: Learning Compact Ranking Models With High Performance for Recommender System","arxiv_id":"1809.07428","date":"2018-09-19","proceeding":null,"authors":["Jiaxi Tang","Ke Wang"],"abstract":"We propose a novel way to train ranking models, such as recommender systems,\nthat are both effective and efficient. Knowledge distillation (KD) was shown to\nbe successful in image recognition to achieve both effectiveness and\nefficiency. We propose a KD technique for learning to rank problems, called\n\\emph{ranking distillation (RD)}. Specifically, we train a smaller student\nmodel to learn to rank documents/items from both the training data and the\nsupervision of a larger teacher model. The student model achieves a similar\nranking performance to that of the large teacher model, but its smaller model\nsize makes the online inference more efficient. RD is flexible because it is\northogonal to the choices of ranking models for the teacher and student. We\naddress the challenges of RD for ranking problems. The experiments on public\ndata sets and state-of-the-art recommendation models showed that RD achieves\nits design purposes: the student model learnt with RD has a model size less\nthan half of the teacher model while achieving a ranking performance similar to\nthe teacher model and much better than the student model learnt without RD.","url_abs":"http://arxiv.org/abs/1809.07428v1","url_pdf":"http://arxiv.org/pdf/1809.07428v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ranking-distillation-learning-compact-ranking","repo_url":"https://github.com/graytowne/rank_distill","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"LGPL-3.0"}}],"tasks":[{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"},{"task_slug":"recommendation-systems","task_name":"Recommendation Systems"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.07428","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}