{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reinforcement-learning-to-rank-in-e-commerce","title":"Reinforcement Learning to Rank in E-Commerce Search Engine: Formalization, Analysis, and Application","arxiv_id":"1803.00710","date":"2018-03-02","proceeding":null,"authors":["Yujing Hu","Qing Da","An-Xiang Zeng","Yang Yu","Yinghui Xu"],"abstract":"In e-commerce platforms such as Amazon and TaoBao, ranking items in a search\nsession is a typical multi-step decision-making problem. Learning to rank (LTR)\nmethods have been widely applied to ranking problems. However, such methods\noften consider different ranking steps in a session to be independent, which\nconversely may be highly correlated to each other. For better utilizing the\ncorrelation between different ranking steps, in this paper, we propose to use\nreinforcement learning (RL) to learn an optimal ranking policy which maximizes\nthe expected accumulative rewards in a search session. Firstly, we formally\ndefine the concept of search session Markov decision process (SSMDP) to\nformulate the multi-step ranking problem. Secondly, we analyze the property of\nSSMDP and theoretically prove the necessity of maximizing accumulative rewards.\nLastly, we propose a novel policy gradient algorithm for learning an optimal\nranking policy, which is able to deal with the problem of high reward variance\nand unbalanced reward distribution of an SSMDP. Experiments are conducted in\nsimulation and TaoBao search engine. The results demonstrate that our algorithm\nperforms much better than online LTR methods, with more than 40% and 30% growth\nof total transaction amount in the simulation and the real application,\nrespectively.","url_abs":"http://arxiv.org/abs/1803.00710v3","url_pdf":"http://arxiv.org/pdf/1803.00710v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reinforcement-learning-to-rank-in-e-commerce","repo_url":"https://github.com/UnibucProjects/DeepRLRecommenderSystem","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.00710","atlas_url":"https://app.syntology.ai/?focus=1803.00710","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}