{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dcm-bandits-learning-to-rank-with-multiple","title":"DCM Bandits: Learning to Rank with Multiple Clicks","arxiv_id":"1602.03146","date":"2016-02-09","proceeding":null,"authors":["Sumeet Katariya","Branislav Kveton","Csaba Szepesvári","Zheng Wen"],"abstract":"A search engine recommends to the user a list of web pages. The user examines\nthis list, from the first page to the last, and clicks on all attractive pages\nuntil the user is satisfied. This behavior of the user can be described by the\ndependent click model (DCM). We propose DCM bandits, an online learning variant\nof the DCM where the goal is to maximize the probability of recommending\nsatisfactory items, such as web pages. The main challenge of our learning\nproblem is that we do not observe which attractive item is satisfactory. We\npropose a computationally-efficient learning algorithm for solving our problem,\ndcmKL-UCB; derive gap-dependent upper bounds on its regret under reasonable\nassumptions; and also prove a matching lower bound up to logarithmic factors.\nWe evaluate our algorithm on synthetic and real-world problems, and show that\nit performs well even when our model is misspecified. This work presents the\nfirst practical and regret-optimal online algorithm for learning to rank with\nmultiple clicks in a cascade-like click model.","url_abs":"http://arxiv.org/abs/1602.03146v2","url_pdf":"http://arxiv.org/pdf/1602.03146v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dcm-bandits-learning-to-rank-with-multiple","repo_url":"https://github.com/wchen408/4803RA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1602.03146","atlas_url":"https://app.syntology.ai/?focus=1602.03146","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}