{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-bidding-by-reinforcement-learning","title":"Real-Time Bidding by Reinforcement Learning in Display Advertising","arxiv_id":"1701.02490","date":"2017-01-10","proceeding":null,"authors":["Han Cai","Kan Ren","Wei-Nan Zhang","Kleanthis Malialis","Jun Wang","Yong Yu","Defeng Guo"],"abstract":"The majority of online display ads are served through real-time bidding (RTB)\n--- each ad display impression is auctioned off in real-time when it is just\nbeing generated from a user visit. To place an ad automatically and optimally,\nit is critical for advertisers to devise a learning algorithm to cleverly bid\nan ad impression in real-time. Most previous works consider the bid decision as\na static optimization problem of either treating the value of each impression\nindependently or setting a bid price to each segment of ad volume. However, the\nbidding for a given ad campaign would repeatedly happen during its life span\nbefore the budget runs out. As such, each bid is strategically correlated by\nthe constrained budget and the overall effectiveness of the campaign (e.g., the\nrewards from generated clicks), which is only observed after the campaign has\ncompleted. Thus, it is of great interest to devise an optimal bidding strategy\nsequentially so that the campaign budget can be dynamically allocated across\nall the available impressions on the basis of both the immediate and future\nrewards. In this paper, we formulate the bid decision process as a\nreinforcement learning problem, where the state space is represented by the\nauction information and the campaign's real-time parameters, while an action is\nthe bid price to set. By modeling the state transition via auction competition,\nwe build a Markov Decision Process framework for learning the optimal bidding\npolicy to optimize the advertising performance in the dynamic real-time bidding\nenvironment. Furthermore, the scalability problem from the large real-world\nauction volume and campaign budget is well handled by state value approximation\nusing neural networks.","url_abs":"http://arxiv.org/abs/1701.02490v2","url_pdf":"http://arxiv.org/pdf/1701.02490v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"real-time-bidding-by-reinforcement-learning","repo_url":"https://github.com/han-cai/rlb-dp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.02490","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}