{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-reinforcement-learning-of-marked","title":"Deep Reinforcement Learning of Marked Temporal Point Processes","arxiv_id":"1805.09360","date":"2018-05-23","proceeding":"NeurIPS 2018 12","authors":["Utkarsh Upadhyay","Abir De","Manuel Gomez-Rodriguez"],"abstract":"In a wide variety of applications, humans interact with a complex environment\nby means of asynchronous stochastic discrete events in continuous time. Can we\ndesign online interventions that will help humans achieve certain goals in such\nasynchronous setting? In this paper, we address the above problem from the\nperspective of deep reinforcement learning of marked temporal point processes,\nwhere both the actions taken by an agent and the feedback it receives from the\nenvironment are asynchronous stochastic discrete events characterized using\nmarked temporal point processes. In doing so, we define the agent's policy\nusing the intensity and mark distribution of the corresponding process and then\nderive a flexible policy gradient method, which embeds the agent's actions and\nthe feedback it receives into real-valued vectors using deep recurrent neural\nnetworks. Our method does not make any assumptions on the functional form of\nthe intensity and mark distribution of the feedback and it allows for\narbitrarily complex reward functions. We apply our methodology to two different\napplications in personalized teaching and viral marketing and, using data\ngathered from Duolingo and Twitter, we show that it may be able to find\ninterventions to help learners and marketers achieve their goals more\neffectively than alternatives.","url_abs":"http://arxiv.org/abs/1805.09360v2","url_pdf":"http://arxiv.org/pdf/1805.09360v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-reinforcement-learning-of-marked","repo_url":"https://github.com/Networks-Learning/tpprl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"marketing","task_name":"Marketing"},{"task_slug":"point-processes","task_name":"Point Processes"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.09360","atlas_url":"https://app.syntology.ai/?focus=1805.09360","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}