{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recogym-a-reinforcement-learning-environment","title":"RecoGym: A Reinforcement Learning Environment for the problem of Product Recommendation in Online Advertising","arxiv_id":"1808.00720","date":"2018-08-02","proceeding":null,"authors":["David Rohde","Stephen Bonner","Travis Dunlop","Flavian vasile","Alexandros Karatzoglou"],"abstract":"Recommender Systems are becoming ubiquitous in many settings and take many\nforms, from product recommendation in e-commerce stores, to query suggestions\nin search engines, to friend recommendation in social networks. Current\nresearch directions which are largely based upon supervised learning from\nhistorical data appear to be showing diminishing returns with a lot of\npractitioners report a discrepancy between improvements in offline metrics for\nsupervised learning and the online performance of the newly proposed models.\nOne possible reason is that we are using the wrong paradigm: when looking at\nthe long-term cycle of collecting historical performance data, creating a new\nversion of the recommendation model, A/B testing it and then rolling it out. We\nsee that there a lot of commonalities with the reinforcement learning (RL)\nsetup, where the agent observes the environment and acts upon it in order to\nchange its state towards better states (states with higher rewards). To this\nend we introduce RecoGym, an RL environment for recommendation, which is\ndefined by a model of user traffic patterns on e-commerce and the users\nresponse to recommendations on the publisher websites. We believe that this is\nan important step forward for the field of recommendation systems research,\nthat could open up an avenue of collaboration between the recommender systems\nand reinforcement learning communities and lead to better alignment between\noffline and online performance metrics.","url_abs":"http://arxiv.org/abs/1808.00720v2","url_pdf":"http://arxiv.org/pdf/1808.00720v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recogym-a-reinforcement-learning-environment","repo_url":"https://github.com/criteo-research/reco-gym","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"product-recommendation","task_name":"Product Recommendation"},{"task_slug":"recommendation-systems","task_name":"Recommendation Systems"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.00720","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1808.00720"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/criteo-research/reco-gym","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"84e63e1caf82f985","entry":"ff","repo":"criteo-research/reco-gym","repo_kind":"official","path":"recogym/envs/reco_env_v1.py","file_url":"https://github.com/criteo-research/reco-gym/blob/HEAD/recogym/envs/reco_env_v1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"84e63e1caf82f985"}},{"code_sha256_prefix":"ef2197d5a69de56d","entry":"sig","repo":"criteo-research/reco-gym","repo_kind":"official","path":"recogym/envs/reco_env_v1.py","file_url":"https://github.com/criteo-research/reco-gym/blob/HEAD/recogym/envs/reco_env_v1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ef2197d5a69de56d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}