{"url":"/method/fisher-brc","slug":"fisher-brc","name":"Fisher-BRC","full_name":"Fisher-BRC","full_name_withheld":false,"description_markdown":"**Fisher-BRC** is an actor critic algorithm for offline reinforcement learning that encourages the learned policy to stay close to the data, namely parameterizing the critic as the $\\log$-behavior-policy, which generated the offline dataset, plus a state-action value offset term, which can be learned using a neural network. Behavior regularization then corresponds to an appropriate regularizer on the offset term. A gradient penalty regularizer is used for the offset term, which is equivalent to Fisher divergence regularization, suggesting connections to the score matching and generative energy-based model literature.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2103.08050v1","title":"Offline Reinforcement Learning with Fisher Divergence Critic Regularization","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Policy Gradient Methods","url":"/methods/category/policy-gradient-methods","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/offline-reinforcement-learning-with-fisher","title":"Offline Reinforcement Learning with Fisher Divergence Critic Regularization","date":"2021-03-14","arxiv_id":"2103.08050","n_code_links":2,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/offline-rl","name":"Offline RL","papers":1},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":1},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":1},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/fisher-brc"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}