Papers › Learning Reward Machines for Partially Observable Reinforcement Learning

Learning Reward Machines for Partially Observable Reinforcement Learning

1 Dec 2019NeurIPS 2019 12archive 2025-07-28

Rodrigo Toro Icarte, Ethan Waldie, Toryn Klassen, Rick Valenzano, Margarita Castro, Sheila McIlraith

Reward Machines (RMs), originally proposed for specifying problems in Reinforcement Learning (RL), provide a structured, automata-based representation of a reward function that allows an agent to decompose problems into subproblems that can be efficiently learned using off-policy learning. Here we show that RMs can be learned from experience, instead of being specified by the user, and that the resulting problem decomposition can be used to effectively solve partially observable RL problems. We pose the task of learning RMs as a discrete optimization problem where the objective is to find an RM that decomposes the problem into a set of subproblems such that the combination of their optimal memoryless policies is an optimal policy for the original problem. We show the effectiveness of this approach on three partially observable domains, where it significantly outperforms A3C, PPO, and ACER, and discuss its advantages, limitations, and broader potential.

PaperPDFCode

Code

bitbucket.org/RToroIcarte/lrm officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Partially Observable Reinforcement LearningProblem DecompositionReinforcement LearningReinforcement Learning (RL)reinforcement-learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

A3CACERConvolutionDense ConnectionsEntropy RegularizationExperience ReplayPPOReLURetraceSoftmaxStochastic Dueling NetworkTRPO

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections