Papers › Addressing reward bias in Adversarial Imitation Learning with neutral reward functions

Addressing reward bias in Adversarial Imitation Learning with neutral reward functions

20 Sep 2020arXiv:2009.09467archive 2025-07-28

Rohit Jena, Siddharth Agrawal, Katia Sycara

Generative Adversarial Imitation Learning suffers from the fundamental problem of reward bias stemming from the choice of reward functions used in the algorithm. Different types of biases also affect different types of environments - which are broadly divided into survival and task-based environments. We provide a theoretical sketch of why existing reward functions would fail in imitation learning scenarios in task based environments with multiple terminal states. We also propose a new reward function for GAIL which outperforms existing GAIL methods on task based environments with single and multiple terminal states and effectively overcomes both survival and termination bias.

PaperPDFCode

Code

rohitrango/Reward-bias-in-GAIL mentioned on GitHubtfMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Imitation Learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

GAIL

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections