Papers › Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and...

Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling

5 Oct 2019ICML 2020 1arXiv:1910.02208archive 2025-07-28

Che Wang, Yanqiu Wu, Quan Vuong, Keith Ross

We aim to develop off-policy DRL algorithms that not only exceed state-of-the-art performance but are also simple and minimalistic. For standard continuous control benchmarks, Soft Actor-Critic (SAC), which employs entropy maximization, currently provides state-of-the-art performance. We first demonstrate that the entropy term in SAC addresses action saturation due to the bounded nature of the action spaces, with this insight, we propose a streamlined algorithm with a simple normalization scheme or with inverted gradients. We show that both approaches can match SAC's sample efficiency performance without the need of entropy maximization, we then propose a simple non-uniform sampling method for selecting transitions from the replay buffer during training. Extensive experimental results demonstrate that our proposed sampling scheme leads to state of the art sample efficiency on challenging continuous control tasks. We combine all of our findings into one simple algorithm, which we call Streamlined Off Policy with Emphasizing Recent Experience, for which we provide robust public-domain code.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

AutumnWu/Streamlined-Off-Policy-Learning officialmentioned in papertf report
Fable67/Streamlined-Off-Policy-Learning mentioned on GitHubpytorchMIT report
angelolovatto/raylab mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Continuous Controlcontinuous-control

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AdamDense ConnectionsExperience ReplayReLUSoft Actor Critic

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections