Papers › Deep Reinforcement Learning with Function Properties in Mean Reversion Strategies

Deep Reinforcement Learning with Function Properties in Mean Reversion Strategies

9 Jan 2021arXiv:2101.03418archive 2025-07-28

Sophia Gu

Over the past decades, researchers have been pushing the limits of Deep Reinforcement Learning (DRL). Although DRL has attracted substantial interest from practitioners, many are blocked by having to search through a plethora of available methodologies that are seemingly alike, while others are still building RL agents from scratch based on classical theories. To address the aforementioned gaps in adopting the latest DRL methods, I am particularly interested in testing out if any of the recent technology developed by the leads in the field can be readily applied to a class of optimal trading problems. Unsurprisingly, many prominent breakthroughs in DRL are investigated and tested on strategic games: from AlphaGo to AlphaStar and at about the same time, OpenAI Five. Thus, in this writing, I want to show precisely how to use a DRL library that is initially built for games in a fundamental trading problem; mean reversion. And by introducing a framework that incorporates economically-motivated function properties, I also demonstrate, through the library, a highly-performant and convergent DRL solution to decision-making financial problems in general.

PaperPDFCode

Code

sophiagu/RLF officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Decision MakingDeep Reinforcement LearningReinforcement LearningReinforcement Learning (RL)reinforcement-learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

1x1 ConvolutionAMAbsolute Position EncodingsAccumulating Eligibility TraceAdamAlphaStarAttentionAverage PoolingBPEBatch NormalizationBottleneck Residual BlockConvolutionDense ConnectionsDropoutFeedforward NetworkGated Linear UnitGlobal Average PoolingKaiming InitializationLSTMLabel SmoothingLayer NormalizationLinear LayerMax PoolingMulti-Head AttentionPointer NetworkPosition-Wise Feed-Forward LayerReLUResidual BlockResidual ConnectionScatter ConnectionSigmoid ActivationSoftmaxTD LambdaTanh ActivationTransformerV-trace

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections