Papers › Self-supervised network distillation: an effective approach to exploration in sparse...

Self-supervised network distillation: an effective approach to exploration in sparse reward environments

22 Feb 2023arXiv:2302.11563archive 2025-07-28

Matej Pecháč, Michal Chovanec, Igor Farkaš

Reinforcement learning can solve decision-making problems and train an agent to behave in an environment according to a predesigned reward function. However, such an approach becomes very problematic if the reward is too sparse and so the agent does not come across the reward during the environmental exploration. The solution to such a problem may be to equip the agent with an intrinsic motivation that will provide informed exploration during which the agent is likely to also encounter external reward. Novelty detection is one of the promising branches of intrinsic motivation research. We present Self-supervised Network Distillation (SND), a class of intrinsic motivation algorithms based on the distillation error as a novelty indicator, where the predictor model and the target model are both trained. We adapted three existing self-supervised methods for this purpose and experimentally tested them on a set of ten environments that are considered difficult to explore. The results show that our approach achieves faster growth and higher external reward for the same training time compared to the baseline models, which implies improved exploration in a very sparse reward environment. In addition, the analytical methods we applied provide valuable explanatory insights into our proposed models.

PaperPDFCode

Code

iskandor/snd officialmentioned in paperpytorch report
michalnand/reinforcement_learning mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Atari GamesDecision MakingNovelty DetectionReinforcement Learning (RL)Self-Supervised Learningreinforcement-learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Atari Games Atari 2600 Gravitar SND-VIC Score 6712 #5 of 53 Archive leaderboard report
Atari Games Atari 2600 Gravitar SND-STD Score 4643 #10 of 53 Archive leaderboard report
Atari Games Atari 2600 Gravitar SND-V Score 2741 #14 of 53 Archive leaderboard report
Atari Games Atari 2600 Montezuma's Revenge SND-V Score 21565 #3 of 50 Archive leaderboard report
Atari Games Atari 2600 Montezuma's Revenge SND-VIC Score 7838 #6 of 50 Archive leaderboard report
Atari Games Atari 2600 Montezuma's Revenge SND-STD Score 7212 #7 of 50 Archive leaderboard report
Atari Games Atari 2600 Pitfall! SND-V Score 0 #17 of 23 Archive leaderboard report
Atari Games Atari 2600 Pitfall! SND-VIC Score 0 #18 of 23 Archive leaderboard report
Atari Games Atari 2600 Private Eye SND-VIC Score 17313 #3 of 52 Archive leaderboard report
Atari Games Atari 2600 Private Eye SND-STD Score 15089 #9 of 52 Archive leaderboard report
Atari Games Atari 2600 Private Eye SND-V Score 4213 #15 of 52 Archive leaderboard report
Atari Games Atari 2600 Solaris SND-STD Score 12460 #3 of 23 Archive leaderboard report
Atari Games Atari 2600 Solaris SND-VIC Score 11865 #4 of 23 Archive leaderboard report
Atari Games Atari 2600 Solaris SND-V Score 11582 #5 of 23 Archive leaderboard report
Atari Games Atari 2600 Venture SND-VIC Score 2188 #3 of 55 Archive leaderboard report
Atari Games Atari 2600 Venture SND-STD Score 2138 #4 of 55 Archive leaderboard report
Atari Games Atari 2600 Venture SND-V Score 1787 #11 of 55 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections