Papers › Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity

Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity

26 Oct 2023arXiv:2310.17247archive 2025-07-28

Jack Miller, Charles O'Neill, Thang Bui

In some settings neural networks exhibit a phenomenon known as \textit{grokking}, where they achieve perfect or near-perfect accuracy on the validation set long after the same performance has been achieved on the training set. In this paper, we discover that grokking is not limited to neural networks but occurs in other settings such as Gaussian process (GP) classification, GP regression, linear regression and Bayesian neural networks. We also uncover a mechanism by which to induce grokking on algorithmic datasets via the addition of dimensions containing spurious information. The presence of the phenomenon in non-neural architectures shows that grokking is not restricted to settings considered in current theoretical and empirical studies. Instead, grokking may be possible in any model where solution search is guided by complexity and error.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

jackmiller2003/tiny-gen officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

regression

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Gaussian ProcessLinear RegressionSETSGD

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections