Papers/2609.10657
🧪 Test?View on arXiv

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

Not provided in the abstract

grokkinggeneralizationneural networkshyperparameter tuning
2609.10657
Builder Relevance
70%
1h ago

Abstract

This paper investigates the transition from memorization to generalization in neural networks, specifically focusing on the phenomenon known as grokking.

Reality Card

Core Claim

The study quantitatively characterizes the memorization-to-generalization boundary in hyperparameter space, revealing that data complexity is the dominant factor influencing the transition.

Method / Result

The power-law scaling relation for generalization onset time is given by: T_grok ∝ H^{-0.27} D^{-2.04} η^{-0.50} λ^{-0.64}, with R^2 = 0.732.

Limitations

The study focuses on a specific architecture (two-hidden-layer MLPs) and modular arithmetic, which may limit generalizability to other architectures or tasks.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers