Study the paper
Scaling Laws for Neural Language Models
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Optimal Allocation of Compute Budget
Figure 9: Overfitting and Dataset Size
Figure 9: Overfitting and Dataset Size
The early-stopped test loss depends predictably on the dataset size and model size . For large , performance is a straight power law in . For a smaller fixed , performance stops improving as increases and the model begins to overfit. The extent of overfitting depends predominantly on the ratio .
Sources
S4.F9
Figure 9: The early-stopped test loss L(N,D)𝐿𝑁𝐷L(N,D) depends predictably on the dataset size D𝐷D and model size N𝑁N according to Equation (1.5). Left: For large D𝐷D, performance is a straight power law in N𝑁N. For a smaller fixed D𝐷D, performance stops improving as N𝑁N increases and the model begins to overfit. (The reverse is also true, see Figure 4.) Right: The extent of overfitting depends predominantly on the ratio NαNαD/Dsuperscript𝑁subscript𝛼𝑁subscript𝛼𝐷𝐷N^{\frac{\alpha_{N}}{\alpha_{D}}}/D, as predicted in equation (4.3). The line is our fit to that equation.
| Measure | Value |
|---|
Sources
S4.F9
Figure 9: The early-stopped test loss L(N,D)𝐿𝑁𝐷L(N,D) depends predictably on the dataset size D𝐷D and model size N𝑁N according to Equation (1.5). Left: For large D𝐷D, performance is a straight power law in N𝑁N. For a smaller fixed D𝐷D, performance stops improving as N𝑁N increases and the model begins to overfit. (The reverse is also true, see Figure 4.) Right: The extent of overfitting depends predominantly on the ratio NαNαD/Dsuperscript𝑁subscript𝛼𝑁subscript𝛼𝐷𝐷N^{\frac{\alpha_{N}}{\alpha_{D}}}/D, as predicted in equation (4.3). The line is our fit to that equation.