Study the paper
Scaling Laws for Neural Language Models
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Optimal Allocation of Compute Budget
Figure 14: Optimal Allocation of Compute
Figure 14: Optimal Allocation of Compute
Each value of the compute budget has an associated optimal model size . Optimal model size grows very rapidly with , increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x.
Sources
S6.F14
Figure 14: Left: Each value of the compute budget Cminsubscript𝐶minC_{\rm min} has an associated optimal model size N𝑁N. Optimal model size grows very rapidly with Cminsubscript𝐶minC_{\rm min}, increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x. Right: The batch-adjusted number of optimization steps also grows very slowly, if at all, meaning that most of the growth in data examples processed can be used for increased batch sizes.
| Measure | Value |
|---|
Sources
S6.F14
Figure 14: Left: Each value of the compute budget Cminsubscript𝐶minC_{\rm min} has an associated optimal model size N𝑁N. Optimal model size grows very rapidly with Cminsubscript𝐶minC_{\rm min}, increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x. Right: The batch-adjusted number of optimization steps also grows very slowly, if at all, meaning that most of the growth in data examples processed can be used for increased batch sizes.