Study the paper

Scaling Laws for Neural Language Models

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Optimal Allocation of Compute Budget

Figure 14: Optimal Allocation of Compute

Figure 14: Optimal Allocation of Compute

Each value of the compute budget CminC_{\rm min} has an associated optimal model size NN. Optimal model size grows very rapidly with CminC_{\rm min}, increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x.

Sources

S6.F14

Figure 14: Left: Each value of the compute budget Cminsubscript𝐶minC_{\rm min} has an associated optimal model size N𝑁N. Optimal model size grows very rapidly with Cminsubscript𝐶minC_{\rm min}, increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x. Right: The batch-adjusted number of optimization steps also grows very slowly, if at all, meaning that most of the growth in data examples processed can be used for increased batch sizes.
Reported values
MeasureValue
Sources

S6.F14

Figure 14: Left: Each value of the compute budget Cminsubscript𝐶minC_{\rm min} has an associated optimal model size N𝑁N. Optimal model size grows very rapidly with Cminsubscript𝐶minC_{\rm min}, increasing by 5x for each 10x increase in compute. The number of data examples processed makes up the remainder of the increase, growing relatively modestly by only 2x. Right: The batch-adjusted number of optimization steps also grows very slowly, if at all, meaning that most of the growth in data examples processed can be used for increased batch sizes.