Study the paper
Scaling Laws for Neural Language Models
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Optimal Allocation of Compute Budgets
Optimal model size scaling with compute
Optimal model size scaling with compute
The optimal model size scales with the minimum compute budget as a power-law, where the exponent is approximately .
Sources
S6.E1
N(Cmin)∝(Cmin)0.73.proportional-to𝑁subscript𝐶minsuperscriptsubscript𝐶min0.73N(C_{\rm min})\propto(C_{\rm min})^{0.73}. (6.1)
N(C_{\rm min})\propto(C_{\rm min})^{0.73}.| Measure | Value |
|---|
Sources
S6.E1
N(Cmin)∝(Cmin)0.73.proportional-to𝑁subscript𝐶minsuperscriptsubscript𝐶min0.73N(C_{\rm min})\propto(C_{\rm min})^{0.73}. (6.1)
N(C_{\rm min})\propto(C_{\rm min})^{0.73}.