Study the paper
Scaling Laws for Neural Language Models
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Overview of Scaling Laws for Language Models
At a glance
At a glance
Scaling Laws for Neural Language Models
Sources
section
1 Introduction
Empirical studies of language model performance demonstrate that cross-entropy loss scales as a power-law with model size, dataset size, and the total compute budget. These relationships remain consistent across several orders of magnitude, showing that architectural details (such as depth versus width) have relatively minor effects compared to scale.
Sources
section
1.2 Summary of Scaling Laws
section
2.1 Parameter and Compute Scaling of Transformers