You are reading immutable version 2. The current guide may be newer.

Study the paper

Scaling Laws for Neural Language Models

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Overview of Scaling Laws for Language Models

At a glance

At a glance

Scaling Laws for Neural Language Models

Sources

section

1 Introduction

Empirical studies of language model performance demonstrate that cross-entropy loss scales as a power-law with model size, dataset size, and the total compute budget. These relationships remain consistent across several orders of magnitude, showing that architectural details (such as depth versus width) have relatively minor effects compared to scale.

Sources

section

1.2 Summary of Scaling Laws

section

2.1 Parameter and Compute Scaling of Transformers