Public learning path · 3 papers
Transformer foundations
A compact route from deep residual learning to attention and the scaling behavior of language models.
- For
- Readers building a first rigorous mental model of modern language-model architecture.
- Curator
- DeepStudy editorial
Read in order. The first paper establishes the optimization context, the second introduces the architecture, and the third examines scale.
Reading order
Required · Guide ready
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015
Establish the residual-learning idea that made much deeper networks practical.
Required · Guide ready
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin · 2017
Study the attention-only architecture and connect each equation to its stated role.
Optional · Guide ready
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020
Use the empirical scaling relationships to reason about what changes as model, data, and compute grow.