Public learning path · 3 papers

Transformer foundations

A compact route from deep residual learning to attention and the scaling behavior of language models.

For
Readers building a first rigorous mental model of modern language-model architecture.
Curator
DeepStudy editorial

Read in order. The first paper establishes the optimization context, the second introduces the architecture, and the third examines scale.

Reading order

  1. Required · Guide ready

    Deep Residual Learning for Image Recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun · 2015

    Establish the residual-learning idea that made much deeper networks practical.
  2. Required · Guide ready

    Attention Is All You Need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin · 2017

    Study the attention-only architecture and connect each equation to its stated role.
  3. Optional · Guide ready

    Scaling Laws for Neural Language Models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei · 2020

    Use the empirical scaling relationships to reason about what changes as model, data, and compute grow.