You are reading immutable version 28. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

21 activities

Lesson

At a glance

Introduction to the Transformer Architecture

Introduction to the Transformer Architecture

Lesson

Limitations of Recurrent and Convolutional Architectures

Introduction to the Transformer Architecture

The Sequential Bottleneck of Recurrent Architectures

Lesson

Scaled Dot-Product Attention

Scaled Dot-Product Attention

Mathematical Formulation

Lesson

Scaled Dot-Product Attention

Scaled Dot-Product Attention

The Scaled Dot-Product Attention mechanism computes attention weights by taking the dot product of queries with keys, scaling them by the square root of the key dimension, applying a softmax function, and using the resulting weights to compute a weighted sum of the values.

Visual

How attention redistributes weight

Scaled Dot-Product Attention

Illustrative attention weights computed from a fixed teaching example; these are not paper-reported values.

Visual

Explore attention score scaling

Scaled Dot-Product Attention

Move across fixed illustrative scenarios to see how score scaling changes concentration.

Visual

Scaled Dot-Product Attention: step by step

Scaled Dot-Product Attention

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Multi-Head Attention Mechanism

Multi-Head Attention Mechanism

Multi-Head Attention Mechanism

Lesson

Multi-Head Attention

Multi-Head Attention Mechanism

Multi-Head Attention allows the model to jointly attend to information from different representation subspaces at different positions. Instead of performing a single attention function with queries, keys, and values, the queries, keys, and values are projected multiple times with different, learned linear projections.

Visual

Multi-Head Attention: step by step

Multi-Head Attention Mechanism

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Injecting Order with Positional Encoding

Injecting Order with Positional Encoding

Injecting Order with Positional Encodings

Lesson

Sinusoidal Positional Encoding (Even)

Injecting Order with Positional Encoding

In sequence transduction models like the Transformer, which lack recurrence or convolution, positional encodings are added to the input embeddings to inject information about the order of tokens. This equation defines the sinusoidal positional encoding for even dimensions.

Visual

Sinusoidal Positional Encoding (Even): step by step

Injecting Order with Positional Encoding

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Comparing Self-Attention with Recurrence and Convolution

Comparing Self-Attention with Recurrence and Convolution

Comparing Self-Attention, Recurrent, and Convolutional Layers

Lesson

BLEU scores and training costs on WMT 2014 translation tasks

Empirical Performance and Generalization

The Transformer model achieves state-of-the-art BLEU scores on both English-to-German (EN-DE) and English-to-French (EN-FR) translation tasks while requiring significantly lower training costs compared to previous architectures.

Lesson

F1 scores on English constituency parsing

Empirical Performance and Generalization

The Transformer generalizes well to English constituency parsing, achieving competitive F1 scores on Section 23 of WSJ under both discriminative (WSJ only) and semi-supervised training setups.

Visual

F1 scores on English constituency parsing

Empirical Performance and Generalization

Explore paper-reported results and reconciled deterministic comparisons.

Visual

BLEU scores and training costs on WMT 2014 translation tasks

Empirical Performance and Generalization

Explore paper-reported results and reconciled deterministic comparisons.

Visual

F1 scores on English constituency parsing: F1

Empirical Performance and Generalization

Compare reconciled F1 values reported by the paper.

Visual

BLEU scores and training costs on WMT 2014 translation tasks: BLEU

Empirical Performance and Generalization

Compare reconciled BLEU values reported by the paper.

Quiz

Test your understanding

Comprehensive Assessment

12 questions grounded in this paper section.