You are reading immutable version 39. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

24 activities

Lesson

At a glance

Introduction to the Transformer Architecture

Introduction to the Transformer Architecture

Lesson

Sequential Computation Bottleneck in Recurrent Models

Introduction to the Transformer Architecture

The Sequential Bottleneck of Recurrent Architectures

Lesson

Scaled Dot-Product Attention

Scaled Dot-Product Attention

Mathematical Formulation

Lesson

Scaled Dot-Product Attention

Scaled Dot-Product Attention

The Scaled Dot-Product Attention mechanism computes attention weights over a set of values based on queries and keys. The scaling factor sqrt d k is introduced to prevent the dot products from growing extremely large in magnitude for high-dimensional keys, which would otherwise push the softmax function into regions with extremely small gradients.

Visual

Illustrative teaching scenarios: attention scaling (not paper-reported)

Scaled Dot-Product Attention

Compare fixed illustrative teaching scenarios; these are not paper-reported results.

Visual

Illustrative teaching scenarios: attention concentration (not paper-reported)

Scaled Dot-Product Attention

Trace fixed illustrative teaching scenarios; these are not paper-reported results.

Visual

How attention redistributes weight

Scaled Dot-Product Attention

Illustrative attention weights computed from a fixed teaching example; these are not paper-reported values.

Visual

Explore attention score scaling

Scaled Dot-Product Attention

Move across fixed illustrative scenarios to see how score scaling changes concentration.

Visual

Scaled Dot-Product Attention: step by step

Scaled Dot-Product Attention

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Multi-Head Attention

Multi-Head Attention

Multi-Head Attention Mechanism

Lesson

Multi-Head Attention

Multi-Head Attention

The Multi-Head Attention mechanism allows the model to jointly attend to information from different representation subspaces at different positions. Instead of performing a single attention function with d text model -dimensional queries, keys, and values, we project them h times with different, learned linear projections.

Visual

Multi-Head Attention: step by step

Multi-Head Attention

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Injecting Sequence Order: Positional Encodings

Injecting Sequence Order: Positional Encodings

Sinusoidal Positional Encodings

Lesson

Sinusoidal Positional Encoding (Even Dimensions)

Injecting Sequence Order: Positional Encodings

In non-recurrent architectures like the Transformer, sequence order is not implicitly captured by the network structure. To inject positional information, sinusoidal positional encodings are added to the input embeddings. This equation defines the encoding value for even dimensions of the positional vector.

Visual

Sinusoidal Positional Encoding (Even Dimensions): step by step

Injecting Sequence Order: Positional Encodings

Follow the existing bounded walkthrough in its intended sequence.

Lesson

Empirical Results and Performance

Empirical Results and Performance

Empirical Evaluation and Efficiency

Lesson

Machine Translation Performance and Training Costs

Empirical Results and Performance

The Transformer model achieves state-of-the-art BLEU scores on both English-to-German (EN-DE) and English-to-French (EN-FR) translation tasks while requiring significantly lower training costs compared to previous recurrent or convolutional architectures.

Lesson

English Constituency Parsing Results

Empirical Results and Performance

To evaluate if the Transformer generalizes well to other tasks, it was evaluated on English constituency parsing on Section 23 of WSJ. Despite the lack of task-specific tuning, the Transformer (4 layers) achieves strong results in both WSJ-only discriminative and semi-supervised settings.

Visual

English Constituency Parsing Results

Empirical Results and Performance

Explore paper-reported results and reconciled deterministic comparisons.

Visual

English Constituency Parsing Results: WSJ 23 F1

Empirical Results and Performance

Compare reconciled WSJ 23 F1 values reported by the paper.

Quiz

Test your understanding

Comprehensive Assessment

10 questions grounded in this paper section.

Visual

Paper concept map

Comprehensive Assessment

Follow how prerequisites and the paper’s main ideas connect in teaching order.

Resource

Further learning

Comprehensive Assessment

3 supplementary resources for this paper section.

Resource

Research and implementation context

Comprehensive Assessment

Explore notable related work and public implementation context.