You are reading immutable version 40. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Comprehensive Assessment

Paper concept map

Interactive study view

Paper concept map

Paper evidence

Follow how prerequisites and the paper’s main ideas connect in teaching order.

Prerequisites

  • Basic understanding of deep learning architectures
  • Familiarity with sequence-to-sequence models and recurrent neural networks

Concepts

  • Explain the motivation for the Transformer architecture over recurrent and convolutional models
  • Derive the Scaled Dot-Product Attention formula and explain the scaling factor
  • Analyze the mechanics and benefits of Multi-Head Attention
  • Describe the design and purpose of sinusoidal Positional Encodings
  • Compare the computational complexity and path lengths of different layer types

Outcomes

  • Evaluate the empirical performance of the Transformer on translation and parsing tasks
  1. Basic understanding of deep learning architectures requires Explain the motivation for the Transformer architecture over recurrent and convolutional models
  2. Familiarity with sequence-to-sequence models and recurrent neural networks requires Explain the motivation for the Transformer architecture over recurrent and convolutional models
  3. Explain the motivation for the Transformer architecture over recurrent and convolutional models influences Derive the Scaled Dot-Product Attention formula and explain the scaling factor
  4. Derive the Scaled Dot-Product Attention formula and explain the scaling factor influences Analyze the mechanics and benefits of Multi-Head Attention
  5. Analyze the mechanics and benefits of Multi-Head Attention influences Describe the design and purpose of sinusoidal Positional Encodings
  6. Describe the design and purpose of sinusoidal Positional Encodings influences Compare the computational complexity and path lengths of different layer types
  7. Compare the computational complexity and path lengths of different layer types influences Evaluate the empirical performance of the Transformer on translation and parsing tasks