You are reading immutable version 40. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Scaled Dot-Product Attention

Scaled Dot-Product Attention: step by step

Interactive study view

Scaled Dot-Product Attention: step by step

Illustrative teaching data

Follow the existing bounded walkthrough in its intended sequence.

Step 1 of 4

  1. Step 1

    Compute dot product of Q and K^T to obtain raw attention scores.

  2. Step 2

    Scale the scores by dividing by the square root of d_k.

  3. Step 3

    Apply the softmax function to obtain normalized attention weights.

  4. Step 4

    Multiply the attention weights by the value matrix V to get the final output.