You are reading immutable version 39. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Scaled Dot-Product Attention

Scaled Dot-Product Attention

Scaled Dot-Product Attention

Mathematical Formulation

Sources

S3.E1

Attention​(Q,K,V)=softmax​(Q​KTdk)​V\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V (1)
\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V

The Scaled Dot-Product Attention operates on query, key, and value matrices denoted as QQ, KK, and VV. The attention weights are computed using the dot product of queries and keys, scaled by the square root of the key dimension dkd_k, and normalized using a softmax function.

Sources

S3.E1

Attention​(Q,K,V)=softmax​(Q​KTdk)​V\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V (1)
\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V

Attention(Q,K,V)=softmax(QKTdk)V\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V

Sources

S3.E1

Attention​(Q,K,V)=softmax​(Q​KTdk)​V\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V (1)
\mathrm{Attention}(Q,K,V)=\mathrm{softmax}(\frac{QK^{T}}{\sqrt{d_{k}}})V

The Role of the Scaling Factor

Sources

S3.SS2.SSS1.p5.4

While for small values of dkd_{k} the two mechanisms perform similarly, additive attention outperforms dot product attention without scaling for larger values of dkd_{k} [3]. We suspect that for large values of dkd_{k}, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients 111To illustrate why the dot products get large, assume that the components of qq and kk are independent random variables with mean 0 and variance 11. Then their dot product, q⋅k=∑i=1dkqi​kiq\cdot k=\sum_{i=1}^{d_{k}}q_{i}k_{i}, has mean 0 and variance dkd_{k}.. To counteract this effect, we scale the dot products by 1dk\frac{1}{\sqrt{d_{k}}}.

For large values of dkd_k, the dot products grow extremely large in magnitude. This pushes the softmax function into regions with very small gradients, causing a vanishing gradient problem during training.

Sources

S3.SS2.SSS1.p5.4

While for small values of dkd_{k} the two mechanisms perform similarly, additive attention outperforms dot product attention without scaling for larger values of dkd_{k} [3]. We suspect that for large values of dkd_{k}, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients 111To illustrate why the dot products get large, assume that the components of qq and kk are independent random variables with mean 0 and variance 11. Then their dot product, q⋅k=∑i=1dkqi​kiq\cdot k=\sum_{i=1}^{d_{k}}q_{i}k_{i}, has mean 0 and variance dkd_{k}.. To counteract this effect, we scale the dot products by 1dk\frac{1}{\sqrt{d_{k}}}.

Why do dot products grow large?

Assuming the components of a query qq and key kk are independent random variables with mean 00 and variance 11, their dot product qk=i=1dkqikiq \cdot k = \sum_{i=1}^{d_k} q_i k_i has mean 00 and variance dkd_k. Scaling by 1/dk1/\sqrt{d_k} pulls the variance back to 11.

Sources

S3.SS2.SSS1.p5.4

While for small values of dkd_{k} the two mechanisms perform similarly, additive attention outperforms dot product attention without scaling for larger values of dkd_{k} [3]. We suspect that for large values of dkd_{k}, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients 111To illustrate why the dot products get large, assume that the components of qq and kk are independent random variables with mean 0 and variance 11. Then their dot product, q⋅k=∑i=1dkqi​kiq\cdot k=\sum_{i=1}^{d_{k}}q_{i}k_{i}, has mean 0 and variance dkd_{k}.. To counteract this effect, we scale the dot products by 1dk\frac{1}{\sqrt{d_{k}}}.