Study the paper
Attention Is All You Need
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Injecting Sequence Order via Positional Encodings
Sinusoidal Positional Encodings
Sinusoidal Positional Encodings
Because the Transformer contains no recurrence and no convolution, it is entirely permutation-invariant. To make use of the order of the sequence, we must inject information about the relative or absolute position of the tokens. This is achieved by adding "positional encodings" of dimension directly to the input embeddings.
Sources
S3.SS5.p1.1
Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodeld_{\text{model}} as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].
Implementation detail
Sources
S3.SS5.p2.1
In this work, we use sine and cosine functions of different frequencies:
equation
PE(pos,2i)=sin(pos/100002i/dmodel)\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})equation
PE(pos,2i+1)=cos(pos/100002i/dmodel)\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})Deep dive
These sinusoidal functions use different frequencies along the dimension index . This formulation allows the model to easily learn to attend by relative positions, as for any fixed offset , can be represented as a linear function of .
Sources
S3.SS5.p1.1
Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodeld_{\text{model}} as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].