Study the paper
Attention Is All You Need
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Injecting Order with Positional Encoding
Injecting Order with Positional Encoding
Injecting Order with Positional Encoding
Injecting Order with Positional Encodings
Sources
S3.SS5.p1.1
Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodeld_{\text{model}} as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].
Because the Transformer architecture contains no recurrence and no convolution, it is completely permutation-invariant. To make use of the order of the sequence, we must inject information about the relative or absolute position of the tokens. To achieve this, "positional encodings" are added directly to the input embeddings at the bottoms of the encoder and decoder stacks. These encodings have the same dimension as the embeddings, allowing them to be summed directly.
Sources
S3.SS5.p1.1
Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodeld_{\text{model}} as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].
Implementation detail
The authors use sine and cosine functions of different frequencies to construct these encodings:
Sources
S3.SS5.p2.1
In this work, we use sine and cosine functions of different frequencies:
equation
PE(pos,2i)=sin(pos/100002i/dmodel)\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})equation
PE(pos,2i+1)=cos(pos/100002i/dmodel)\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})Implementation detail
Sources
equation
PE(pos,2i)=sin(pos/100002i/dmodel)\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i)}=sin(pos/10000^{2i/d_{\text{model}}})Implementation detail
Sources
equation
PE(pos,2i+1)=cos(pos/100002i/dmodel)\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})
\displaystyle PE_{(pos,2i+1)}=cos(pos/10000^{2i/d_{\text{model}}})Deep dive
Here, is the position in the sequence, and is the dimension index. This formulation allows the model to easily learn to attend by relative positions, since for any fixed offset , can be represented as a linear function of .
Sources
S3.SS5.p1.1
Since our model contains no recurrence and no convolution, in order for the model to make use of the order of the sequence, we must inject some information about the relative or absolute position of the tokens in the sequence. To this end, we add "positional encodings" to the input embeddings at the bottoms of the encoder and decoder stacks. The positional encodings have the same dimension dmodeld_{\text{model}} as the embeddings, so that the two can be summed. There are many choices of positional encodings, learned and fixed [9].