Study the paper
Attention Is All You Need
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Empirical Evaluation and Generalization
Table 2: BLEU scores and training costs for translation tasks
Table 2: BLEU scores and training costs for translation tasks
The Transformer model achieves state-of-the-art results on translation tasks. On the English-to-German (EN-DE) translation task, the Transformer (big) model achieves a BLEU score of 28.4, while the base model achieves 27.3. On the English-to-French (EN-FR) task, the Transformer (big) model achieves a BLEU score of 41.8, and the base model achieves 38.1.
Sources
S6.T2
Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the English-to-German and English-to-French newstest2014 tests at a fraction of the training cost. Model BLEU Training Cost (FLOPs) EN-DE EN-FR EN-DE EN-FR ByteNet [18] 23.75 Deep-Att + PosUnk [39] 39.2 1.0⋅10201.0\cdot 10^{20} GNMT + RL [38] 24.6 39.92 2.3⋅10192.3\cdot 10^{19} 1.4⋅10201.4\cdot 10^{20} ConvS2S [9] 25.16 40.46 9.6⋅10189.6\cdot 10^{18} 1.5⋅10201.5\cdot 10^{20} MoE [32] 26.03 40.56 2.0⋅10192.0\cdot 10^{19} 1.2⋅10201.2\cdot 10^{20} Deep-Att + PosUnk Ensemble [39] 40.4 8.0⋅10208.0\cdot 10^{20} GNMT + RL Ensemble [38] 26.30 41.16 1.8⋅10201.8\cdot 10^{20} 1.1⋅10211.1\cdot 10^{21} ConvS2S Ensemble [9] 26.36 41.29 7.7⋅10197.7\cdot 10^{19} 1.2⋅10211.2\cdot 10^{21} Transformer (base model) 27.3 38.1 3.3⋅𝟏𝟎𝟏𝟖3.3\cdot 10^{18} Transformer (big) 28.4 41.8 2.3⋅10192.3\cdot 10^{19}
| Measure | Value |
|---|
Sources
S6.T2
Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the English-to-German and English-to-French newstest2014 tests at a fraction of the training cost. Model BLEU Training Cost (FLOPs) EN-DE EN-FR EN-DE EN-FR ByteNet [18] 23.75 Deep-Att + PosUnk [39] 39.2 1.0⋅10201.0\cdot 10^{20} GNMT + RL [38] 24.6 39.92 2.3⋅10192.3\cdot 10^{19} 1.4⋅10201.4\cdot 10^{20} ConvS2S [9] 25.16 40.46 9.6⋅10189.6\cdot 10^{18} 1.5⋅10201.5\cdot 10^{20} MoE [32] 26.03 40.56 2.0⋅10192.0\cdot 10^{19} 1.2⋅10201.2\cdot 10^{20} Deep-Att + PosUnk Ensemble [39] 40.4 8.0⋅10208.0\cdot 10^{20} GNMT + RL Ensemble [38] 26.30 41.16 1.8⋅10201.8\cdot 10^{20} 1.1⋅10211.1\cdot 10^{21} ConvS2S Ensemble [9] 26.36 41.29 7.7⋅10197.7\cdot 10^{19} 1.2⋅10211.2\cdot 10^{21} Transformer (base model) 27.3 38.1 3.3⋅𝟏𝟎𝟏𝟖3.3\cdot 10^{18} Transformer (big) 28.4 41.8 2.3⋅10192.3\cdot 10^{19}