You are reading immutable version 28. The current guide may be newer.

Study the paper

Attention Is All You Need

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Empirical Performance and Generalization

BLEU scores and training costs on WMT 2014 translation tasks

BLEU scores and training costs on WMT 2014 translation tasks

The Transformer model achieves state-of-the-art BLEU scores on both English-to-German (EN-DE) and English-to-French (EN-FR) translation tasks while requiring significantly lower training costs compared to previous architectures.

Sources

S6.T2

Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the English-to-German and English-to-French newstest2014 tests at a fraction of the training cost. Model BLEU Training Cost (FLOPs) EN-DE EN-FR EN-DE EN-FR ByteNet [18] 23.75 Deep-Att + PosUnk [39] 39.2 1.0⋅10201.0\cdot 10^{20} GNMT + RL [38] 24.6 39.92 2.3⋅10192.3\cdot 10^{19} 1.4⋅10201.4\cdot 10^{20} ConvS2S [9] 25.16 40.46 9.6⋅10189.6\cdot 10^{18} 1.5⋅10201.5\cdot 10^{20} MoE [32] 26.03 40.56 2.0⋅10192.0\cdot 10^{19} 1.2⋅10201.2\cdot 10^{20} Deep-Att + PosUnk Ensemble [39] 40.4 8.0⋅10208.0\cdot 10^{20} GNMT + RL Ensemble [38] 26.30 41.16 1.8⋅10201.8\cdot 10^{20} 1.1⋅10211.1\cdot 10^{21} ConvS2S Ensemble [9] 26.36 41.29 7.7⋅10197.7\cdot 10^{19} 1.2⋅10211.2\cdot 10^{21} Transformer (base model) 27.3 38.1 3.3⋅𝟏𝟎𝟏𝟖3.3\cdot 10^{18} Transformer (big) 28.4 41.8 2.3⋅10192.3\cdot 10^{19}
Reported values
MeasureValue
Transformer (base model) EN-DE BLEU27.3 BLEU
Transformer (big) EN-DE BLEU28.4 BLEU
Transformer (base model) EN-FR BLEU38.1 BLEU
Transformer (big) EN-FR BLEU41.8 BLEU
Sources

S6.T2

Table 2: The Transformer achieves better BLEU scores than previous state-of-the-art models on the English-to-German and English-to-French newstest2014 tests at a fraction of the training cost. Model BLEU Training Cost (FLOPs) EN-DE EN-FR EN-DE EN-FR ByteNet [18] 23.75 Deep-Att + PosUnk [39] 39.2 1.0⋅10201.0\cdot 10^{20} GNMT + RL [38] 24.6 39.92 2.3⋅10192.3\cdot 10^{19} 1.4⋅10201.4\cdot 10^{20} ConvS2S [9] 25.16 40.46 9.6⋅10189.6\cdot 10^{18} 1.5⋅10201.5\cdot 10^{20} MoE [32] 26.03 40.56 2.0⋅10192.0\cdot 10^{19} 1.2⋅10201.2\cdot 10^{20} Deep-Att + PosUnk Ensemble [39] 40.4 8.0⋅10208.0\cdot 10^{20} GNMT + RL Ensemble [38] 26.30 41.16 1.8⋅10201.8\cdot 10^{20} 1.1⋅10211.1\cdot 10^{21} ConvS2S Ensemble [9] 26.36 41.29 7.7⋅10197.7\cdot 10^{19} 1.2⋅10211.2\cdot 10^{21} Transformer (base model) 27.3 38.1 3.3⋅𝟏𝟎𝟏𝟖3.3\cdot 10^{18} Transformer (big) 28.4 41.8 2.3⋅10192.3\cdot 10^{19}