Study the paper
Kimi Linear: An Expressive, Efficient Attention Architecture
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Empirical Evaluation and Scaling Laws
Long-Context Benchmark Comparisons
Long-Context Benchmark Comparisons
Table 5 presents the long-context performance of Kimi Linear compared to MLA, GDN-H, and Kimi Linear (RoPE) across several benchmarks at 128k context length.
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT Table 5: Comparisons of Kimi Linear with MLA, GDN-H, and Kimi Linear (RoPE) across long-context benchmarks. The last column reports the overall average ( ↑ ). All models is trained on 1.4T tokens. Best per-column results are bolded . RULER MRCR HELMET-ICL LongBench V2 Frames RepoQA Long Code Arena Avg. Lib Commit MLA 81.3 22.6 88.0 36.1 60.5 63.0 32.8 33.2 52.2 GDN-H 80.5 23.9 85.5 32.6 58.7 63.0 34.7 30.5 51.2 Kimi Linear (RoPE) 78.8 22.0 88.0 35.4 59.9 66.5 31.3 32.5 51.8 Kimi Linear 84.3 29.6 90.0 35.0 58.8 68.5 37.1 32.7 54.5 20 40 60 80 100 20 35 50 65 Train Accuracy MLA@1.4T Kimi Linear@1.4T (a) 20 40 60 80 100 70 78 86 94 MATH 500 Test Accuracy MLA@1.4T Kimi Linear@1.4T (b) 20 40 60 80 100 10 15 20 25 AIME 2025 Accuracy MLA@1.4T Kimi Linear@1.4T (c) Figure 6: The training and test accuracy curves for Kimi Linear@1.4T and MLA@1.4T during Math RL training. Kimi Linear consistently outperforms the full attention baseline by a sizable margin during the whole RL process. Long Context Performance Evaluation We evaluate the long-context performance of Kimi Linear against three baseline models—MLA, GDN-…
| Measure | Value |
|---|
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT Table 5: Comparisons of Kimi Linear with MLA, GDN-H, and Kimi Linear (RoPE) across long-context benchmarks. The last column reports the overall average ( ↑ ). All models is trained on 1.4T tokens. Best per-column results are bolded . RULER MRCR HELMET-ICL LongBench V2 Frames RepoQA Long Code Arena Avg. Lib Commit MLA 81.3 22.6 88.0 36.1 60.5 63.0 32.8 33.2 52.2 GDN-H 80.5 23.9 85.5 32.6 58.7 63.0 34.7 30.5 51.2 Kimi Linear (RoPE) 78.8 22.0 88.0 35.4 59.9 66.5 31.3 32.5 51.8 Kimi Linear 84.3 29.6 90.0 35.0 58.8 68.5 37.1 32.7 54.5 20 40 60 80 100 20 35 50 65 Train Accuracy MLA@1.4T Kimi Linear@1.4T (a) 20 40 60 80 100 70 78 86 94 MATH 500 Test Accuracy MLA@1.4T Kimi Linear@1.4T (b) 20 40 60 80 100 10 15 20 25 AIME 2025 Accuracy MLA@1.4T Kimi Linear@1.4T (c) Figure 6: The training and test accuracy curves for Kimi Linear@1.4T and MLA@1.4T during Math RL training. Kimi Linear consistently outperforms the full attention baseline by a sizable margin during the whole RL process. Long Context Performance Evaluation We evaluate the long-context performance of Kimi Linear against three baseline models—MLA, GDN-…