Study the paper
Kimi Linear: An Expressive, Efficient Attention Architecture
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Empirical Evaluation and Scaling Laws
Pretrain Performance Comparison
Pretrain Performance Comparison
Table 3 compares the pretraining performance of Kimi Linear against the full-attention MLA baseline and the hybrid GDN baseline (GDN-H) across general, math & code, and Chinese benchmarks, all trained on 1.4T tokens.
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT • General Knowledge: Kimi Linear scores highest on all of the key benchmarks like BBH, MMLU and HellaSwag. • Reasoning: It leads in math (GSM8K) and most code tasks (CRUXEval). However, it scores slightly lower on EvalPlus compared to GDN-H. • Chinese Tasks: Kimi Linear achieves the top scores on CEval and CMMLU. In summary, Kimi Linear demonstrated the strongest performance, positioning it as a strong alternative to full-attention architectures at short context pretraining. Table 3: Performance comparison of Kimi Linear with the full-attention MLA baseline and the hybrid GDN baseline, all after the same pretraining recipe. Kimi Linear consistently outperforms both MLA and GDN-H on short-context pretrain evaluations. Best per-column results are bolded . Type Base MLA GDN-H Kimi Linear Trained Tokens 1.4T 1.4T 1.4T General HellaSwag 81.7 82.2 82.9 ARC-challenge 64.6 66.5 67.3 Winogrande 78.1 77.9 78.6 BBH 71.6 70.6 72.9 MMLU 71.6 72.2 73.8 MMLU-Pro 47.2 47.9 51.0 TriviaQA 68.9 70.1 71.7 Math & Code GSM8K 83.7 81.7 83.9 MATH 54.7 54.1 54.7 EvalPlus 59.5 63.1 60.2 CRUXEval-I-cot 51.6 56.0 56.6 CRUXEval-O-…
| Measure | Value |
|---|
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT • General Knowledge: Kimi Linear scores highest on all of the key benchmarks like BBH, MMLU and HellaSwag. • Reasoning: It leads in math (GSM8K) and most code tasks (CRUXEval). However, it scores slightly lower on EvalPlus compared to GDN-H. • Chinese Tasks: Kimi Linear achieves the top scores on CEval and CMMLU. In summary, Kimi Linear demonstrated the strongest performance, positioning it as a strong alternative to full-attention architectures at short context pretraining. Table 3: Performance comparison of Kimi Linear with the full-attention MLA baseline and the hybrid GDN baseline, all after the same pretraining recipe. Kimi Linear consistently outperforms both MLA and GDN-H on short-context pretrain evaluations. Best per-column results are bolded . Type Base MLA GDN-H Kimi Linear Trained Tokens 1.4T 1.4T 1.4T General HellaSwag 81.7 82.2 82.9 ARC-challenge 64.6 66.5 67.3 Winogrande 78.1 77.9 78.6 BBH 71.6 70.6 72.9 MMLU 71.6 72.2 73.8 MMLU-Pro 47.2 47.9 51.0 TriviaQA 68.9 70.1 71.7 Math & Code GSM8K 83.7 81.7 83.9 MATH 54.7 54.1 54.7 EvalPlus 59.5 63.1 60.2 CRUXEval-I-cot 51.6 56.0 56.6 CRUXEval-O-…