Study the paper
Kimi Linear: An Expressive, Efficient Attention Architecture
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
10 activities
At a glance
Introduction to Kimi Linear and KDAOverview of Kimi Linear
LessonIntroduction to Kimi Linear and KDA
Introduction to Kimi Linear and KDAKimi Delta Attention (KDA) and Hybrid Architecture
LessonMathematical Formulation and Chunkwise Parallelism of KDA
Mathematical Formulation and Chunkwise Parallelism of KDAMathematical Formulation of Kimi Delta Attention (KDA)
LessonEfficiency Analysis and DPLR Comparison
Efficiency Analysis and DPLR ComparisonComparing KDA with General DPLR Formulations
LessonKDA as Learnable Position Embeddings
KDA as Learnable Position EmbeddingsLinear Attention and Multiplicative Positional Encodings
LessonPretrain Performance Comparison
Empirical Evaluation and Scaling LawsTable 3 compares the pretraining performance of Kimi Linear against the full-attention MLA baseline and the hybrid GDN baseline (GDN-H) across general, math & code, and Chinese benchmarks, all trained on 1.4T tokens.
LessonLong-Context Benchmark Comparisons
Empirical Evaluation and Scaling LawsTable 5 presents the long-context performance of Kimi Linear compared to MLA, GDN-H, and Kimi Linear (RoPE) across several benchmarks at 128k context length.
QuizTest your understanding
Comprehensive Assessment on Kimi Linear12 questions grounded in this paper section.
ResourceFurther learning
Comprehensive Assessment on Kimi Linear3 supplementary resources for this paper section.
ResourceResearch and implementation context
Comprehensive Assessment on Kimi LinearExplore notable related work and public implementation context.