Study the paper

Kimi Linear: An Expressive, Efficient Attention Architecture

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

10 activities

Lesson

At a glance

Introduction to Kimi Linear and KDA

Overview of Kimi Linear

Lesson

Introduction to Kimi Linear and KDA

Introduction to Kimi Linear and KDA

Kimi Delta Attention (KDA) and Hybrid Architecture

Lesson

Mathematical Formulation and Chunkwise Parallelism of KDA

Mathematical Formulation and Chunkwise Parallelism of KDA

Mathematical Formulation of Kimi Delta Attention (KDA)

Lesson

Efficiency Analysis and DPLR Comparison

Efficiency Analysis and DPLR Comparison

Comparing KDA with General DPLR Formulations

Lesson

KDA as Learnable Position Embeddings

KDA as Learnable Position Embeddings

Linear Attention and Multiplicative Positional Encodings

Lesson

Pretrain Performance Comparison

Empirical Evaluation and Scaling Laws

Table 3 compares the pretraining performance of Kimi Linear against the full-attention MLA baseline and the hybrid GDN baseline (GDN-H) across general, math & code, and Chinese benchmarks, all trained on 1.4T tokens.

Lesson

Long-Context Benchmark Comparisons

Empirical Evaluation and Scaling Laws

Table 5 presents the long-context performance of Kimi Linear compared to MLA, GDN-H, and Kimi Linear (RoPE) across several benchmarks at 128k context length.

Quiz

Test your understanding

Comprehensive Assessment on Kimi Linear

12 questions grounded in this paper section.

Resource

Further learning

Comprehensive Assessment on Kimi Linear

3 supplementary resources for this paper section.

Resource

Research and implementation context

Comprehensive Assessment on Kimi Linear

Explore notable related work and public implementation context.