Study the paper

Kimi Linear: An Expressive, Efficient Attention Architecture

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

Comprehensive Assessment on Kimi Linear

Test your understanding

Check your understanding

1. How does Kimi Delta Attention (KDA) mathematically refine the scalar decay mechanism of Gated DeltaNet (GDN) to improve memory decay control and positional awareness?
2. In the recurrent formulation of Kimi Delta Attention (KDA), what is the correct mathematical expression for updating the state matrix $S_t$ given $S_{t-1}$, key $k_t$, value $v_t$, scalar gate $\beta_t$, and diagonalized gate $\text{Diag}(\alpha_t)$?
3. Write down the recurrence relation for the auxiliary vector $w_r[t]$ used in the WY representation of KDA to avoid matrix inversion, and explain the role of the diagonal cumulative decay term $\text{Diag}(\gamma_{i \to r}[t])$.
4. What is the primary hardware motivation for applying the UT transform in the chunkwise parallelization of KDA?
5. Explain how KDA avoids the numerical instability and high computational overhead of secondary chunking in the general DPLR formulation by binding variables $a$ and $b$ to the key $k$.
6. By how much does the operator efficiency of KDA improve compared to the general DPLR formulation due to the reduction of second-level chunk matrix computations and matrix multiplications?
7. Describe the layerwise hybridization strategy of Kimi Linear, including the specific ratio of KDA layers to Multi-Head Latent Attention (MLA) layers.
8. What position encoding design is used for the global MLA layers in Kimi Linear, and which layers are responsible for encoding positional information?
9. Summarize the performance hierarchy of Kimi Linear, GDN-H, and MLA across the pretraining/SFT stages versus the long-context evaluation stage.
10. What decoding throughput acceleration factor does Kimi Linear achieve compared to full MLA at a 1M context length?
11. Explain how linear attention with a gated delta rule can be interpreted as a form of multiplicative positional encoding, and contrast its transition matrix properties with those of RoPE.
12. Which of the following constraints imposed by RoPE is relaxed by the data-dependent transition matrix in gated delta rule linear attention?