You are reading immutable version 1. The current guide may be newer.

Study the paper

Kimi Linear: An Expressive, Efficient Attention Architecture

Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.

All activities

KDA as Learnable Position Embeddings

KDA as Learnable Position Embeddings

KDA as Learnable Position Embeddings

Linear Attention and Multiplicative Positional Encodings

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…

Standard attention mechanisms are inherently agnostic to sequence order and rely on explicit positional encodings like Rotary Position Embedding (RoPE). RoPE applies multiplicative transformations to queries and keys, which can be analyzed through a generalized attention formulation:

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…

st,i=qt(j=i+1tRj)kis_{t,i} = q_t^\top \left( \prod_{j=i+1}^t R_j \right) k_i

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…

where RjR_j is a block-diagonal rotation matrix. Similarly, linear attention with a gated delta rule can be expressed in a comparable recurrent formulation:

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…

ot=i=1t(qt(j=i+1tAj(Iβjkjkj))ki)vio_t = \sum_{i=1}^t \left( q_t^\top \left( \prod_{j=i+1}^t A_j (I - \beta_j k_j k_j^\top) \right) k_i \right) v_i

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
Deep dive

From this perspective, the Gated Delta Network (GDN) acts as a multiplicative positional encoding where the transition matrix is data-dependent, relaxing the strict orthogonality constraints of RoPE.

While standard GDN uses a per-head scalar decay and lacks per-dimensional diversity, KDA introduces a learnable channel-wise gate Diag(αt)\text{Diag}(\alpha_t). This functions analogously to RoPE's assignment of different rotation frequencies to different dimensions, providing a highly expressive, learnable, and fine-grained positional encoding directly within the linear attention mechanism.

Sources

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…

block

Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT Table 6: An overview of attention mechanisms in their mathematically equivalent recurrent ( o t ) and parallel ( O ) forms. We omitted the normalization term and β t to achieve a more concise representation. The function ϕ refers to the infinite-dimensional feature space corresponding to the exponential kernel, i.e., ϕ ( q ) ⊤ ϕ ( k ) = exp( q ⊤ k ) . Recurrent form Parallel form SA [99] t P j =1 exp q ⊤ t k j  v j exp QK ⊤  ⊙ M  V SA + RoPE [88] t P j =1 exp q ⊤ t t Q s = j +1 R s ! k j ! v j  exp  R ( Q ) R ( K ) ⊤  ⊙ M  V LA [101] t P j =1 q ⊤ t k j  v j QK ⊤ ⊙ M  V Mamba2 [16] t P j =1 q ⊤ t t Q s = j +1 α s ! k j ! v j QK ⊤ ⊙ A ⊙ M  V GLA [114] t P j =1 q ⊤ t t Q s = j +1 Diag ( α s ) ! k j ! v j  ( Q ⊙ Γ ) K Γ  ⊤ ⊙ M  V DeltaNet [84] t P j =1 q ⊤ t t Q s = j +1 I − k s k ⊤ s  ! k j ! v j QK ⊤ ⊙ M  I + KK ⊤ ⊙ M −  − 1 V FoX [58] t P j =1 exp q ⊤ t k j  t Q s = j +1 α s ! v j exp QK ⊤  ⊙ A ⊙ M  V DeltaFormer [125] t P j =1 ϕ ( q t ) ⊤ t Q s = j +1  I − ϕ ( k s ) ϕ ( w s ) ⊤  ! ϕ ( k j ) ! v j exp QK ⊤  ⊙ M  I + exp WK ⊤  ⊙ M −  − 1 V PaTH-FoX [115] t P j =1 exp q ⊤ t t Q s…