Study the paper
Kimi Linear: An Expressive, Efficient Attention Architecture
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
KDA as Learnable Position Embeddings
KDA as Learnable Position Embeddings
KDA as Learnable Position Embeddings
Linear Attention and Multiplicative Positional Encodings
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
Standard attention mechanisms are inherently agnostic to sequence order and rely on explicit positional encodings like Rotary Position Embedding (RoPE). RoPE applies multiplicative transformations to queries and keys, which can be analyzed through a generalized attention formulation:
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
where is a block-diagonal rotation matrix. Similarly, linear attention with a gated delta rule can be expressed in a comparable recurrent formulation:
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
Deep dive
From this perspective, the Gated Delta Network (GDN) acts as a multiplicative positional encoding where the transition matrix is data-dependent, relaxing the strict orthogonality constraints of RoPE.
While standard GDN uses a per-head scalar decay and lacks per-dimensional diversity, KDA introduces a learnable channel-wise gate . This functions analogously to RoPE's assignment of different rotation frequencies to different dimensions, providing a highly expressive, learnable, and fine-grained positional encoding directly within the linear attention mechanism.
Sources
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT 4K 128K 256K 512K 1M 0 20 40 60 2 . 9 × 2 . 3 × Prefilling Length Latency (s) MLA GDN-H Kimi Linear (a) 4K 128K 256K 512K 1M 5 10 15 2 . 2 × 1 . 8 × Decoding Length TPOT (ms) MLA GDN-H Kimi Linear (b) Figure 7: (a) The prefilling time of MLA (full attention), hybrid GDN-H and our Kimi Linear. (b) The time per output token (TPOT) for MLA, GDN-H and Kimi Linear during decoding. (We use batch size = 1 here for tests.) performance curves are virtually indistinguishable, confirming that our method maintains high efficiency. The hybrid Kimi Linear model demonstrates a clear efficiency advantage over the MLA baseline as sequence length increases. While its performance is comparable to MLA at shorter lengths (4k–16k), it becomes significantly faster from 128k onwards. This efficiency gap widens dramatically at scale, with Kimi Linear outperforming MLA by a factor of 2 . 3 for 512k sequences and 2 . 9 for 1M sequences. As shown in Figure 1b, Kimi Linear fully demonstrates its advantages during the decoding phase. For decoding at 1M context length, Kimi Linear is 6 × faster than full attention. 6 Discussions 6.1…
block
Kimi Linear: An Expressive, Efficient Attention Architecture T ECHNICAL R EPORT Table 6: An overview of attention mechanisms in their mathematically equivalent recurrent ( o t ) and parallel ( O ) forms. We omitted the normalization term and β t to achieve a more concise representation. The function ϕ refers to the infinite-dimensional feature space corresponding to the exponential kernel, i.e., ϕ ( q ) ⊤ ϕ ( k ) = exp( q ⊤ k ) . Recurrent form Parallel form SA [99] t P j =1 exp q ⊤ t k j v j exp QK ⊤ ⊙ M V SA + RoPE [88] t P j =1 exp q ⊤ t t Q s = j +1 R s ! k j ! v j exp R ( Q ) R ( K ) ⊤ ⊙ M V LA [101] t P j =1 q ⊤ t k j v j QK ⊤ ⊙ M V Mamba2 [16] t P j =1 q ⊤ t t Q s = j +1 α s ! k j ! v j QK ⊤ ⊙ A ⊙ M V GLA [114] t P j =1 q ⊤ t t Q s = j +1 Diag ( α s ) ! k j ! v j ( Q ⊙ Γ ) K Γ ⊤ ⊙ M V DeltaNet [84] t P j =1 q ⊤ t t Q s = j +1 I − k s k ⊤ s ! k j ! v j QK ⊤ ⊙ M I + KK ⊤ ⊙ M − − 1 V FoX [58] t P j =1 exp q ⊤ t k j t Q s = j +1 α s ! v j exp QK ⊤ ⊙ A ⊙ M V DeltaFormer [125] t P j =1 ϕ ( q t ) ⊤ t Q s = j +1 I − ϕ ( k s ) ϕ ( w s ) ⊤ ! ϕ ( k j ) ! v j exp QK ⊤ ⊙ M I + exp WK ⊤ ⊙ M − − 1 V PaTH-FoX [115] t P j =1 exp q ⊤ t t Q s…