Study the paper
Kimi Linear: An Expressive, Efficient Attention Architecture
Lessons, visuals, quizzes, flashcards, and resources—organized in teaching order.
Comprehensive Assessment on Kimi Linear
Research and implementation context
Supplementary context
Research and implementation context
Research lineage
Notable related work, not a complete survey
Predecessors
Gated Delta Networks: Improving Mamba2 with Delta Rule
This paper introduces Gated Delta Networks (Gated DeltaNet), which improves Mamba2 with the Delta Rule, establishing the foundation for subsequent linear attention and delta rule optimizations.
Successors
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
Gated DeltaNet-2 directly builds upon Gated DeltaNet by decoupling the erase and write operations in linear attention, representing a direct architectural successor.