Published hybrid-attention comparison [ftip-00KU]

Kimi Linear combines recurrent KDA layers with periodic full MLA attention. Its published comparison shows task-dependent score changes and lower long-context decoding times. The selected scores in Example [ftip-00LC], the memory estimate in Example 4, and the discovery calculation in Example 9 connect these observations to a post-training question: when can a cheaper attempt compensate for a possible lower per-attempt success probability?