Linear attention has spent years as the thing I’d reach for only after giving up on quality to buy throughput. The Kimi Linear tech report is interesting because it claims the opposite trade — matching or beating full attention while cutting the memory you pay for at serving time.

KV cache is the quiet line item that decides how many concurrent agent sessions a single GPU can hold. If a hybrid linear stack really gives back three-quarters of it without a quality tax, the math for long-horizon agentic workloads — where context just keeps growing — shifts more than any single benchmark suggests. The HN discussion is already poking at the “under fair comparisons” caveat, which is exactly the right thing to poke at. Does the drop-in claim survive contact with a real 1M-context agent loop, or does the delta rule start forgetting things full attention would have kept?