The interesting result in Cloudflare’s writeup on serving Kimi K2.6 and GLM 5.2 isn’t how much memory they saved. It’s that quantization helped in exactly one half of the inference loop and hurt in the other — so the right answer wasn’t a single “quantized” checkpoint.

This is the part of LLM Ops that never shows up on a model card: the win comes from matching numeric precision to the prefill/decode split, not from picking one “smaller” weight file. The HN thread has good skepticism about whether FP8 KV cache holds up on long-context recall rather than aggregate benchmarks — the right thing to be nervous about.

If your serving stack still quantizes uniformly across both phases, what is that symmetry actually costing you at P99?

tags: [ llm-ops ] [ ai-infrastructure ]