The headline — 125B parameters at 100 tokens/sec on a single RTX 4090 — is doing a lot of work. Strata runs Qwen 3.8 Flash Next, and what makes it fit isn’t miracle compression; it’s that the model is a sparse MoE with roughly 6B active parameters per token. 4-bit weights plus custom kernel scheduling keep the resident set inside 24GB. The interesting engineering is the scheduling, not the parameter count.

The HN discussion has the usual throughput one-upmanship, which is exactly why reproducibility matters more than the peak figure.

That skepticism is the dominant note across the coverage. DEV argues the 100 tok/s headline collapses once cooling, power stability, quantization choices, and memory fragmentation enter the picture. Developers Digest is blunter — the 100 figure is an estimate and the 124 tok/s claim is a single unverified report, with measured speeds landing between 53 and 94 depending on quantization. PromptZone is the most upbeat, calling it worth testing while noting throughput drops on 3090/4080 cards and that 4-bit costs real accuracy on long-context work. The trend is consistent: everyone agrees the feat is real, and almost no one expects the headline number to survive contact with a second machine.

tags: [ llm-ops ] [ ai-infrastructure ]