Fitting all 304B parameters of DeepSeek V4 Flash into a single MI300X’s 192GB of HBM without quantization is the headline of this repo, but it’s the least interesting number in it. Memory capacity is a spec-sheet win. The serving economics underneath are where the actual engineering is.

None of this shows up in a benchmark table, which is exactly why single-GPU serving posts are worth reading closely. The HN discussion is mostly about whether MI300X is finally viable for inference — the more useful question is whether that 168-to-542 concurrency curve holds at your P99, or collapses the moment the offload tier starts thrashing.

tags: [ llm-ops ] [ ai-infrastructure ]