The number that’s hard to ignore in this MI300X worklog: 192GB of HBM3 per card against the H100’s 80GB, comparable FP8 compute, list price roughly half — and you can rent one on-demand today while H100 capacity is sold out. The silicon isn’t the bottleneck. The stack is.

The post is an honest log of getting DeepSeek-V4-Flash to serve on AMD when vLLM simply doesn’t, and the failure modes are the ones that quietly eat infra teams alive:

This is the part of LLM ops nobody demos. Picking an accelerator off a spec sheet is easy; the real cost shows up in the weeks of kernel-level yak-shaving before the model serves a single token at target latency. The HN discussion splits between “AMD is finally viable” and “this proves it still isn’t” — both are right, depending on whether you have an engineer willing to write the worklog.

My bet: the teams that come out ahead in the next 18 months of the compute crunch aren’t the ones with the best NVIDIA allocation — they’re the ones who treated portability as an eval target before the shortage forced their hand. Is your inference stack one vendor outage away from a standstill?

tags: [ ai-infrastructure ] [ llm-ops ]