jamesob’s local-LLM writeup is the most honest hardware post I’ve read this year, because the takeaway isn’t “buy the new thing.” It’s: spend on VRAM, buy everything else secondhand. The ~$40k rig runs GLM-5.2 quantized on 4x RTX 6000 Blackwell and lands “pretty close to Opus” — but the base system wrapped around those GPUs is a $500 secondhand EPYC Milan and DDR4 ECC.

The economics argument on the HN thread mostly misses this. The question was never “is local cheaper than the API” (usually it isn’t). It’s whether you need the weights on your own iron for governance reasons — and if you do, this is the actual bill of materials rather than a hand-wave.

At enterprise-fleet scale the calculus flips again: you’re not power-limiting 3090s, you’re amortizing a cluster and worrying about utilization. But the single-rig version is a useful forcing function — it makes the VRAM-per-dollar math impossible to fudge. When does self-hosting stop being a hobby and start being a compliance line item for your team?