The pitch for “decision models” like Jev is seductive: LLM-judge quality for classifier money, no GPU, typed probabilities instead of a wall of text. Red Hat’s AI safety team actually benchmarked that claim across prompt-injection and content-safety guardrails, nine approaches deep — and the pitch doesn’t survive contact with the latency column.

The operational read for anyone shipping guardrails: these checks run on every request, so a 300ms remote call versus a 33ms in-process model is a P99 decision, not a leaderboard footnote. The HN thread has the familiar tension between people who want one model for everything and people who’ve been burned by exactly that.

When a 2019-era classifier still lands in the top three for well-defined risks, maybe the question isn’t which model replaces the judge — it’s why we reached for a judge at all.

tags: [ llm-ops ] [ enterprise-ai ] [ industry ]