The pitch for “decision models” like Jev is seductive: LLM-judge quality for classifier money, no GPU, typed probabilities instead of a wall of text. Red Hat’s AI safety team actually benchmarked that claim across prompt-injection and content-safety guardrails, nine approaches deep — and the pitch doesn’t survive contact with the latency column.
- 🎯 A 125M-param classifier won on speed: Granite Guardian at ~33ms and DeBERTa at ~54ms, both CPU-only, while Jev’s remote API sat around 348–360ms
- 📊 Accuracy was a wash at the top: Jev led content safety (86.2%), but a 35B LLM judge led prompt injection (89.3%) with DeBERTa right behind it
- ⚡ “Cheaper and faster” didn’t hold: decision models didn’t reliably beat LLM-as-a-judge or open alternatives on the metrics that bite in the hot path
- ⚠️ Zero-shot is the real niche: these models earn their keep when you have no labeled data for a novel risk, not as a blanket classifier replacement
- 🔍 Prompt phrasing moved the results enough that a sloppy eval could rank these models in almost any order you like
The operational read for anyone shipping guardrails: these checks run on every request, so a 300ms remote call versus a 33ms in-process model is a P99 decision, not a leaderboard footnote. The HN thread has the familiar tension between people who want one model for everything and people who’ve been burned by exactly that.
When a 2019-era classifier still lands in the top three for well-defined risks, maybe the question isn’t which model replaces the judge — it’s why we reached for a judge at all.