The cost that quietly wrecks an eval or guardrail stack isn’t accuracy — it’s that every criterion is another model call. Jev goes straight at that: instead of a generative judge spending a full decoding pass per criterion, or a fixed-label classifier like Llama Guard scoring one thing per call, it answers many typed questions about a single input — each with a calibrated probability — in one pass.

Where this lands in a real stack: a cheap first-layer screen for high-volume routing, uncertain cases escalated to a heavier judge or a person, and label disagreements mined as an audit signal. The single-pass, multi-question shape is what makes any of that viable at QPS, and the full setup sits on the HF paper page.

Early write-ups read cautiously positive. Analytics Made Simple calls RLCD “a real shift, not just another model release” for treating inference as a typed function call, while naming the hard limits — text/JSON only, an option cap, no free-form generation. redreamality is more guarded: fine as a first-layer monitor, but don’t equate a calibrated probability with real-world risk, and expect thresholds to move between benchmarks. Both agree the cost curve is the story; where they part is how far to trust the number that comes back.

tags: [ llm-ops ] [ research ]