“Know when it’s wrong” is the quiet unlock in hybrid inference — and Cactus Hybrid (repo) makes the case that the confidence signal, not the small model, is the hard part.

This is the agentic routing/fallback pattern most teams bolt on with a second LLM call to “check the answer.” A calibrated in-model probe is cheaper and lives where the hidden states already are. The HN discussion is bullish; my question is operational — how stable is that 0.85 line as the input distribution drifts away from what the probe was trained on?