If you run agents against a third-party inference endpoint, you’re trusting that the model behind the URL is the one you benchmarked — and that it stays that model at 2am under load. Ventor-QTest is a black-box audit that stops you from taking that on faith, and it needs no probability information from the provider.

It treats hosted model routing as a stochastic process and probes it two ways:

That last point is the one I’d take straight to production. A provider silently swapping in a quantized or rerouted model can look fine on short QA while quietly wrecking long-horizon agentic runs — exactly the setting where you’re least able to eyeball each output. The HF paper page links the open-source implementation if you want to point it at your own vendors.

We pour effort into eval harnesses for our own models and spend almost nothing verifying that the API we rent still serves what the contract says. Which of your production routes could pass a short-prompt eval today and still be drifting under long agentic load?

tags: [ llm-ops ] [ ai-infrastructure ] [ research ]