Doug Turnbull’s post flips classification inside out: stop asking the LLM to be right, and start asking it to be fluent.

Instead of stuffing a thousand-label taxonomy into the prompt and constraining the output, you ask a small, cheap model to invent plausible-but-fake categories for the query — then embed those hallucinations with something like MiniLM and nearest-neighbor them onto your real vocabulary.

The HN discussion is worth a scroll if you’ve ever fought a provider’s structured-output ceiling. My take: this is the rare trick that gets cheaper and more robust at once. If a 22M-parameter embedder plus a throwaway generation can stand in for your classifier calls, what else in your pipeline is quietly over-modeled?

tags: [ rag ] [ conversational-ai ] [ llm-ops ]