Ternlight ships a 7 MB embedding model that runs in the browser on CPU — and the size is the least interesting part. What it changes is where the retrieval boundary sits.

The npm-install-and-go demo is genuinely three lines, and the HN discussion gets into recall tradeoffs versus server-side models. My question: what’s the smallest embedder that still holds R@10 on your corpus — because “runs in a browser tab” is worthless if it can’t find the right chunk?