- π― The quiet tax in production RAG is memory, and Turbovec goes after it with Googleβs data-oblivious TurboQuant β a Rust index with Python bindings that needs no training step.
- π It compresses a 1536-dim embedding 16Γ (6,144 β 384 bytes at 2-bit), so a 10M-document corpus that eats 31GB as float32 fits in about 4GB of RAM.
- β‘ Benchmarks land 3.4Γ faster than FAISS IndexPQFastScan at 4-bit and ~20% at 2-bit, on hand-written SIMD kernels (NEON on ARM, AVX-512 VNNI on x86).
- π‘ No codebook to train means vectors index on ingest with no rebuilds as the corpus grows, and deletes are O(1) (~1Β΅s) instead of FAISS-style full repacking.
- β It ships LangChain, LlamaIndex, and Haystack bindings, so it slots into an existing retrieval stack rather than asking you to rebuild one β the kind of drop-in that actually gets tried instead of admired from a distance.
- π The catch: this is a young, largely single-author project, not a battle-tested library β Iβd run it against my own recall eval harness before trusting it in prod.
The HN discussion and the early write-ups are running hot. MarkTechPost frames it as a practical 16Γ win that beats FAISS on ARM with no codebook training, DuckDB Lab leads on the 1/8-the-RAM comparison, and Data Science in Your Pocket keeps circling back to the training-free simplicity; even Search Engine Land covered the underlying TurboQuant as a genuine speed improvement. The reception is near-uniformly positive β which, for a fresh quantization scheme, usually means the recall-versus-compression tradeoff just hasnβt been stress-tested widely enough yet.