A search regression in turbovec, found with PubMed embeddings | PharmaTools.AI
← Publications · Notes
Note · Open source · Evaluation · 7 Oct 2026

A search regression in turbovec, found with PubMed embeddings

turbovec is a fast, compact vector index for retrieval: it stores embeddings at 2 or 4 bits per dimension and searches them directly. Version 1.1.0 added a “staged” search, which shortlists candidates with a cheap first pass and then scores only the shortlist properly. The docs reported that this matched a full scan almost exactly, and asked users to check on their own data. So I did, using biomedical data.

What I tested

I took 101,000 PubMed abstracts embedded with NCBI’s MedCPT model, plus the project’s own OpenAI benchmark data as a control. I compared the staged search against a full scan over 3,000 queries at both bit widths.

What I found

On the OpenAI data, the two searches agreed 100% of the time, as documented. On PubMed, at k=10 (the usual setting for RAG), they returned identical top-10 results for only 82.6% of queries. Top-10 recall dropped from 0.944 to 0.931, and for 0.37% of queries the true best match was missing entirely. In one example, a corrigendum’s nearest neighbour was the paper it corrects. Asking for 64 results found it at rank 1, but asking for 10 left it out.

The cause was the shape of the data. MedCPT vectors are tightly bunched (mean pairwise cosine 0.65, against 0.09 for OpenAI), so the cheap first pass struggled to separate true neighbours. Subtracting the corpus’s average vector restored much of the agreement, which pointed to the shared direction as a cause.

What happened next

I posted a reproducible report with scripts and data. The maintainer reproduced it within hours, confirmed it was a regression introduced in 1.1.0, and designed a fix. The fix makes the first pass correct for exactly that shared direction. He also took up my offer to test a general-purpose model. bge-small turned out to be affected even more (75.6% agreement), so the problem wasn’t specific to biomedical models.

The fix shipped in turbovec 1.1.2. I verified it independently:

BeforeAfter
MedCPT, 4-bit, k=10 agreement82.6%99.1%
bge-small, 4-bit, k=10 agreement75.6%96.6%
True best match missing0.37–0.53%0%
OpenAI (control)100%100%

MedCPT is now one of the corpora in turbovec’s docs, and I’m contributing benchmark tests so the project checks bunched data on every release.

What I took from it

I ran the investigation with Claude as a research partner. The experiment design, judgement calls and conclusions are mine. Scripts and results: turbovec-biomed-bench.