Retrieval Model
Understand the two retrieval branches KuraDB runs, how semantic search narrows its candidates, and how query embeddings are cached.
Dual retrieval
A search can execute keyword retrieval, semantic retrieval, or both.
| Strategy | Query preparation | Retrieval | Ranking |
|---|---|---|---|
| Keyword | gse tokenization, trimming, lowercase, deduplication | SQLite LOWER(content) LIKE across active content |
Number of matched tokens, then row ID |
| Semantic | Query embedding from cache or OpenAI | Two-stage in-memory cosine search, then SQLite hydration | Cosine score, with a 0.3 minimum |
When both branches run, they execute concurrently and remain separate in the response. KuraDB does not merge their rankings into a single score. An error in either branch fails the whole request. Query handling is covered in detail in Search and Retrieval.
Source-level vector filtering
Semantic search derives one normalized vector per source from its chunk vectors. It first ranks all source vectors, keeps approximately 5% with a minimum of 20 candidates, then calculates chunk-level cosine similarity only within those sources. If source vectors are unavailable, it falls back to searching all chunk vectors. See Semantic search for the stage-by-stage algorithm.
Query cache
openai.Cache maps the exact query string to an embedding. On a miss, semantic search calls OpenAI and stores the vector in memory. In the daemon, an OnSet callback asynchronously persists the encoded vector to global.db with a five-second timeout. Startup preloads valid entries without re-triggering persistence.
kura mcp preloads the same table but registers no OnSet callback, so query embeddings it computes live only for that session.