# Retrieval Model

Understand the two retrieval branches KuraDB runs, how semantic search narrows its candidates, and how query embeddings are cached.

## Dual retrieval

A search can execute keyword retrieval, semantic retrieval, or both.

| Strategy | Query preparation | Retrieval | Ranking |
|---|---|---|---|
| Keyword | gse tokenization, trimming, lowercase, deduplication | SQLite `LOWER(content) LIKE` across active content | Number of matched tokens, then row ID |
| Semantic | Query embedding from cache or OpenAI | Two-stage in-memory cosine search, then SQLite hydration | Cosine score, with a `0.3` minimum |

When both branches run, they execute concurrently and remain separate in the response. KuraDB does not merge their rankings into a single score. An error in either branch fails the whole request. Query handling is covered in detail in [Search and Retrieval](/search-and-retrieval).

## Source-level vector filtering

Semantic search derives one normalized vector per source from its chunk vectors. It first ranks all source vectors, keeps approximately 5% with a minimum of 20 candidates, then calculates chunk-level cosine similarity only within those sources. If source vectors are unavailable, it falls back to searching all chunk vectors. See [Semantic search](/search-and-retrieval-semantic) for the stage-by-stage algorithm.

## Query cache

`openai.Cache` maps the exact query string to an embedding. On a miss, semantic search calls OpenAI and stores the vector in memory. In the daemon, an `OnSet` callback asynchronously persists the encoded vector to `global.db` with a five-second timeout. Startup preloads valid entries without re-triggering persistence.

`kura mcp` preloads the same table but registers no `OnSet` callback, so query embeddings it computes live only for that session.
