Semantic Search
How the semantic branch embeds a query, narrows candidates with source vectors, ranks chunks, and hydrates hits from SQLite.
Semantic query embedding
The semantic branch first checks openai.Cache with the exact query string. On a cache miss it sends a one-item batch to OpenAI (text-embedding-3-small, 512 dimensions) and stores the vector in memory.
| Process | What happens after a cache miss |
|---|---|
| Daemon | The cache's OnSet callback persists the vector to global.db asynchronously, with a 5-second timeout |
kura mcp |
No OnSet callback is registered; the vector lives in memory for that session only |
Both processes preload query-cache rows from global.db at startup, keeping only blobs whose byte length equals openai.Dim() * 4. Preload bypasses OnSet, so it does not rewrite the same entry.
Two-stage vector search
Each database bucket stores chunk vectors and one derived vector per source: the L2-normalized sum of that source's same-dimension chunk vectors, rebuilt whenever its chunks change.
Stage 1: source candidates
KuraDB computes cosine similarity between the query and every same-dimension source vector. The candidate count is:
min(max(number of sources / 20, 20), number of sources)
This keeps about 5% of a large source set while retaining at least 20 when available.
Stage 2: chunk ranking
KuraDB collects chunk IDs from candidate sources and calculates same-dimension cosine similarity for each chunk. Work is split across min(CPU count - 1, chunk count / 200, chunk count) workers, never fewer than one, so each worker handles at least 200 chunks when the set is large enough. Hits are sorted by descending score and truncated to topK, which equals the request limit.
If no source vectors exist, search falls back to ranking every chunk vector directly.
Score filtering and hydration
The search core removes semantic hits whose score is below 0.3. It then fetches the remaining IDs from SQLite with dismiss = FALSE, restores vector-ranking order, and drops any ID that no longer resolves to an active row.
This final hydration keeps SQLite authoritative even when in-memory state changes concurrently. How the surviving rows are grouped and returned is described in Search Results.