# Semantic Search

How the semantic branch embeds a query, narrows candidates with source vectors, ranks chunks, and hydrates hits from SQLite.

## Semantic query embedding

The semantic branch first checks `openai.Cache` with the exact query string. On a cache miss it sends a one-item batch to OpenAI (`text-embedding-3-small`, 512 dimensions) and stores the vector in memory.

| Process | What happens after a cache miss |
|---|---|
| Daemon | The cache's `OnSet` callback persists the vector to `global.db` asynchronously, with a 5-second timeout |
| `kura mcp` | No `OnSet` callback is registered; the vector lives in memory for that session only |

Both processes preload query-cache rows from `global.db` at startup, keeping only blobs whose byte length equals `openai.Dim() * 4`. Preload bypasses `OnSet`, so it does not rewrite the same entry.

## Two-stage vector search

Each database bucket stores chunk vectors and one derived vector per source: the L2-normalized sum of that source's same-dimension chunk vectors, rebuilt whenever its chunks change.

### Stage 1: source candidates

KuraDB computes cosine similarity between the query and every same-dimension source vector. The candidate count is:

```text
min(max(number of sources / 20, 20), number of sources)
```

This keeps about 5% of a large source set while retaining at least 20 when available.

### Stage 2: chunk ranking

KuraDB collects chunk IDs from candidate sources and calculates same-dimension cosine similarity for each chunk. Work is split across `min(CPU count - 1, chunk count / 200, chunk count)` workers, never fewer than one, so each worker handles at least 200 chunks when the set is large enough. Hits are sorted by descending score and truncated to `topK`, which equals the request limit.

If no source vectors exist, search falls back to ranking every chunk vector directly.

## Score filtering and hydration

The search core removes semantic hits whose score is below `0.3`. It then fetches the remaining IDs from SQLite with `dismiss = FALSE`, restores vector-ranking order, and drops any ID that no longer resolves to an active row.

This final hydration keeps SQLite authoritative even when in-memory state changes concurrently. How the surviving rows are grouped and returned is described in [Search Results](/search-and-retrieval-results).
