Data Consistency
Understand which store is authoritative for each piece of KuraDB state and how chunks and embeddings stay consistent with it.
SQLite as source of truth
Each database stores parsed chunks and embeddings in {db}/data.db. SQLite controls whether a row is active, pending embedding, or dismissed. In-memory caches improve query speed but are disposable:
| State | Durable source | Rebuild behavior |
|---|---|---|
| Chunk content | file_data.content |
Never sourced from vector cache |
| Embedding | file_data.embedding |
Loaded into vector buckets at startup |
| Query embedding | query_cache.embedding in global.db |
Preloaded into openai.Cache |
| Source vectors | Derived in memory | Rebuilt by summing and normalizing chunk vectors |
| Watch snapshot | record.json |
Loaded before the next directory comparison |
Search results are always hydrated from SQLite with dismiss = FALSE, so a vector that is still in memory cannot return content SQLite no longer considers active.
Chunk lifecycle
A parsed file produces numbered chunks keyed by (source, chunk). Upsert first marks existing active rows for the source as dismissed, then inserts or updates current chunks in one transaction. If a chunk's content is unchanged, its embedding is retained; changed content clears the embedding and resets is_embed to FALSE.
Deleted files are not physically removed. The watcher calls Dismiss, which marks matching rows with dismiss = TRUE. Every active query path filters those rows. The write path is described step by step in Storage writes.
Embedding consistency
The OpenAI client requests text-embedding-3-small vectors with 512 dimensions. Encoded vectors use little-endian float32, so one valid blob is 2,048 bytes. KuraDB enforces consistency at several points:
- OpenAI responses must contain the requested number of vectors.
- Every returned vector must have exactly 512 elements.
- Query-cache preload skips blobs whose byte length differs from
openai.Dim() * 4. - Vector search and source-vector rebuilds ignore vectors whose dimensions differ from the query or the source.
- Embedding updates include the original content in the SQL predicate, preventing stale results from being applied after a file changes.