v0.6.0

Data Consistency

Understand which store is authoritative for each piece of KuraDB state and how chunks and embeddings stay consistent with it.

SQLite as source of truth

Each database stores parsed chunks and embeddings in {db}/data.db. SQLite controls whether a row is active, pending embedding, or dismissed. In-memory caches improve query speed but are disposable:

State Durable source Rebuild behavior
Chunk content file_data.content Never sourced from vector cache
Embedding file_data.embedding Loaded into vector buckets at startup
Query embedding query_cache.embedding in global.db Preloaded into openai.Cache
Source vectors Derived in memory Rebuilt by summing and normalizing chunk vectors
Watch snapshot record.json Loaded before the next directory comparison

Search results are always hydrated from SQLite with dismiss = FALSE, so a vector that is still in memory cannot return content SQLite no longer considers active.

Chunk lifecycle

A parsed file produces numbered chunks keyed by (source, chunk). Upsert first marks existing active rows for the source as dismissed, then inserts or updates current chunks in one transaction. If a chunk's content is unchanged, its embedding is retained; changed content clears the embedding and resets is_embed to FALSE.

Deleted files are not physically removed. The watcher calls Dismiss, which marks matching rows with dismiss = TRUE. Every active query path filters those rows. The write path is described step by step in Storage writes.

Embedding consistency

The OpenAI client requests text-embedding-3-small vectors with 512 dimensions. Encoded vectors use little-endian float32, so one valid blob is 2,048 bytes. KuraDB enforces consistency at several points:

中文