v0.6.0

Parsing

See which parser KuraDB selects for each changed file and which files never enter the pipeline.

Parser dispatch

A changed non-directory entry is matched by lowercase extension first, then by text sniffing:

Input Parser path Chunking behavior
.pdf go_pkg_parser.PDF Document parser chunks
.docx go_pkg_parser.Docx Document parser chunks
.pptx go_pkg_parser.PPTX Presentation parser chunks
.csv, .tsv go_pkg_parser.CSV through parseTabular Header-aware groups of five data rows
.xlsx go_pkg_parser.XLSX through parseTabular Header-aware groups of five data rows
Other valid UTF-8 text go_pkg_parser.Markdown Markdown/text chunks

Tabular chunks render each row as [n] header=value, ..., using colN when a header cell is empty and dropping cells beyond the header width. A sheet with fewer than two rows (header plus data) produces no chunks.

If a parser returns an error, the failure is logged and nothing is written, so rows from the previous successful parse stay active.

Skipped inputs

Rule Examples
Excluded by name .DS_Store
Image extensions .jpg, .png, .gif, .webp, .heic, .svg, .tiff, .avif
Video and audio extensions .mp4, .mov, .mkv, .mp3, .wav, .flac, .m4a, .opus
Archive and package extensions .zip, .tar, .gz, .7z, .rar, .iso, .dmg, .deb, .apk
Binary, bytecode, and library extensions .exe, .dll, .so, .dylib, .class, .jar, .pyc, .wasm
Database, design, and font extensions .db, .sqlite, .psd, .sketch, .fig, .ttf, .woff2
Fails text sniffing Empty files, files with a NUL byte or invalid UTF-8 in the first 8,192 bytes

Skipping keeps image base64 and unsupported binary data out of the text embedding pipeline. A skipped file is not dismissed: if a previously indexed text file is emptied, it fails sniffing and its earlier chunks remain searchable until the file is deleted or receives new text.

中文