Parsing
See which parser KuraDB selects for each changed file and which files never enter the pipeline.
Parser dispatch
A changed non-directory entry is matched by lowercase extension first, then by text sniffing:
| Input | Parser path | Chunking behavior |
|---|---|---|
.pdf |
go_pkg_parser.PDF |
Document parser chunks |
.docx |
go_pkg_parser.Docx |
Document parser chunks |
.pptx |
go_pkg_parser.PPTX |
Presentation parser chunks |
.csv, .tsv |
go_pkg_parser.CSV through parseTabular |
Header-aware groups of five data rows |
.xlsx |
go_pkg_parser.XLSX through parseTabular |
Header-aware groups of five data rows |
| Other valid UTF-8 text | go_pkg_parser.Markdown |
Markdown/text chunks |
Tabular chunks render each row as [n] header=value, ..., using colN when a header cell is empty and dropping cells beyond the header width. A sheet with fewer than two rows (header plus data) produces no chunks.
If a parser returns an error, the failure is logged and nothing is written, so rows from the previous successful parse stay active.
Skipped inputs
| Rule | Examples |
|---|---|
| Excluded by name | .DS_Store |
| Image extensions | .jpg, .png, .gif, .webp, .heic, .svg, .tiff, .avif |
| Video and audio extensions | .mp4, .mov, .mkv, .mp3, .wav, .flac, .m4a, .opus |
| Archive and package extensions | .zip, .tar, .gz, .7z, .rar, .iso, .dmg, .deb, .apk |
| Binary, bytecode, and library extensions | .exe, .dll, .so, .dylib, .class, .jar, .pyc, .wasm |
| Database, design, and font extensions | .db, .sqlite, .psd, .sketch, .fig, .ttf, .woff2 |
| Fails text sniffing | Empty files, files with a NUL byte or invalid UTF-8 in the first 8,192 bytes |
Skipping keeps image base64 and unsupported binary data out of the text embedding pipeline. A skipped file is not dismissed: if a previously indexed text file is emptied, it fails sniffing and its earlier chunks remain searchable until the file is deleted or receives new text.