The document engine for
AI applications
Turn PDF, Office, HTML, email and scans into a queryable, citable intermediate representation — then query it, pack a context window, recompile or redact it, and ground every answer in the exact page region.
Not a text extractor. A document compiler.
Every format becomes one intermediate representation you can query, pack, recompile, search and verify — deterministically, no model in the loop.
A queryable, citable IR
Every block gets a stable id like p2-b7. Select by type, heading, page or text — and resolve any citation back to the exact region on the page.
p1-b0p1-b1p1-b4Context, packed
Fit the highest-value blocks into a token budget — in reading order, with citations.
Recompile & redact
Deterministic IR→IR passes, then re-emit as IR, Markdown or text.
Search the library
Full-text across every stored doc — ranked, block-level, cited.
p1-b3…total revenue grew…Pixel-faithful pages
Born-digital text redrawn from the file's own embedded fonts.
Verifiable fidelity
A 0–1 score for how faithfully each document was reconstructed.
One pipeline. Every arrow is an interface.
Watch a document move from raw bytes to one intermediate representation — the same IR everything downstream reads.
Parse
Format parsers lower each document into positioned primitives — text with geometry, embedded fonts, bookmarks and links.
Not a black box — the playground streams every stage live as your file parses, page by page.
Built from the bytes up. Not a wrapper.
Most “document AI” is a thin client over a cloud API or a neural model. Cognita reads the bytes itself — the PDF object graph, the font outlines, the page raster — in one deterministic process. Nothing here calls out to a model or a server; your file never leaves the box.
// decode applies a PDF stream's filter chainswitch name {case "FlateDecode": return inflate(data)case "LZWDecode": return lzw(data)case "ASCII85Decode": return ascii85(data)case "RunLengthDecode": return runLength(data)case "CCITTFaxDecode": return ccitt(data, params)}