Introduction
Cognita is a document-understanding engine. Send a file, get structured, AI-native JSON — an Intermediate Representation (headings, paragraphs, tables, lists, images, reading order, bounding boxes) plus ready-to-use Markdown, plain text and retrieval-ready chunks. Output is deterministic: identical input yields byte-identical output.
All endpoints are served under a single base URL:
Authentication
Every route except GET /health requires authentication. Create an account, generate an API key from your dashboard, and pass it in the X-API-Key header (a Bearer token also works).
Create an accountManage your key
Quickstart
Upload a document as multipart form data (field file) — or send the raw bytes as the request body. Both work on every parsing endpoint.
Parse a document
/v1/parseReturns the full IR plus rendered Markdown, text, metadata and one image per page.
| Parameter | Values | Default | ||
|---|---|---|---|---|
| include | comma list of document,markdown,text,metadata,page_images | all | Select response fields. | |
| image_data | omit | included | Strip embedded image bytes from the IR (dimensions/MIME stay). | |
| languages | e.g. eng+deu | server default | OCR language override for scanned pages. |
Response (200):
RAG chunks
/v1/chunkStructure-aware chunks that never split mid-block, carry heading breadcrumbs, and cite the exact pages and block IDs they came from. Accepts a file upload or a previously returned IR document as application/json.
| Parameter | Values | Default | ||
|---|---|---|---|---|
| max_chars | integer | 2000 | Target maximum characters per chunk. | |
| overlap | integer | 0 | Character overlap between adjacent chunks. |
Response (200):
Export / re-render
/v1/export?format=markdown|html|text|jsonRe-render an IR document (the document object from a parse response) to another format.
Stored documents
Parse once, then retrieve, export, chunk or fetch page images later.
Async jobs
For large files or long parses, submit a job instead of holding a request open. You get a job id immediately, poll for status, then fetch the result. The finished result is also saved to your library, so its document_id works with every stored-document route above.
/v1/jobsSame body and query parameters as /v1/parse. Returns 202 with the job; counts against your daily limit.
A job moves through pending → running → succeeded | failed.
Job object:
The IR
Every output is rendered from one common Intermediate Representation:
Block IDs are deterministic (p2-b7 = page 2, reading order 7). Geometry survives for PDFs and scans, so answers can be highlighted on the page image.
Rate limits
Accounts on the free tier get 4 successful parses per day, counted per user and reset at 00:00 UTC. Every parse response carries the current budget:
Once the limit is reached, parses return 429 Too Many Requests until the reset time.
Errors
Errors are uniform JSON with an HTTP status code:
| Status | Meaning | ||
|---|---|---|---|
| 401 | missing or invalid API key | ||
| 413 | upload too large | ||
| 415 | unsupported file format | ||
| 422 | unparseable document | ||
| 429 | daily parse limit reached |