One place that has read everything
Point Certant at the documents you already have. It reads them, works out how they connect, and shows the page behind every answer.
The builder: sources in on the left, a searchable model out on the right.
Six steps, file to answer
The same six steps the product runs, in the order a document travels through them.
Drag sideways to see the whole diagram
Fifty contracts, one consistent model
Certant reads every passage and pulls out the things that matter: people, organisations, clauses, salary bands, claim types. It then reconciles them, so "Staff", "Worker" and "Employee" become one searchable thing.
What people build with it
A research corpus you can ask
A water authority indexes its whole library of research. "What do we know about PFAS in NSW?" comes back as one answer across many papers, each with its citation.
Two rule sets, side by side
A bank loads the new prudential standard beside its existing policies. The model surfaces every clause that changed and every gap, with the page each came from.
HR answers without the queue
An association loads every benefits, leave and procedure document. Employees answer their own questions, with a link to the exact policy section.
What comes with every knowledgebase
Weak matches never reach the model
The cost is attributed per knowledgebase, and the providers are listed in Specifications below.
A quick first pass over a large knowledgebase
Certant narrows the field before retrieval starts, so size does not slow a question down.
Every citation opens the page
Click any citation, jump to the exact page, with the relevant passage highlighted. Relevance can be scored sentence by sentence.
Built for one desk or the whole company
File-based for a single user, or PostgreSQL for production scale. Both are listed further down.
Click a citation, land on the page.
Ranking is configurable per knowledgebase.
What your engineers will ask about
Below the line, the real names. Every one of these is a setting, not a rewrite.
Readers
GPU-grade layout and table extraction, local or hosted. Pipeline, VLM and hybrid backends. Best on PDFs, scans and complex layouts.
IBM's structured document parser, local CPU or remote API, with optional Granite Vision for picture descriptions. Best on Word, slides and structured PDFs.
Multilingual OCR with chart and formula recognition, configurable layout threshold. Best on CJK, right-to-left and mixed-script documents.
MIT-licensed open-source vision OCR that runs offline anywhere, with an optional faster variant. Fully sovereign.
Retrieval modes
Pick the mode per query, per chatbot, per agent.
Pure vector lookup. Fastest path. For simple factual questions where graph reasoning is overkill.
Walk the knowledge graph one hop from named entities. Precise, context-aware follow-ups.
Aggregate across the entire graph. For strategic, big-picture questions.
Local plus global, combined. Depth and breadth in one pass.
Graph context fused with vector chunks, then reranked. The default mode.
Same evidence, brief answer. Operational replies from rich context.
Keyword search (BM25) fused with vector search by reciprocal rank fusion, no entity extraction. Built for live chat.
Model only, no retrieval. For meta questions about the knowledgebase itself.
Specifications
Token windows from 128 to 4,096, or structure-aware splitting.
Qwen3-Embedding-4B by default, 2,560 dimensions, batched. Bring your own.
A single sequential worker with distributed locking, and model-assisted summarisation when an entity's descriptions pile up.
File-based (NanoVectorDB and NetworkX) for one user, or PostgreSQL with pgvector and Apache AGE for production.
Cohere, Jina, Aliyun or DeepInfra, with cost attributed per knowledgebase.
Your documents never leave the building
Every reader, model and store on this page has a version that runs inside your own walls.
Long service leave accrues after seven years of continuous service, and can be taken in two blocks with the employer's agreement.
Every answer opens the document at the page it came from.
