One place that has read everything

Point Certant at the documents you already have. It reads them, works out how they connect, and shows the page behind every answer.

More than 120,000 pages processed in production.

The Certant knowledgebase builder
Sources in Builds the model Embeds it Ready to ask

The builder: sources in on the left, a searchable model out on the right.

The pipeline

Six steps, file to answer

The same six steps the product runs, in the order a document travels through them.

1 Bring documents in 2 Read every file 3 Split it sensibly 4 Match on meaning 5 Work out the links 6 Reconcile duplicates

Drag sideways to see the whole diagram

Bring in your documents. Drag files in, or pull from S3, SharePoint or Confluence. PDFs, Word, slides, spreadsheets, scans and images.
Read every file properly. A different reader for each kind of file, so tables stay tables and handwriting still gets read. The four readers are listed further down.
Split it sensibly. A table or a clause is never cut in half.
Match on meaning. Two passages that say the same thing in different words are found as related, not just ones that share a word. You can bring your own model for this.
Work out how it connects. People, employers, clauses, sites, claim types: named and linked to each other. Let Certant decide the types, define them yourself, or enforce them strictly.
Reconcile duplicates. "Staff", "Worker" and "Employee" end up as one thing rather than three.
What comes out

Fifty contracts, one consistent model

Certant reads every passage and pulls out the things that matter: people, organisations, clauses, salary bands, claim types. It then reconciles them, so "Staff", "Worker" and "Employee" become one searchable thing.

You choose how strict it is. Let Certant decide the types, define them yourself, or enforce them.
Duplicates get reconciled. One entity at a time, with a lock, so two documents describing the same thing end up as one.
One model, every way of asking. Search, chatbots, agents and dashboards all read from it.
Knowledge graph
A knowledge graph built from uploaded documents
Entity browser
Browsing one entity and its relationships
Who uses it

What people build with it

Research

A research corpus you can ask

A water authority indexes its whole library of research. "What do we know about PFAS in NSW?" comes back as one answer across many papers, each with its citation.

Compliance

Two rule sets, side by side

A bank loads the new prudential standard beside its existing policies. The model surfaces every clause that changed and every gap, with the page each came from.

HR

HR answers without the queue

An association loads every benefits, leave and procedure document. Employees answer their own questions, with a link to the exact policy section.

What comes with it

What comes with every knowledgebase

Reranking

Weak matches never reach the model

The cost is attributed per knowledgebase, and the providers are listed in Specifications below.

PageIndex

A quick first pass over a large knowledgebase

Certant narrows the field before retrieval starts, so size does not slow a question down.

PDF visualiser

Every citation opens the page

Click any citation, jump to the exact page, with the relevant passage highlighted. Relevance can be scored sentence by sentence.

Storage backends

Built for one desk or the whole company

File-based for a single user, or PostgreSQL for production scale. Both are listed further down.

A citation opened at the page, with the passage highlighted

Click a citation, land on the page.

Choosing how results are ranked

Ranking is configurable per knowledgebase.

For your IT team

What your engineers will ask about

Below the line, the real names. Every one of these is a setting, not a rewrite.

Readers

MinerU

GPU-grade layout and table extraction, local or hosted. Pipeline, VLM and hybrid backends. Best on PDFs, scans and complex layouts.

Docling

IBM's structured document parser, local CPU or remote API, with optional Granite Vision for picture descriptions. Best on Word, slides and structured PDFs.

PaddleOCR

Multilingual OCR with chart and formula recognition, configurable layout threshold. Best on CJK, right-to-left and mixed-script documents.

GLM-OCR

MIT-licensed open-source vision OCR that runs offline anywhere, with an optional faster variant. Fully sovereign.

Retrieval modes

Pick the mode per query, per chatbot, per agent.

Naive

Pure vector lookup. Fastest path. For simple factual questions where graph reasoning is overkill.

Local

Walk the knowledge graph one hop from named entities. Precise, context-aware follow-ups.

Global

Aggregate across the entire graph. For strategic, big-picture questions.

Hybrid

Local plus global, combined. Depth and breadth in one pass.

Mix

Graph context fused with vector chunks, then reranked. The default mode.

Mix-concise

Same evidence, brief answer. Operational replies from rich context.

Conversational

Keyword search (BM25) fused with vector search by reciprocal rank fusion, no entity extraction. Built for live chat.

Bypass

Model only, no retrieval. For meta questions about the knowledgebase itself.

Specifications

Chunk size

Token windows from 128 to 4,096, or structure-aware splitting.

Embedding model

Qwen3-Embedding-4B by default, 2,560 dimensions, batched. Bring your own.

Graph merge

A single sequential worker with distributed locking, and model-assisted summarisation when an entity's descriptions pile up.

Storage

File-based (NanoVectorDB and NetworkX) for one user, or PostgreSQL with pgvector and Apache AGE for production.

Reranking

Cohere, Jina, Aliyun or DeepInfra, with cost attributed per knowledgebase.

Your documents never leave the building

Every reader, model and store on this page has a version that runs inside your own walls.

Certant Cloud, your cloud, or fully air-gapped.

AnswerExample

Long service leave accrues after seven years of continuous service, and can be taken in two blocks with the employer's agreement.

Enterprise agreement · clause 22.1 · page 63 Leave policy v4.pdf

Every answer opens the document at the page it came from.