How to Test Document AI Before Trusting Answers
A practical method for evaluating document AI with traceable citations, refusal tests, human review and inspectable calculations.
Fluent answers are easy to demonstrate. Defensible answers take a much harder test.
When I evaluate document AI, I want to know whether a reviewer can move from a response back to the evidence, understand the limits of that evidence, and see how the system handled anything unresolved. That standard matters whenever an answer may influence a policy decision, a member entitlement, a compliance review or an operational action.
This is the evaluation method I would use before trusting a system with consequential document work. It is a proposed test framework, not a reported benchmark.
Start with a question you can mark
An open-ended demonstration gives you plenty to discuss and very little to mark. “Summarise these documents” may be useful during exploration, but it does not give a reviewer a clear pass or fail condition.
For an evaluation, choose a narrow question tied to a real decision. Before running the system, write down what a defensible answer would contain:
- the answer itself,
- the passages that support it,
- the qualifications attached to those passages,
- the information the documents do not establish, and
- the conditions that would require clarification or refusal.
Also decide which errors make the response unusable. A clumsy sentence and a missing eligibility condition should not receive the same treatment. The test needs to reflect the consequences of getting the answer wrong.
I would begin with a small set of decision-relevant questions, not a large pile of documents and a vague request for a demonstration. A narrow test makes it possible to explain why each response passed or failed.
Check whether the citation proves the claim
A citation should let you verify a specific statement. Finding the same words, or the same number, somewhere on a page is not enough.
Imagine a policy passage that describes an entitlement, followed by a condition that limits who qualifies. An answer can reproduce the entitlement accurately while omitting the condition. The words may look familiar and the citation may point to the right document, but the answer is still misleading.
For each material claim, I would ask a reviewer to check:
- Does the cited passage support the wording of the answer?
- Does the answer preserve the conditions attached to the passage?
- Is the reference precise enough to find without searching the whole document?
- Does the response identify relevant uncertainty or missing context?
The useful measure is the work left for the reviewer. A citation has value when it reduces the time required to check the claim and makes disagreement easy to investigate.
Test a question the documents cannot answer
A trustworthy system needs a boundary. I would deliberately include a question for which the supplied documents do not provide enough evidence.
The expected behaviour should be decided before the test. Depending on the workflow, the right response may be to explain what is missing, ask for a particular document, or stop. It should not be an unsupported conclusion presented with confidence.
I would accept a response that says the supplied material does not establish the answer and identifies the unresolved requirement. That may send the work back to a person, but it keeps the decision connected to the available evidence.
This case belongs in the acceptance criteria. Otherwise, an evaluation can reward a complete-looking answer when the correct result was to leave the question unresolved.
Preserve disagreement between documents
Many document collections contain different versions, overlapping authorities or instructions that do not agree. I would test that deliberately.
Retrieving both passages is one requirement. Deciding which passage governs the answer is another. The system should have an explicit basis for that decision, such as a documented precedence rule, an effective date or a named authority.
If no such rule exists, the expected response should identify the disagreement and show the relevant passages. It should not quietly turn an assumption about authority into part of the answer.
This is where I separate retrieval from judgement. A system may find the relevant evidence without having the authority to decide which document controls the situation. A reviewer needs to see that distinction.
Make calculations inspectable
For a numerical answer, I want to see the inputs and the operation, not only the final figure.
A source reference beside the result can still leave too much checking to do. I would ask the system to identify where each input came from, show the units, explain how it combined the values and make any material assumption visible.
The calculation should be reproducible without turning the answer into a page of unnecessary prose. Straightforward arithmetic needs a short explanation. It still needs enough information for another person to check whether the operation answered the right question.
A mathematically correct operation can fail if it combines values from different scopes or ignores a qualification in the source. I would test the meaning of the inputs before judging the arithmetic.
Define the human reviewer’s job
“Human review” is not an acceptance criterion on its own. The evaluation needs to say what the reviewer checks and what authority they have to accept, reject or escalate the output.
I would separate two decisions:
- Is the answer supported by the available evidence?
- Is acting on that answer appropriate in this situation?
Those decisions may require different knowledge or authority. A response can pass the evidence check while leaving an operational decision unresolved.
The reviewer should be able to inspect the relevant source without repeating the entire retrieval task. During the evaluation, I would record where review stops being an evidence check and becomes a decision that belongs to a person.
Measure the work left after the response
If the purpose of document AI is to reduce document work, stop the clock after the answer has been checked, not when text first appears.
I would record response time separately from review time. Corrections, unresolved questions and escalations should remain visible. A slow response with clear evidence presents a different problem from a fast response that requires extensive manual checking.
The evaluation should also record why reviewers reject answers. Grouping failures by requirement gives the next test a purpose. You can see whether a change improved citation precision, refusal behaviour, conflict handling or calculation traceability.
I would avoid publishing a time-saving or accuracy claim from an exercise that only checked whether the final answer looked plausible. The test has to measure the work the proposed workflow is meant to change.
Questions to settle before a trial
Does every sentence need a citation?
I would require evidence for claims drawn from the documents, especially claims that could affect a decision. The practical standard is that a reviewer can identify the support for each material claim without guessing which reference applies.
Should the system refuse whenever there is uncertainty?
I would distinguish a qualified answer from a conclusion the evidence cannot support. The evaluation should define which uncertainties permit a qualified response and which require clarification or a stop.
Can a correct answer fail the evaluation?
Yes. If the system cannot show adequate support, I would mark the evidence requirement as failed even when the answer happens to match the reference. The point is to test whether the response is defensible, not whether it can sometimes guess correctly.
Where should a limited-budget trial begin?
Choose a narrow document set and one decision-relevant question. Establish the expected answer and supporting passages, then add an unanswerable case and a conflict case. Broaden the exercise only after you can explain why each response passed or failed.
The standard I would use
Before trusting document AI, I would want to answer four questions:
- Can a reviewer trace each material claim to the evidence?
- Does the system show what the documents do not establish?
- Can it preserve conflicts instead of hiding them?
- Can a person reproduce the reasoning behind a consequential result?
If the answer to any of these is no, the system may still be useful for exploration. I would not treat it as ready for decisions that need to stand up to checking.
Frequently asked questions
Does every sentence need a citation?
I would require evidence for claims drawn from the documents, especially claims that could affect a decision. The practical standard is that a reviewer can identify the support for each material claim without guessing which reference applies.
Should the system refuse whenever there is uncertainty?
I would distinguish a qualified answer from a conclusion the evidence cannot support. The evaluation should define which uncertainties permit a qualified response and which require clarification or a stop.
Can a correct answer fail the evaluation?
Yes. If the system cannot show adequate support, I would mark the evidence requirement as failed even when the answer happens to match the reference. The point is to test whether the response is defensible.
Where should a limited-budget trial begin?
Choose a narrow document set and one decision-relevant question. Establish the expected answer and supporting passages, then add an unanswerable case and a conflict case. Broaden the exercise only after you can explain why each response passed or failed.



