Certant StrataOntologyData ModellingLLM

Teach it once: the corrections ledger

A Year column that meant years of experience broke our auto-built data model. Now one English sentence fixes it, and the fix survives every re-ingest.

Daniel Voyce··9 min read

Certant Strata is meant to build the data model from a pile of PDFs, with no analyst designing it first. On a 1,222-page corpus of Australian enterprise agreements (the run is in Building a data model from PDFs instead of designing one) that held until the pay rates, where the explorer showed five rows for "DDSO 1" with nothing to tell them apart.

On page 146 of the 155-page Aruma agreement, a column headed Year means years of experience within the classification. Our profiler labelled it a fiscal year, synthesis mapped it to a fiscal_year property with a TRY_CAST(x AS DATE), and the cast turned every value into NULL, even though extraction had captured all of them.

The table also carries four date-suffixed rate groups and synthesis canonicalised the first, so the unified view gave a DDSO 2 weekly rate of $1,330.33 instead of the July 2023 figure of $1,370.25. That is one 3% pay rise out of date, and nothing flagged it.

What we built lets an operator say what is wrong in one sentence, with the fix surviving every future ingest.

Where the Year column broke

We traced it stage by stage on the live KB on 12 June 2026, and only one stage got it wrong.

Stage What it did Verdict
Extraction Year / Year VARCHAR, values '1' to '8', clean in raw_extracted correct
Profiler, statistics inferred_type: BIGINT, cardinality 10, null_rate 0.0 correct
Profiler, LLM annotation semantic_type: "date_fiscal_year", is_temporal: true, while its own description reads "The sequential year number of the agreement period (1-5)" the error
Synthesis Trusted the semantic type, mapped the column to fiscal_year with TRY_CAST(x AS DATE) propagated
View TRY_CAST('3' AS DATE) returns NULL on every row the visible symptom

Source: the evidence chain in FRD_STRATA_MODEL_CORRECTIONS.md §1.

The annotation model wrote a correct description, then filed the column under a contradictory label. Its controlled vocabulary (CORE_SEMANTIC_TYPES in profiler.py) has four temporal terms, four financial terms and nothing for an ordinal step or a tenure counter, so it picked the nearest option.

The mislabel caused three modelling problems in one table: the column needed re-binding to a new property, the grain was wrong (it should be classification plus experience year plus effective date), and the four Rate Effective FPPOOA / 1-Jul-2X column groups needed unpivoting into rows.

Why the fix cannot live in the prompt

The obvious fix was to put the correction into the next synthesis prompt. An audit of every existing human-input seam ruled that out.

Seam Storage Durable? Respected on re-synthesis?
Schema-match confirm/reject graph nodes, match_method="manual" yes yes, active_matches filters
Instance-key override KB metadata strata_instance_keys yes yes, resolved at read time
Semantic-type vocabulary KB metadata strata_semantic_types yes yes, extends the vocabulary
Property / transformation edits none no seam exists next synthesis overwrites
ObjectType / LinkType structure edits none no seam exists no

Source: the durability audit in FRD_STRATA_MODEL_CORRECTIONS.md §3.

Synthesis is incremental by prompt: it gets the current active version and is told to extend it while keeping names stable. A model can ignore that, and re-fed its own output plus a soft hint it tends to re-litigate and flip-flop (arXiv 2310.01798). A correction held only in a prompt would eventually be undone by a re-ingest, unnoticed, because the number would still look plausible.

So human corrections are deterministic constraints applied by code, and the ledger is authoritative. The prompt also carries a NON-NEGOTIABLE HUMAN CORRECTIONS section so the model builds around the pinned facts, but nothing depends on it obeying.

In dbt a human owns the model source. In Strata the machine owns it and regenerates freely, so the human layer is an overlay replayed deterministically on every rebuild. A binding's effective definition is synthesis output plus the latest active directive, and a UI badge says which side owns it.

Describing the fix in plain English

The dialog opens from the Data Model page, any Object 360 instance page, and under any Ask answer. You type plain English:

"In the salary schedules, the Year column means years of experience within the classification, it is not a fiscal year or a date."

Strata returns typed correction directives (rebind_property, add_property, an environment variable, unpivot, extend_vocabulary), each quoting its measured evidence.

Certant Strata "Fix the model" dialog showing typed correction proposals, each with the column evidence quoted underneath and an amber warn-and-allow callout
The proposal cards from the live enterprise-agreement KB, captured 12 June 2026. Each card is a typed directive citing its evidence; the amber note is the evidence check pushing back.

The evidence check exists because language models tend to agree with whatever is asserted confidently (arXiv 2310.13548). Every proposal must cite supporting profile statistics, and when they contradict the user the system says so in amber and records the disagreement in the directive's provenance. We chose warn-and-allow because the operator is the domain authority and can proceed anyway.

In that screenshot the proposer found no profile stats for the column, so it warned that the evidence was thin instead of quoting the values 1 to 8, 0% date-castable chain the design intends. It still flagged the claim for review instead of accepting it, but that is the weaker version of the feature.

Previewing and applying the fix

The preview compiles the candidate view body in memory and runs it read-only with a LIMIT, showing the change list and real sample rows with nothing persisted and no version minted.

Preview panel listing six change lines with a table of sample rows from the corrected view
The preview for the Schedule J bundle, 12 June 2026: the change list plus real rows from the candidate view, compiled in memory and persisting nothing.

Apply mints a new ontology version through the existing version-store and view-generation machinery, so the previous version stays immutable and rollback is a pointer move.

Object Explorer showing CompensationRate rows with years_of_experience and effective_date columns populated
The same explorer after apply, 12 June 2026. The five unexplained DDSO rows have become years_of_experience by effective_date, and the four annual pay rises are real rows.

The acceptance test seeds a throwaway KB with the exact mis-bound projection synthesis produced, then drives the loop through the HTTP API against ground truth copied verbatim from the corpus:

Check Before correction After correction
Rows in the CompensationRate view 9 36 (9 rows x 4 rate groups)
fiscal_year NULL on every row still a column, all-NULL, unmapped rather than dropped
years_of_experience does not exist populated, part of the instance key
Explorer filter classification_code = DDSO 1 5 unexplained rows 20 instances (5 experience years x 4 effective dates)
DDSO 2, 4 years, 1-Jul-23 weekly rate unreachable in the governed layer $1,370.25
Instance key classification only [classification_code, years_of_experience, effective_date]

Source: maintained_tests/strata/test_corrections_schedule_j.py, which asserts each of these against ground-truth cells from the Aruma raw layer.

Three more cells show the unpivot produced distinct values per effective date: DDSO 1 at 1 year is $52,455 annual at 1-Jul-22 and $57,320 at 1-Jul-25, and DDSO 2 at 1 year is $35.48 an hour at 1-Jul-24. Before the correction the view carried only the 1-Jul-22 group, so the other two sat in the raw layer, out of reach of the governed views that Ask and the monitors read.

The DDSO 2 lookup now returns one row: a composite-key instance fetch for DDSO 2, four years, effective 2023-07-01 gives weekly_rate_aud 1370.25, with one provenance record pointing at Schedule J on page 146. Click it and the PDF opens on that page.

What the ledger records

Corrected types and properties carry a corrected chip, and fixed property rows a Human-corrected marker. I want that visible, because a model that shows what a human certified is easier to trust.

Data Model cards showing a corrected chip on CompensationRate and Human-corrected markers on individual property rows
Provenance badges on the corrected type and its properties, 12 June 2026. Origin is one of synthesized, AI-proposed and human-approved, or human-corrected.

The Corrections tab is the ledger: every directive, who applied it, its evidence, and a Revoke button.

The Corrections tab listing applied directives with author, evidence and Revoke buttons
The append-only corrections ledger on the same KB, 12 June 2026. Nothing is deleted; revoking appends an event.

Append-only covers the ledger's writes; a correction can still be withdrawn. Each directive is active, revoked or superseded. Revoking appends a revoke event referencing the directive id, replay skips it, and the next version mint hands the binding back to the machine. Editing appends a new directive that supersedes the old one, and replay applies only the latest active directive per key. As in git, history is immutable and undo is a new commit.

The test also revokes all six directives, re-mints, and asserts the view returns to 9 rows with years_of_experience gone and the bug back, while the ledger keeps all six revoked events.

Replaying corrections after re-synthesis

Application is pure code with no model in the loop: apply_directives(projection, ledger) -> projection', run after every synthesis round and before the version is stored. When new documents land and synthesis re-runs, the corrections replay on top by construction.

The test runs the ledger replay over the original mis-bound projection and gets the corrected shape: years_of_experience bound, fiscal_year unmapped, the unpivot attached with its four column groups, and the composite grain intact.

The ledger and replay engine landed with 41 plus 12 pytest tests in a module suite of 378 with no regressions, the API with 16 tests plus a live test, and the UI with 22 new jest tests (209 total). The end-to-end acceptance case is 24 checks against real ground truth.

What it still cannot do

We scoped eight correction classes: five shipped whole, one in part, and two not at all.

Class Example In v1?
add_property expose a profiled but unmapped column yes
rebind_property the Year mis-cast yes
rename / retype property fiscal_year to years_of_experience retype and description only; rename is gated
an environment variable grain fix yes
reassign_table table attached to the wrong type v1.1
edit_link broken join condition yes
split / merge types one type that should be two deferred
unpivot the date-suffixed rate groups shipped with this release

Source: the correction taxonomy in FRD_STRATA_MODEL_CORRECTIONS.md §5, with v1 scope fixed by decision D5.

Rename has a high blast radius: monitors, instance-key overrides, saved execution plans and resolver prompts all key off property_name. Doing it properly needs Foundry-style move-edit migrations, so rename stays gated until those exist.

The bug register lists the unpivot as partly fixed: it is ground-truth-proven on Schedule J but fires only when a human asks. Automatic detection of date-suffixed column groups has not been built.

The profiler's core vocabulary still has no ordinal or tenure term. The decision record says the core list should gain one; what shipped is the per-KB route, an extend_vocabulary directive writing into that KB's strata_semantic_types. A fresh KB of enterprise agreements will repeat the mistake until corrected there.

The propose step's latency has not been measured. The sales-demo script expects 10 to 30 seconds for a real model call on that KB, which is an estimate.

The whole loop, broken projection to revoke, is in maintained_tests/strata/test_corrections_schedule_j.py.

Build a brain for your business.

Certant turns your documents, data and processes into agents, dashboards and assistants you can actually trust.