Teach it once: the corrections ledger
A Year column that meant years of experience broke our auto-built data model. Now one English sentence fixes it, and the fix survives every re-ingest.
Certant Strata is meant to build the data model from a pile of PDFs, with no analyst designing it first. On a 1,222-page corpus of Australian enterprise agreements (the run is in Building a data model from PDFs instead of designing one) that held until the pay rates, where the explorer showed five rows for "DDSO 1" with nothing to tell them apart.
On page 146 of the 155-page Aruma agreement, a column headed Year means years of experience within the classification. Our profiler labelled it a fiscal year, synthesis mapped it to a fiscal_year property with a TRY_CAST(x AS DATE), and the cast turned every value into NULL, even though extraction had captured all of them.
The table also carries four date-suffixed rate groups and synthesis canonicalised the first, so the unified view gave a DDSO 2 weekly rate of $1,330.33 instead of the July 2023 figure of $1,370.25. That is one 3% pay rise out of date, and nothing flagged it.
What we built lets an operator say what is wrong in one sentence, with the fix surviving every future ingest.
Where the Year column broke
We traced it stage by stage on the live KB on 12 June 2026, and only one stage got it wrong.
| Stage | What it did | Verdict |
|---|---|---|
| Extraction | Year / Year VARCHAR, values '1' to '8', clean in raw_extracted |
correct |
| Profiler, statistics | inferred_type: BIGINT, cardinality 10, null_rate 0.0 |
correct |
| Profiler, LLM annotation | semantic_type: "date_fiscal_year", is_temporal: true, while its own description reads "The sequential year number of the agreement period (1-5)" |
the error |
| Synthesis | Trusted the semantic type, mapped the column to fiscal_year with TRY_CAST(x AS DATE) |
propagated |
| View | TRY_CAST('3' AS DATE) returns NULL on every row |
the visible symptom |
Source: the evidence chain in FRD_STRATA_MODEL_CORRECTIONS.md §1.
The annotation model wrote a correct description, then filed the column under a contradictory label. Its controlled vocabulary (CORE_SEMANTIC_TYPES in profiler.py) has four temporal terms, four financial terms and nothing for an ordinal step or a tenure counter, so it picked the nearest option.
The mislabel caused three modelling problems in one table: the column needed re-binding to a new property, the grain was wrong (it should be classification plus experience year plus effective date), and the four Rate Effective FPPOOA / 1-Jul-2X column groups needed unpivoting into rows.
Why the fix cannot live in the prompt
The obvious fix was to put the correction into the next synthesis prompt. An audit of every existing human-input seam ruled that out.
| Seam | Storage | Durable? | Respected on re-synthesis? |
|---|---|---|---|
| Schema-match confirm/reject | graph nodes, match_method="manual" |
yes | yes, active_matches filters |
| Instance-key override | KB metadata strata_instance_keys |
yes | yes, resolved at read time |
| Semantic-type vocabulary | KB metadata strata_semantic_types |
yes | yes, extends the vocabulary |
| Property / transformation edits | none | no seam exists | next synthesis overwrites |
| ObjectType / LinkType structure edits | none | no seam exists | no |
Source: the durability audit in FRD_STRATA_MODEL_CORRECTIONS.md §3.
Synthesis is incremental by prompt: it gets the current active version and is told to extend it while keeping names stable. A model can ignore that, and re-fed its own output plus a soft hint it tends to re-litigate and flip-flop (arXiv 2310.01798). A correction held only in a prompt would eventually be undone by a re-ingest, unnoticed, because the number would still look plausible.
So human corrections are deterministic constraints applied by code, and the ledger is authoritative. The prompt also carries a NON-NEGOTIABLE HUMAN CORRECTIONS section so the model builds around the pinned facts, but nothing depends on it obeying.
In dbt a human owns the model source. In Strata the machine owns it and regenerates freely, so the human layer is an overlay replayed deterministically on every rebuild. A binding's effective definition is synthesis output plus the latest active directive, and a UI badge says which side owns it.
Describing the fix in plain English
The dialog opens from the Data Model page, any Object 360 instance page, and under any Ask answer. You type plain English:
"In the salary schedules, the Year column means years of experience within the classification, it is not a fiscal year or a date."
Strata returns typed correction directives (rebind_property, add_property, an environment variable, unpivot, extend_vocabulary), each quoting its measured evidence.

The evidence check exists because language models tend to agree with whatever is asserted confidently (arXiv 2310.13548). Every proposal must cite supporting profile statistics, and when they contradict the user the system says so in amber and records the disagreement in the directive's provenance. We chose warn-and-allow because the operator is the domain authority and can proceed anyway.
In that screenshot the proposer found no profile stats for the column, so it warned that the evidence was thin instead of quoting the values 1 to 8, 0% date-castable chain the design intends. It still flagged the claim for review instead of accepting it, but that is the weaker version of the feature.
Previewing and applying the fix
The preview compiles the candidate view body in memory and runs it read-only with a LIMIT, showing the change list and real sample rows with nothing persisted and no version minted.

Apply mints a new ontology version through the existing version-store and view-generation machinery, so the previous version stays immutable and rollback is a pointer move.

The acceptance test seeds a throwaway KB with the exact mis-bound projection synthesis produced, then drives the loop through the HTTP API against ground truth copied verbatim from the corpus:
| Check | Before correction | After correction |
|---|---|---|
| Rows in the CompensationRate view | 9 | 36 (9 rows x 4 rate groups) |
fiscal_year |
NULL on every row | still a column, all-NULL, unmapped rather than dropped |
years_of_experience |
does not exist | populated, part of the instance key |
Explorer filter classification_code = DDSO 1 |
5 unexplained rows | 20 instances (5 experience years x 4 effective dates) |
| DDSO 2, 4 years, 1-Jul-23 weekly rate | unreachable in the governed layer | $1,370.25 |
| Instance key | classification only | [classification_code, years_of_experience, effective_date] |
Source: maintained_tests/strata/test_corrections_schedule_j.py, which asserts each of these against ground-truth cells from the Aruma raw layer.
Three more cells show the unpivot produced distinct values per effective date: DDSO 1 at 1 year is $52,455 annual at 1-Jul-22 and $57,320 at 1-Jul-25, and DDSO 2 at 1 year is $35.48 an hour at 1-Jul-24. Before the correction the view carried only the 1-Jul-22 group, so the other two sat in the raw layer, out of reach of the governed views that Ask and the monitors read.
The DDSO 2 lookup now returns one row: a composite-key instance fetch for DDSO 2, four years, effective 2023-07-01 gives weekly_rate_aud 1370.25, with one provenance record pointing at Schedule J on page 146. Click it and the PDF opens on that page.
What the ledger records
Corrected types and properties carry a corrected chip, and fixed property rows a Human-corrected marker. I want that visible, because a model that shows what a human certified is easier to trust.

The Corrections tab is the ledger: every directive, who applied it, its evidence, and a Revoke button.

Append-only covers the ledger's writes; a correction can still be withdrawn. Each directive is active, revoked or superseded. Revoking appends a revoke event referencing the directive id, replay skips it, and the next version mint hands the binding back to the machine. Editing appends a new directive that supersedes the old one, and replay applies only the latest active directive per key. As in git, history is immutable and undo is a new commit.
The test also revokes all six directives, re-mints, and asserts the view returns to 9 rows with years_of_experience gone and the bug back, while the ledger keeps all six revoked events.
Replaying corrections after re-synthesis
Application is pure code with no model in the loop: apply_directives(projection, ledger) -> projection', run after every synthesis round and before the version is stored. When new documents land and synthesis re-runs, the corrections replay on top by construction.
The test runs the ledger replay over the original mis-bound projection and gets the corrected shape: years_of_experience bound, fiscal_year unmapped, the unpivot attached with its four column groups, and the composite grain intact.
The ledger and replay engine landed with 41 plus 12 pytest tests in a module suite of 378 with no regressions, the API with 16 tests plus a live test, and the UI with 22 new jest tests (209 total). The end-to-end acceptance case is 24 checks against real ground truth.
What it still cannot do
We scoped eight correction classes: five shipped whole, one in part, and two not at all.
| Class | Example | In v1? |
|---|---|---|
add_property |
expose a profiled but unmapped column | yes |
rebind_property |
the Year mis-cast | yes |
| rename / retype property | fiscal_year to years_of_experience |
retype and description only; rename is gated |
an environment variable |
grain fix | yes |
reassign_table |
table attached to the wrong type | v1.1 |
edit_link |
broken join condition | yes |
| split / merge types | one type that should be two | deferred |
unpivot |
the date-suffixed rate groups | shipped with this release |
Source: the correction taxonomy in FRD_STRATA_MODEL_CORRECTIONS.md §5, with v1 scope fixed by decision D5.
Rename has a high blast radius: monitors, instance-key overrides, saved execution plans and resolver prompts all key off property_name. Doing it properly needs Foundry-style move-edit migrations, so rename stays gated until those exist.
The bug register lists the unpivot as partly fixed: it is ground-truth-proven on Schedule J but fires only when a human asks. Automatic detection of date-suffixed column groups has not been built.
The profiler's core vocabulary still has no ordinal or tenure term. The decision record says the core list should gain one; what shipped is the per-KB route, an extend_vocabulary directive writing into that KB's strata_semantic_types. A fresh KB of enterprise agreements will repeat the mistake until corrected there.
The propose step's latency has not been measured. The sales-demo script expects 10 to 30 seconds for a real model call on that KB, which is an estimate.
The whole loop, broken projection to revoke, is in maintained_tests/strata/test_corrections_schedule_j.py.


