Certant StrataMonitorsApprovalsOntology

From a threshold breach to an approved action

How a Strata monitor detects a breach, counts the blast radius from the ontology, and parks the fix in an approvals inbox until a person signs it off.

Daniel Voyce··10 min read

An unattended monitor run used to reach the human-in-the-loop gate in our agent runtime, wait up to 600 seconds for a person who was never going to arrive, then quietly take the reject path.

We closed that gap with detection that fires once per change, a blast radius counted in SQL (it used to come from a language model), a suggested action, and an approval gate a person has to pass before any external write happens. We also shipped a bug on the way, covered at the end.

Detection runs on a 5-minute clock

Strata monitors sit over the governed ontology views: two shipped types and, since segment 056, four selectable ones in the create form. A Celery-beat sweep evaluates every enabled monitor every 5 minutes, so new table data from a document, a re-extraction or an agent feed is picked up on the next tick.

The Monitors panel with the edit-monitor form and the demo monitors below
The Monitors panel on the Aqua Valley demo knowledgebase, with the seeded turbidity ceiling (BP-07) and site-coverage (BP-10) monitors. Captured for the Strata feature manual, 23 June 2026.

The interval setting used to do nothing: MonitorDefinition declared check_interval_seconds=3600, the beat hardwired 300.0, and an hourly monitor ran twelve times an hour. Segment 056 skips each monitor until its own interval has elapsed; shorter intervals round up to 5 minutes, and the form says so.

Each monitor stores a SHA256 fingerprint of the breaching rows and re-fires only when that set changes. The demo turbidity monitor fired at 3 breaching rows and again when the drift document pushed it to 6, and that was two alerts in total with the sweep running twelve times an hour.

The fingerprint was json.dumps(rows, sort_keys=True), which sorts keys inside each row but not the row list, and the breach query has no ORDER BY, so reordered rows gave a new fingerprint. One canonical helper now derives it for every monitor type; the worst case on deploy was each monitor re-firing once. Edits that affect evaluation (type, object, property, operator, threshold, rule) re-arm the fingerprint, so a tightened threshold fires on the next tick; cosmetic edits (name, target, enabled) keep it.

Thresholds compared against a floor in the agreement

Until segment 056 a threshold was one float against one property: WHERE TRY_CAST("prop" AS DOUBLE) > 5.0. That cannot express the flagship compliance rule, "pay rate below the Award floor", because the floor is data in an enterprise agreement PDF's schedule of minimum rates, varying by classification and level.

An analyst models the floor once. The schedule extracts into its own ObjectType, a governed_by LinkType joins compensation rates to it, and a dynamic_threshold rule in the monitor's existing rule field compares hourly_rate against the linked min_hourly_rate. Page 146: the number nobody was going to find walks through that monitor, its config and the compiled SQL, so I won't repeat it here.

The comparison compiles to a correlated EXISTS, since a bare JOIN emits a copy of each local row per matching floor row, inflating the count and destabilising the fingerprint. A row is breach only when both operands cast to a finite number and the comparison holds, no_data when its operand will not cast or no castable floor row exists, otherwise compliant, so a missing floor never passes as compliance. A non-finite bound (NaN or infinity) is rejected when the monitor is created.

Four monitors with status chips: three in breach, one compliant
The four monitor kinds with live operational status read from monitor_state. From the strata/056 browser walkthrough, July 2026.

In the live acceptance run (isolated VM), the award-floor monitor fired exactly one alert with the two below-floor rows, E1 at 20.00 and E3 at 17.50, carrying the Schedule-J provenance of the floor documents. An unknown expression property returned 400, a withheld classification came back no_data(right_missing), and the deadline monitor fired on the row overdue since 2020 and ignored the one due in 2099. Segment 056 also carries 52 unit tests.

Completeness: alerting on data that never arrived

The second shipped type detects absence: a grid of every group_by value against every period in the view, where an empty cell is a gap.

Completeness monitor editor with its inline coverage explainer
The completeness editor, with the inline explainer that spells out what coverage means before you save the rule. Captured for the Strata feature manual, 23 June 2026.

The demo alert reads "coverage 93.8% (12 gap(s) across 4 group(s) × 48 period(s)); cast casualties: 2": 4 sites by 48 months, 12 empty cells. Cast casualties are rows whose measure failed numeric conversion (two OCR-mangled cells, 5..2 and N/A*); they stay in the views as NULLs and get counted here. An optional expected_periods list pins the period axis, which is how the monitor knows 2024 should have had rows for the Western Bore site.

Blast radius as a count the engine computed

The selling point is "a monitor fired, here are the people affected". That number used to come from the BP-13 Impact Advisor agent, prompted with the ontology's cardinality, so any count was the model's own.

Alert card with the impact narrative, three suggested actions and an approved badge
The alert-to-decision card: the breach, the advisor's narrative walked through the LinkTypes, the suggested actions (notify_compliance, schedule_retest, flag_dashboard) and the approved badge written back after a human decided. Captured for the Strata feature manual, 23 June 2026.

The narrative follows the ontology's topology (WaterQualityObservation → OrganisationDimension (observed_by) and → MonitoringSiteDimension (observed_at)) to name the intake operator, downstream sites, the department with the remediation budget and the risk of a reportable event. Only the counts were ungrounded. The audit before segment 055 called it a wiring gap: grounded counting already ran on the Object 360 page, but not on the alert path.

In segment 055 the topology walker gained a depth (clamped at 2), a direction and a per-hop chain of join conditions. The executor counts distinct terminal instance keys rather than physical rows, and total_affected is the union of the primary target type's keys across all selected paths, so an instance reached by two paths counts once. Budgets of 32 blast paths, 128 count queries and 4,000 ms of wall clock cap the work; beyond them, hops degrade to count: null with a reason, so a diamond-shaped ontology cannot run unbounded joins.

The live suite named in the impact-rules FRD ran these checks in an isolated VM against the real ontology:

Live check (strata/055, isolated VM) Result
Distinct logical instances, not physical rows 2 sites, not 3 rows
Grounded count against a hand-run COUNT(DISTINCT key) Equal
Two physical rows sharing one key Counts 1
Composite-key hop (from strata/054) Counts 1 target instance
Declared primary target type total_affected=2 plus a deterministic union digest

The counts commit in the same DuckDB writer transaction as the alert row and the fire-once claim; the notification goes out after commit, and only if this evaluation won the claim. A crash between "insert alert" and "record fingerprint" cannot duplicate an alert, and if a beat tick races "Evaluate now", the loser's compare-and-set updates zero rows.

Operations Cockpit showing an exception queue and a grounded "2 AwardFloor affected" headline with the breaching rows
The read-only Operations Cockpit: three open alerts across one Strata knowledgebase, with the grounded "2 AwardFloor affected (grounded, distinct instances)" headline and the breaching rows E1 and E3 traced back to comp-fy25.pdf. From the strata/061 verification run, alert timestamps 12 July 2026.

The cockpit renders numbers only from the committed impact payload; without one it shows the narration, numbers stripped, under "grounded impact unavailable".

The approval gate

The advisor agent also proposes actions and POSTs its narrative, suggestions and decision back onto the alert. Rejections are recorded with decision: rejected and zero external writes, since compliance teams want those on the record too.

Unattended runs have no stream queue, so the human-in-the-loop node emitted into a null queue and blocked on a future for up to min(timeoutSeconds, 600) seconds before rejecting. The gate only worked for a developer running the agent in the builder.

Segment 059 added an unattended branch. With no stream queue the node snapshots a manifest, moves the durable ledger row from pending to awaiting_approval by compare-and-set, and raises ExecutionPaused, which unwinds the run and returns {"status": "awaiting_approval", "ledger_id": ...}. The ledger is SQLite in WAL mode on a volume shared by the API and the Celery workers, so it survives restarts and every worker sees it.

The manifest freezes the agent and ActionType content hashes, params, variables, effect plan, prompt and digest, completed nodes and switch selections; approval replays exactly that, re-planning and re-fetching nothing. An ActionType with a second downstream gate or a non-serialisable output is refused at registration with a 422, as it could never resume deterministically.

Approvals inbox card for reissuing an award rate, with Approve and Reject buttons
A parked action in the approvals inbox: requested by system:monitor:mon_awardfloor, ledger a record ID, one declared effect. From the strata/059 verification run, 12 July 2026.

Each row records the requester (system:monitor:{monitor_id} for a monitor), the execution owner and, once someone decides, the approver, so the record shows that a monitor asked and which user approved. The same inbox takes the AI FDE's parked actions, covered in An agent with write access to your ontology.

We describe effects as compensatable, because undoing an external write means issuing a second one: a void does not erase an invoice or its consequences, and a stale before-image can overwrite someone's newer change. Each effect has an append-only state log, prepared → dispatching → applied → compensating → compensated, plus outcome_unknown for "sent the write, crashed before recording the result", which a reconcile pass resolves. The Xero proof creates a DRAFT invoice and compensates it to DELETED (Xero only permits VOIDED from AUTHORISED), with the ledger's effect key as the Idempotency-Key.

The loop passes end to end against a recorded stub connector. The real Xero DRAFT-to-DELETED run is skip-marked and environment-gated, and segment 060 stays IMPLEMENTED until it passes.

The bug: alerts from deleted monitors looked live

Alert history is immutable, so a deleted monitor's alerts stay for the audit trail, but the UI never said so: an empty monitors list could sit above three alerts from no visible rule.

Empty monitors list above three alerts badged "monitor deleted"
The fix on screen: no monitors, three retained alerts, each badged "monitor deleted" and with the now-meaningless Snooze control hidden. Captured 16 August 2026.

The panel reads monitor definitions and alerts from two stores (now concurrently, not back to back) and left-outer-joins the alerts onto the monitor list in the client. An alert whose monitor has gone survives the join and gets the badge, and Snooze is hidden because snoozing a deleted monitor does nothing. Deleting now asks first and states the consequence.

Delete-monitor confirmation dialog with Keep monitor and Delete monitor buttons
The confirmation dialog. When the monitor has fired, the same dialog names the alert count and explains that the history is kept and labelled. Captured 16 August 2026.

The walkthrough video shows the empty list beside badged alert history, then a monitor created, deleted through the dialog, and the toast.

Tracing that found a bigger problem. MonitorDefinition was a node in the Apache AGE graph alongside ObjectType, LinkType and SchemaMatch. Those three are relationship-shaped; a monitor is flat CRUD config with no edges. The split had already caused a duplicate-alert bug, fixed in segment 056, and left mute and snooze with no existence check. Monitor definitions moved to a dedicated Postgres table, lazily created like the KB metadata table, with the ontology code untouched. We backfilled and verified all 15 pre-existing monitors across 8 knowledgebases, then removed the stale graph nodes.

Segment status and open gaps

Status and test counts from the v3.0 master tracker in the PRD:

Segment Scope Status
strata/055 Grounded multi-hop impact counts DONE, 52 unit + live
strata/056 Dynamic, deadline and lifecycle monitors DONE, 52 unit + live
strata/058 ActionType registry + execution ledger DONE, 43 unit + 26 live checks
strata/059 Approvals inbox, unattended park DONE, 39 unit + live
strata/060 Compensation and write-back IMPLEMENTED, 21 unit + live stub; real Xero proof deferred
strata/061 Read-only Operations Cockpit IMPLEMENTED, 44 unit (32 jest + 12 pytest)
strata/065 Cockpit action leg DONE, 29 jest; consumer-only over the committed ledger APIs

Rate-of-change monitors are proven at the compile and expression level; forecast-based drift was out of scope. Impact stops at two hops. Completeness alerts carry no instance counts, since a gap is a (group, period) cell. Segment 060's live proof is the stub connector. Not yet measured: what the 5-minute sweep costs at a large monitor count, and the time from a breach landing to the card reaching an inbox.

Build a brain for your business.

Certant turns your documents, data and processes into agents, dashboards and assistants you can actually trust.