Certant Strata and Palantir Foundry / AIP, feature by feature
I graded our ontology product against the nine things a Foundry / AIP engagement delivers. It scored three partials and six reds, and it leads at the document end.
In July I graded Certant Strata against the nine capabilities that make a Foundry / AIP demo land, using our own code and Palantir's public documentation. It scored three partials and six reds, with no full greens. That document is STRATA_TRUTH_AND_FOUNDRY_FEASIBILITY.md in the repo, and this article is mostly it, brought forward to what has shipped since.
The nine capabilities I graded against
The list came from what a Foundry / AIP supply-chain twin demo shows, and the reason in each row comes from our own code.
| Foundry / AIP capability | Strata verdict at the time of writing | Why |
|---|---|---|
| Live digital twin | red | Batch and snapshot only; feeds were cron poll-and-replace, so best case is as-of-last-ingest |
| Data harmony across live systems | red | No CDC, no live connectors, and no cross-system entity resolution |
| Real-time click-to-act write-back | red | Actions were opaque agent graphs, no typed contract, no dry run, no durable ledger |
| Object-level economics and rollups | red | Property has no formula or derived field; views substitute one raw column |
| Chained scenario simulation | red | Repo-wide grep for `simulat |
| Ontology feeding LLM reasoning | partial | Object model is real and query-backed; blast-radius narration was ungrounded prose |
| Master-data correction flywheel | partial | Corrects how data is understood, never the values themselves |
| Modelling the idiosyncratic business | partial | Strong on messy data, no process or workflow model in the ontology |
| Scale | red | Largest proven run was ~320 table shapes over ~1,920 rows |
The document summed it up in one line. Strata is a batch, per-KB, document-derived analytics modeller; Foundry / AIP is a live, multi-system, write-back operational twin; the gaps are the distance between them. The v3.0 programme has since closed four of those gaps.
Where Strata is ahead
Our advantage is at the front of the pipeline, which has six phases: extract, profile, match, synthesise, views, resolve. Tables come out of OCR HTML into per-knowledgebase DuckDB, get deterministic statistics plus one temperature-0 annotation call each, get matched across documents and years by mutual top-k nearest-neighbour clustering over column embeddings, and then a single LLM call turns the active matches into ObjectTypes and LinkTypes as immutable versioned projections. Views are deterministic SQL over the raw tables, and every identifier is a deterministic hash, so re-runs are idempotent and human decisions survive them (Building a data model from PDFs instead of designing one has the detail).
Nobody else in the field builds the model from the documents. I checked every competitor profile first: Databricks Genie needs the structured estate, Snowflake's Semantic Views work exclusively with structured relational tables, Microsoft's Data Agents do not support .pdf, .docx or .txt, and the formal knowledge-graph vendors tag document text into an ontology a human authored first. Palantir's own February 2026 documentation describes entity extraction from documents that would directly populate Ontology objects as being developed, not generally available.
Extraction has run to the cent against a 414-page scanned enterprise agreement holding 257 tables, and a GLM-OCR end-to-end check matched 10 out of 10 cells against ground truth.
Natural language goes to an ontology-grounded execution plan over ontology.* views only, capped at three joins, with joins forced to follow stored LinkType conditions so the model cannot invent a join key. Ask something the model cannot answer, such as an employee churn rate when there is no staff data, and you get a "No matching data model" card with the resolver's reasoning instead of a number.

Those 13 blueprints, BP-01 through BP-13, are the Foundry / AIP playbook as one-click agent templates, with two more under "Strata examples". The wizard collects only what it needs and hands you a wired canvas. The write-back blueprints, BP-11 for Xero and BP-12 for iMIS, stay locked until the matching connector is configured. Their external writes are mocked in tests, and the live Xero proof is still a tracked follow-up.
An alert only re-fires when the SHA256 fingerprint of the triggering row set changes, which mirrors Automate's threshold-crossed semantics. In v3.0 the threshold stopped being a constant float: a monitor can compare a property against another property, a joined column, an aggregate, its own previous value with LAG(), or a date. The flagship case, "pay rate below the Award floor", reads the floor from another table. It compiles to correlated EXISTS SQL, and fired live on the real enterprise-agreement corpus.

For provenance, every entity with a designated instance key gets a URL and a page: properties, related instances traversed through real LinkType joins, source documents, and an activity feed of alerts whose rows match it.

measured_at LinkType to 60 water-quality observations and computes chips over the linked measures. The completeness monitor's FY2024 coverage gap appears on the instance itself. Captured 11 June 2026.The provenance chip on a value opens the source PDF at its page: on the demo corpus, page 146 of a 414-page agreement, landing on Schedule J, the salary table the instance rows were extracted from (Page 146: the number nobody was going to find tells that story). In v3.0 the extractor also started capturing the table-block bounding box the OCR already carried, so eligible answers hover-crop to the region, labelled "source table region". We do not claim exact cell geometry. The interface says which scope it is showing, and any answer that involved a join is marked as not exact.

What v3.0 changed, and what we still do not claim
v3.0 closed four gaps. Grounded blast radius replaced the narrated kind: a declarative impact rule per ObjectType names which LinkType traversals to count, and the engine computes COUNT(DISTINCT instance-key tuple) per hop over one bounded snapshot, so the advisor agent narrates over computed numbers. Actions became a versioned, content-hashed ActionType declaring params, typed effects, its bound agent and an approval policy, with a server-forced dry run and a durable execution ledger. The unattended approval loop closed, so a monitor firing unattended parks the action in an inbox instead of blocking for 600 seconds and falling through to a no-write reject. Write-back is now compensatable (deliberately not "reversible"): a DRAFT invoice can be walked back to DELETED with the ledger effect key as the connector idempotency key, and a partial outcome is reported as partially_compensated.
The v3.0 test records show 120 unit tests plus 22 live checks for composite joins, and 43 unit tests plus 26 live checks for typed actions, all run in an isolated VM against the real queue and the real per-knowledgebase DuckDB.
The v3.0 charter also lists what we will not say, and it binds sales language as well as code. It rules out claims of data-volume scale, cross-system entity resolution, scenario simulation, computed properties and a value-level correction flywheel. The phrases "live digital twin", "entity-resolved" and "billions of rows" are retired until the deferred work activates, along with "every answer exact", "cell highlighting" and "Foundry requires cloud".
Where Foundry / AIP is clearly ahead
Scale is the biggest gap. What we have proven is table cardinality: roughly 320 table shapes over about 1,920 fact rows in the largest run anywhere, with live tests between 60 and 257 tables. Each knowledgebase is one embedded single-file DuckDB behind one write lock, so parallelism exists only across knowledgebases, and the concurrency work fixed correctness by serialising, keeping the single writer. We have never run anything near millions of rows.
Cross-system entity resolution does not exist. Sources sharing a knowledgebase get unified columns, but the views are UNION ALL for facts and SELECT DISTINCT for dimensions, so the same real-world entity arriving from two systems stacks as two rows and is never merged into a golden record.
There is no scenario engine either. Impact traversal is bounded hops with counts, the resolver emits one read-only SELECT capped at three joins, and the LinkType model has nowhere to store an effect rule. We cannot chain price to demand to capacity to procurement and compare the options.
Palantir also has OSDK, and we have no SDK at all. Our structured path is Strata Feeds and a handful of connectors, and the dashboards are functional but young.
Deployment on customer hardware
Strata runs on one box the customer owns, and we have proven the full stack disconnected on it: local Qwen weights served by SGLang, embedded DuckDB, the graph in Apache AGE, answers with citations, and tcpdump showing zero packets out with the external model endpoints blocked. (What air-gapped actually means has the audit.) That configuration has no cloud control plane and no forward-deployed engineer on site. The AI FDE is a drawer in the product that parks anything touching the live model in the approvals inbox as a frozen manifest (An engineer in a drawer shows it at work). It is interactive only and bound to one knowledgebase per session in this version.
How I would choose between them
If the requirement is a live operational twin over many transactional systems, at retail scale, with entity resolution and scenario planning, buy Foundry / AIP. We are not close, and six of the nine rows above say so in code.
If the requirement is an estate of contracts, agreements, compliance schedules and rate tables nobody has been able to compute over, with data that cannot leave the building and no warehouse or ontologist on staff, Strata fits. It builds the model out of the documents, traces every number back to the page it came from, watches it with rules whose thresholds are themselves data, and gates any write behind a human approval. That is a much smaller claim than a digital twin, and it is the one our own gap register supports.


