How the platform is put together — read at whichever altitude you need.
Requirements filed
84
Delivered
15
Alpha (prototype)
12
beside Delivered, not in it
In progress
4
Proposals open
42
of 88 with a header
Next gate
M4Persistent Sensing
Q3 2026
Live at v3.12 (2026-09-03). Every figure comes from the catalogue.
The product document makes specific commitments to procurement: human oversight as a design requirement, a triage cycle under four hours, no vendor lock-in at any layer — and an assistant that brings work forward rather than waiting to be asked. This is the machinery that makes those commitments enforceable rather than aspirational — and the one place it is not finished yet.
Sections are collapsed — the summaries are the overview.
The loop we sell, and the number it turns onThe delay we sell against is not analysis time — it is the hand-offs between systems that do not share a record.
The delay we sell against is not analysis time — it is the hand-offs between systems that do not share a record. Removing the hand-offs is an architectural property, not an optimisation, and the same record is what makes each campaign improve the next.
The assistant does not wait to be askedThe platform does not only answer faster — it notices what nobody asked about.
The platform does not only answer faster — it notices what nobody asked about. This has already happened in practice: a cross-check surfaced an asset carrying more unreviewed evidence and more design-side corroboration than the one under investigation, with no finding ever raised against it. The loop back is the part that keeps it welcome: acceptance rate is how the feature measures whether it is earning attention or spending it.
Why “no vendor lock-in at any layer” holdsHardware- and system-agnosticism is a consequence of the architecture, not a promise on a slide.
Hardware- and system-agnosticism is a consequence of the architecture, not a promise on a slide. Every boundary is defined once, so a new sensor, historian or downstream system is a connector written against a contract — which is also why a pilot configures rather than integrates.
How “human in the loop” stops being a policy“Observe, reason, recommend” and “human oversight is a design requirement” are commitments we have already put in front of procurement.
“Observe, reason, recommend” and “human oversight is a design requirement” are commitments we have already put in front of procurement. This is what turns them into properties of the system rather than rules people are asked to follow — the assistant has no route to facility data of its own.
What the platform is made ofFive capabilities, each with the requirements delivered over the requirements filed.
Five capabilities, each with the requirements delivered over the requirements filed — and, beside that strict count, how many are on alpha and how many are in progress, so a capability with work under way never reads as untouched. This is the CEO altitude; the CTO tab opens each one into its themes, and the Engineering tab opens each theme into its requirements.
Asset identity, spatial registration, tag reconciliation, engineering context, photorealistic 3D scene (3DGS), AI assistant brain (persistent memory and knowledge)
Verification queue, recommendation packets, inspection plan handoff
0 / 10+6 alpha+0 in progress
Commitment by commitmentEach promise the product document makes to a customer, and the thing that keeps it true.
Each promise the product document makes to a customer, and the thing that keeps it true.
What we have committed to
What makes it true
Status
Human oversight as a design requirement; observe, reason, recommend only
The assistant has no path to data of its own. It acts as the signed-in engineer, sees only what they see, and every action is recorded against their name.
Built
Triage to work order in under four hours, against five to ten days today
One loop and one record end to end, so nothing is re-keyed between systems. The delay we sell against is hand-offs, not analysis time.
Built
Hardware-agnostic, system-agnostic, no vendor lock-in at any layer
Every boundary is defined once, so a new sensor, historian or downstream system is a connector written against a contract rather than a rebuild.
Built
A 90-day pilot rather than a multi-quarter data-modelling effort
The same definitions generate the connectors, so a pilot configures what already exists instead of integrating from scratch.
Built
Partner integrity firms hold the review seat
Partners reach the same capabilities under their own identity and their own permissions — no separate, looser access path to build or audit.
Built
AI outputs constrained to defined schemas, with out-of-range values flagged rather than passed on
Definitions are enforced by the build, not by review. This is the machinery the hallucination-mitigation story rests on.
Built for data and events
An assistant that is proactive and curious — bringing work forward rather than waiting to be asked
Ranked by attention debt, filtered by the same confidence gate as a requested answer, delivered into a digest the engineer paces. It surfaces and asks; it never acts. Now four requirements in the catalogue with dates — the digest and the ranking in Q1 2027, coverage prompts in Q2 2027. The fourth, the assistant asking its own question, is filed deliberately undated: it waits on a contract change, not on capacity.
Committed
Specialist detection models called on demand as plug-and-play tools — one of the four “why now” enablers
The tool definition: one description of a capability, projected to every surface at once. Designed, with the first test written to fail.
The gap
What this needs from you
Nothing to approve. One useful input: which route matters commercially next — a partner API, a customer-facing command line, or deeper reach inside the assistants their teams already use.
And one thing to know: the single unfinished piece sits under a capability we already list as a reason the platform is possible now. It is the next thing we close.
What ships, in what order, and what each new evidence stream unlocks?
The product document's 84 requirements are one catalogue, read here the way a product manager reads it: the integrity chain grows one evidence modality at a time, each milestone has a product gate, and every status below is the catalogue's own — none is typed into this page.
Sections are collapsed — the summaries are the overview.
One chain, entered one modality at a timeThe six-stage integrity chain is modality-agnostic at entry; each evidence stream unlocks a different subset of the later stages.
The chain — observed anomaly → damage-mechanism review → engineering decision → risk → recommended action — is one product. What changes over the roadmap is which evidence enters it. The ladder is the order the platform adds them; the status column is the anchoring requirement's own.
Evidence modality
Milestone
What it adds to the chain
Status
Inspection imagery — RGB and thermal campaign captures
M4
Stages 1–4 and 6: attributed intake, candidate screening, evidence-grounded findings, the engineer's decision, actions and reports. Closes with the engineer's priority, not a risk score.
Alpha (prototype)
Field survey readings — gas with ambient temperature, humidity, pressure
M4
The first operating-condition input: the ambient reference for thermal screening, environmental factors for mechanism plausibility, a second independent source per asset. Attached to assets as nearby context, never as attribution.
Q3 Target
OGI imagery
M5
Leak evidence at Stage 4. Verification needs footage no existing campaign carries.
Q4 Target
Process telemetry — SCADA and historians via OPC UA
M5
Turns Stage 2 into true operating-window checks and makes Stage 5 — quantified risk — possible.
Q4 Target
Thickness measurements — via the IDMS read direction
M4
Measured corrosion rates and remaining life at Stage 5.
Q3 Target
Robot patrol evidence
M6
Repeat visits of the same modalities; trend across passes.
Q1 2027 Target
Two boundaries keep this honest. Quantified risk (Stage 5) is not derivable from imagery and stays a Q4 commitment for every modality. And the imagery entry runs as a functional prototype on the alpha lineage, exercised so far on a handful of findings rather than a campaign.
The milestone loopEight milestones, each with a product gate; 4 closed, 4 open, next is Persistent Sensing (Q3 2026).
Milestone
Target
Product gate
Status
M0 · Platform Foundation
Jul 2025
3D Viewer + Authentication + Image Gallery + Operator Dashboard, validated on first real-world RGB inspection dataset
done
M1 · AI Foundation
Dec 2025
Multimodal AI Pipeline + Natural Language Interface + Automated Task Coordination + Machine Vision Engine prototype, validated with thermal imagery and gas sensor readings
done
M2 · App MVP
Mar 2026
Unified 3D Viewer + AI Chat operator interface; Avoided-Cost Pilot case study (MVP v0.2)
done
M3 · AI Q2 Delivery
Jun 2026
MVP v0.3 — Contextual data chat in persona-tailored workspaces (Data Explorer & Integrity Engineer); validated agent reliability (evaluation tests + failure recovery)
done
M4 · Persistent Sensing
Q3 2026
MVP v0.4 — Closed loop on inspection evidence (Alpha) + CAD/P&ID Engineering Context; IDMS Workflow Integration; Persona Sign-off (IE & DE); ROI Calculator
active
M5 · Engineering Context & Enterprise
Q4 2026
MVP v0.5 — Closed loop with process context; Engineering-context depth + enterprise readiness; Quantified Inspection Prioritization (API 581 RBI); 10 Signed Letters of Intent (LOIs); SOC 2 Type II Sign-off; First 3 Production Subscriptions; Customer Cloud Tenant Package (air-gapped: H1 2027)
planned
M6 · Repeat Visit
Q1 2027
A facility’s second and third patrol cycles ingest, rank, and reach an engineer as a digest they did not ask for — with acceptance rate recorded
planned
M7 · Continuous Coverage
Q2 2027
20 h/day coverage sustained at one equipped site, with the coverage plan produced by the platform rather than by a person
planned
Quarter labels are committed work; half labels are directional. A milestone closes at the status vocabulary's own bar — Alpha (prototype) closes a milestone on the alpha lineage, Delivered needs the production trunk — so a gate met on alpha reads as met, with promotion tracked as the next milestone's work.
Proposals as active changes42 open of 88 with a header; decisions waiting first, then accepted, then partially implemented, then implemented — each with the requirements it would move.
Status
Decision requested
Moves
Proposal
At target
proposed
2026-07-19
—
Single AI Engine Manifest
—
proposed
2026-08-14
—
read asset-tag placards from imagery — OCR-once on the existing skill pipeline
—
proposed
2026-08-18
—
one image viewer, opened in place, carrying the gallery you came from
—
proposed
2026-08-20
—
backport the toolchain, regenerate the rest, copy almost nothing
—
proposed
2026-09-02
—
a shared UI package between KAP's web/ and KavApps' frontend
—
proposed
2026-09-02
—
a TDD approach to designing kav-ai and kav-ui
—
proposed
2026-09-02
—
an analysis turn's image reference — two contracts, both written down
—
proposed
2026-09-02
—
what the cross-engine benchmark can learn from the AI team's ADK evals
—
proposed
2026-09-10
FR-APP-03
FR-APP-03 is two requirements — split it, credit the session row, gate the rest on evidence from the engines we run
0 / 1
proposed
2026-09-10
—
Reduce AI Startup, Asset Loading, and Chat History Latency
—
proposed
2026-09-14
—
a CAD twin of every photo — verified vessels as control points, then render the model from the camera
—
proposed
2026-09-15
—
one reconstruction for camera poses, Gaussian splats, and CAD alignment
—
proposed
2026-09-17
—
an asset is a tag — one resolver for the identifier spaces, not an asset object
—
proposed
2026-09-18
—
placards as control points
—
proposed
2026-09-22
—
one film script for the CAD twin, with the pose guard in one place
—
proposed
2026-09-23
—
anchor the walk, not the frames — flight C in four numbers
Committing H1 2027 — M6 and M7 out of the Roadmap Horizon
10 / 10
Drafts, not yet proposed (6):
AI annotation suggestions — analyze an image, accept what is rightFR-AI-10FR-AI-11FR-AI-12FR-AI-13FR-AI-14FR-SEC-05
the backend as a CLI — generate the commands from the spec, and measure what the spec claims
Documenting the AI assistant — the charter, the tools, the prompts, and the behavior that is a test
Documenting the REST API — the design sheet, the spec, and the things a spec cannot say
Alpha → Main — Identifying the Gap and Transferring the FeaturesFR-VIS-02FR-APP-06
One Tool Object — and a Test That the Model Receives What We Declared
88 with a header of 114 proposals; 42 open. Not listed: 0 superseded, 0 withdrawn, 21 deferred. At target counts the requirements a proposal names whose catalogue status already equals the one it asks for; open means proposed or accepted.
Where every requirement stands84 requirements by delivery state — 31 delivered, on alpha or in progress; 42 targeted this year.
State
Requirements
Share
Delivered
15
18%
Alpha (prototype)
12
14%
In progress
4
5%
Q3 2026
10
12%
Q4 2026
32
38%
2027 quarters
9
11%
Horizon
2
2%
The same buckets the status video narrates. Alpha is counted beside Delivered and never folded into it; the CTO and Engineering tabs open each capability into the requirements behind these numbers.
What the product does, in the user's words59 user stories across 7 milestones; 27 of 27 delivered and alpha requirements have one. Every one of them has a story.
Every story names the requirement it is written for, in the persona's words, with Given / When / Then criteria a test or a certification row can check. Stories for the delivered foundations (M0–M2) were written after delivery, on 2026-09-10, and say so; a story is written only where the behaviour can be named — a requirement without one is a requirement whose behaviour nobody has yet put into words, which is the gap this line reports.
M0 · Platform Foundation5 storiesUS-M0-01Ingest a campaign's imagery with its provenance intactData ExplorerFR-EVI-01Delivered
As a Data Explorer, I want to load a campaign's RGB inspection imagery into one dataset with each image's capture time and GPS position preserved, so that everything downstream — the scene, the gallery, an engineer's finding — can say where and when a picture was taken.
Given a set of drone images with EXIF headers, when I ingest them into a dataset, then each image record carries its capture timestamp and GPS position from the header, and the dataset reports how many images it holds.
Given an image with no GPS in its header, when it is ingested, then it is kept in the dataset and marked as unpositioned — never dropped silently, never given a fabricated position.
Given two datasets in one organisation, when I browse either, then I see only that dataset's images; an image belongs to exactly one dataset.
US-M0-02See the campaign in spaceIntegrity EngineerFR-APP-07Delivered
As an Integrity Engineer, I want a browser-based 3D geospatial scene of the site with the campaign's imagery placed where it was captured, so that I can orient a picture to the plant without leaving the application.
Given a dataset with positioned images, when I open its 3D view, then the scene renders in the browser with each positioned image marked at its GPS location.
Given a marked image in the scene, when I select it, then I see the image and its metadata in place, and can continue to the gallery or the viewer from there.
Given a dataset whose images carry no position, when I open the 3D view, then the scene loads with a clear statement that nothing could be placed, rather than an empty globe with no explanation.
US-M0-03Know where every campaign standsData ExplorerFR-APP-08Delivered
As a Data Explorer, I want one dashboard listing my organisations, their campaigns and the anomalies found so far, so that I can see at a glance which campaign needs attention without opening each one.
Given membership of one or more organisations, when I open the dashboard, then I see every campaign I can access with its image count and anomaly count, and nothing from organisations I am not a member of.
Given a campaign on the dashboard, when I select it, then I land in that campaign's workspace with the same counts.
Given an organisation with no campaigns yet, when I open the dashboard, then it is listed as empty rather than omitted.
US-M0-04Review what the analysis found, picture by pictureOn-call Integrity EngineerFR-APP-09Delivered
As an On-call Integrity Engineer, I want a gallery of a campaign's images with the detected anomalies drawn as bounding boxes and their annotations alongside, so that I can review findings against the evidence rather than against a list.
Given a campaign with analysed images, when I open its gallery, then each image shows its anomaly bounding boxes with the annotation's category and confidence.
Given the gallery, when I filter by anomaly category, then only images carrying that category remain and the count updates to match.
Given an image with no detections, when I open it, then it is shown clean with an explicit "no anomalies detected" state — an empty overlay is not an error.
US-M0-05Sign in once; see only my organisationsOperations SupervisorFR-SEC-04Delivered
As an Operations Supervisor, I want every user to sign in to one account whose organisation memberships and roles decide what they can see and change, so that a contractor on one site cannot reach another site's data by accident or by crafting a request.
Given a signed-in user, when any page or API route loads campaign data, then only rows from organisations that user belongs to are returned — enforced in the database by row-level security, not only in the interface.
Given a user with a viewer role in an organisation, when they attempt a write (an annotation, a finding, an export), then it is refused with a permission error and nothing changes.
Given an expired or tampered session token, when a request is made, then it is rejected and the user is sent to sign in again — never served another user's data.
M1 · AI Foundation4 storiesUS-M1-01Bring thermal, OGI and gas evidence into the same campaignData ExplorerFR-EVI-02Delivered
As a Data Explorer, I want thermal images, OGI clips and gas readings ingested into the same campaign as the RGB imagery, each carrying its modality, so that an engineer can see a location's evidence across sensors instead of across folders.
Given thermal images alongside RGB in a campaign, when they are ingested, then each carries its modality and — where the capture pairs them — a link to its RGB counterpart, so the pair can be viewed together.
Given gas readings for a campaign, when they are ingested, then each reading keeps its concentration, unit, timestamp and position, and can be listed for the campaign.
Given a modality the pipeline does not recognise, when ingestion runs, then the files are reported as unhandled with their names — not ingested as RGB, not lost.
US-M1-02Ask the data a question in plain languageIntegrity EngineerFR-AI-05Delivered
As an Integrity Engineer, I want to ask a question about my inspection data in plain language and get an answer drawn from it, so that finding "how many anomalies in this campaign" does not require knowing where that number lives.
Given a question about a campaign's images or anomalies, when I ask it, then the assistant answers from that campaign's data and shows what it looked at.
Given a question the assistant cannot ground in data, when it answers, then it says so rather than producing a number with nothing behind it.
Given the prototype's known limits, when an answer is wrong, then I can see the query it ran — the criterion this prototype met was inspectable, not reliable; reliability is the bar of M3's single-turn query and its evaluation gate.
US-M1-03Let the assistant plan the work, not just answerIntegrity EngineerFR-APP-10Delivered
As an Integrity Engineer, I want a question that needs several steps — find the dataset, pull its images, run an analysis, write it up — carried out by the assistant as a sequence, so that I ask once and get a result rather than driving each tool by hand.
Given a request that needs more than one tool, when the assistant handles it, then a planner chooses the steps and an executor runs them in order, and the answer reflects every step's output.
Given a step that fails, when the sequence runs, then the assistant reports which step failed and what it had so far, rather than presenting a partial result as complete.
Given a request outside the tools' reach, when the planner considers it, then it says what it cannot do instead of routing to the nearest tool anyway.
US-M1-04See gas readings on the sceneOn-call Integrity EngineerFR-APP-12Delivered
As an On-call Integrity Engineer, I want a campaign's gas readings drawn on the 3D scene as a colour-mapped heatmap, so that a high reading is a place on the plant, not a row in a table.
Given a campaign with positioned gas readings, when I enable the gas overlay, then a heatmap renders on the scene with concentration mapped to colour and a legend stating the scale and unit.
Given the overlay, when I select a hot spot, then I see the readings behind it with their values, units and times.
Given a campaign with no gas readings, when I enable the overlay, then the scene says there is nothing to draw rather than rendering an empty layer.
M2 · App MVP3 storiesUS-M2-01Ask for datasets and get a list, whichever engine answersData ExplorerFR-AI-06Delivered
As a Data Explorer, I want an answer about which datasets exist to arrive as dataset cards I can open, so that the assistant's first answer is a starting point and not a paragraph to read names out of.
Given a question that asks for a collection of datasets, when any certified engine answers, then the stream carries exactly one structured dataset list on a clean run — started, the list, finished — and the interface renders it as cards.
Given a dataset card, when I select it, then I land in that dataset's workspace.
Given an engine that answers such a question in prose only, when it is certified, then it fails the row — the product commits to the surface, not to any one engine's habit.
US-M2-02Ask for images and get a gallery I can click throughIntegrity EngineerFR-AI-07Delivered
As an Integrity Engineer, I want an answer that shows images to arrive as a gallery whose every image opens in the viewer, so that "show me ten images from the unit audit" ends with me looking at the tenth image, not at a list of file names.
Given an image-browsing question, when any certified engine answers, then the answer carries each image's identity — id, dataset slug, filename, type — and the interface resolves thumbnails and opens the viewer from it.
Given the identities in the answer, when the interface renders them, then every image has its dataset slug and no answer carries a signed storage URL — the presentation is the interface's, the identity is the engine's.
Given an image the user's organisation cannot see, when an answer would include it, then it is absent — the gallery is filtered by the same row-level security as every other read.
US-M2-03Hand the recommendations to the people who do the workOperations SupervisorFR-OPS-01Alpha (prototype)
As an Operations Supervisor, I want a campaign's work orders and recommendations exported as a markdown packet and a CSV, so that the maintenance planner can take them into their own system without re-typing.
Given a campaign with recommendations, when I export, then I receive a markdown packet listing each recommendation with its asset, priority and evidence reference, and a CSV with one row per recommendation and the same fields.
Given the CSV, when it is opened in a spreadsheet, then every column is labelled and every row resolves to a recommendation in the platform by its id.
Given a campaign with nothing to recommend, when I export, then I get an empty packet that says so — not an error, not a packet from another campaign.
M3 · AI Q2 Delivery5 storiesUS-M3-01Workspace-scoped natural-language queryIntegrity EngineerFR-APP-02DeliveredFR-APP-03Q3 Target
As an Integrity Engineer, I want to ask natural-language questions scoped to my workspace and campaign so that I can find relevant findings without manual filtering.
Given a selected workspace and campaign, when I ask a single-turn question, then the answer draws only on that workspace's data. (Multi-turn context retention is FR-APP-03, targeted for M4 — accuracy at 2+ turns is currently below release bar.)
Given a query that implies a high-consequence action, when the assistant proposes it, then it requires explicit human-in-the-loop confirmation before any action is taken.
Given no grounded data supports an answer, when the assistant responds, then it states it cannot find supporting data rather than fabricating a result.
US-M3-08Graceful degradation on backend faultsIntegrity EngineerFR-APP-13Delivered
As an Integrity Engineer, I want the chat assistant to fail safely when a backend dependency errors so that I get a clear message instead of a broken or hallucinated answer.
Given a Supabase RLS denial or query timeout, when it occurs, then the assistant returns an empty/no-access result rather than crashing or leaking another workspace's data.
Given a Gemini rate-limit or timeout, when it occurs, then the request retries/backs off and the user sees a clear "try again" message rather than a stalled UI.
Given an expired or tampered JWT, when a request is made, then it is rejected with a re-authentication prompt, not a silent failure.
US-M3-09Trustworthy chat backed by accuracy gatesIntegrity EngineerFR-APP-14Delivered
As an Integrity Engineer, I want the chat assistant's classification/query/report accuracy validated against a benchmark before each release so that I can trust its answers in day-to-day use.
Given a new build of the classifier, executor, or reporter agent, when it is proposed for release, then it must pass the accuracy benchmark thresholds before shipping.
Given a benchmark regression (e.g. below-target accuracy on a query category), when detected, then the release is blocked until resolved.
Given multi-turn conversation accuracy remains below target, when a user starts a new conversation, then the assistant is scoped to single-turn queries only for M3 rather than silently degrading on follow-ups.
As a Data Explorer, I want the chat workspace tailored to my role so that I see the framing and results relevant to exploration rather than engineering sign-off tasks.
Given a Data Explorer persona, when I open a workspace, then the chat surface presents exploration-oriented framing (discovery, browsing) distinct from the Integrity Engineer's action-oriented framing.
Given an Integrity Engineer persona, when the same underlying data is queried, then results are framed toward verification/action rather than raw exploration.
Given multiple workspaces I have access to, when I switch between them, then the persona-tailored framing is preserved per workspace.
US-M3-11Query firewall against injection and malicious inputIntegrity EngineerFR-APP-16Delivered
As an Integrity Engineer, I want the chat assistant's generated SQL to be firewalled against injection and malicious input so that a crafted question can never modify data or escape my workspace's scope.
Given a generated query, when it contains DML/DDL (INSERT/UPDATE/DELETE/DROP/ALTER) or an injection pattern (e.g. SLEEP, CHR/ASCII obfuscation), then it is rejected before execution.
Given a query attempting to reference another workspace's data, when it is generated, then the workspace scope is enforced and the cross-workspace reference is blocked.
Given a rejected query, when the assistant responds, then the user sees a safe refusal rather than a raw database error or partial result.
As an Integrity Engineer, I want OGI, calibrated thermal, and gas readings ingested and rendered natively so that I can review anomalies across modalities in one place.
Given a campaign with OGI/thermal/gas data, when it is ingested, then each modality is parsed, associated to its asset, and viewable without external tools.
Given a modality file is malformed or unsupported, when ingestion runs, then it is rejected with a clear reason and does not block the other modalities.
US-M4-03Geo-tagged assets & imagery in 3DData ExplorerFR-VIS-02Q3 Target
As a Data Explorer, I want geo-tagged assets and imagery registered in the 3D scene so that I can see findings in physical context.
Given geo-tagged captures, when I open the unit, then assets and images register to the same coordinate frame.
Given a large scene, when I navigate, then the viewer streams tiles and stays interactive (no full-model load stall).
US-M4-08Click-through from 3D element to engineering identityIntegrity EngineerFR-CAD-01In Progress (Q2)FR-CAD-07In Progress (Q2)
As an Integrity Engineer, I want to select a 3D element and see its standard engineering identity so that I can move from a visual anomaly to its asset record without manual cross-referencing.
Given an IFC4 model for the unit, when I select an equipment item, then its IFC class, tag, material/lining, and source are shown.
Given a legacy CAD tag and an operator/DEXPI tag for the same asset, when I open either, then both resolve to the same asset record via the dual-tagging cross-reference.
Given a tag the dual-tagging rules cannot resolve, when resolution fails, then the item is flagged for manual mapping rather than silently mismatched.
US-M4-09P&ID structure from a clicked assetIntegrity EngineerFR-CAD-06In Progress (Q2)
As an Integrity Engineer, I want logical P&ID structure ingested via DEXPI so that a clicked asset shows its nozzles, connected lines, and connections.
Given a DEXPI P&ID for the unit, when I view an equipment item, then its nozzles (with service), connected pipe runs, and source-to-target connections are listed.
Given a non-compliant CAD export, when the DEXPI file is ingested, then it is sanitized and parsed rather than failing outright.
US-M4-13Focus the chat on an assetIntegrity EngineerFR-APP-17Q3 TargetFR-CAD-07In Progress (Q2)FR-VIS-02Q3 Target
As an Integrity Engineer, I want to set an asset — by its engineering tag, e.g. AB-106 — as my chat focus so that follow-up questions resolve against that asset without restating context each time.
Given a tag in either its legacy CAD or operator/DEXPI form, when I set it as focus, then the scope shows the one resolved asset record (tag, class, material) via the dual-tagging cross-reference.
Given a tag that does not resolve, when I set focus, then I get an explicit "unknown asset" response with nearest candidate tags — never a silent empty scope.
Given an active asset focus, when I ask a question that names no asset, then the answer is scoped to the focused asset and states that scope.
US-M4-14Anomalies on an asset and its vicinityIntegrity EngineerFR-APP-17Q3 Target
As an Integrity Engineer, I want to ask for all anomalies detected on the focused asset or within a stated distance around it so that I can review everything found at that location across modalities in one answer.
Given a focused asset with associated findings, when I ask "what anomalies were detected on this asset", then annotations linked to it (by tag or spatial association) are returned with evidence references.
Given a stated radius (e.g. "within 5 m"), when I ask about the area around the asset, then findings within that distance of the asset's coordinates are included, each labeled with its distance.
Given no findings exist for the asset or radius, when I ask, then the answer states that none were detected rather than fabricating results.
As an Integrity Engineer, I want to export the active workspace selection and defect findings to PDF/Word so that I can share a defensible record without re-keying.
Given a set of selected findings, when I export, then the document includes asset IDs, evidence references, severity, and recommended actions.
Given an export is generated, when I open it, then content matches what is shown on screen (no missing or placeholder fields).
As an Integrity Engineer, I want a defined bidirectional IDMS integration so that an approved finding becomes planned work without re-keying — under human sign-off.
Given an engineer-approved finding, when I hand it off, then a work item is created in the target IDMS with asset, evidence, and recommended action.
Given Kav AI's corrosion rate disagrees with the IDMS record, when synced, then both values are surfaced for the engineer — the existing record is not silently overwritten.
Given a Critical finding or a Remaining-Life change, when handed off, then it requires engineer sign-off (canonical HITL rule).
As an Operations Supervisor, I want a single procurement vehicle covering platform plus partner integrity-engineering services so that I can buy the closed loop without stitching two contracts — while keeping the option to contract separately.
Given a partner-integrated engagement, when scoped, then the partner provides the HITL Integrity Engineer seat and the proposal is a single vehicle.
Given a customer that prefers separate contracts, when requested, then the platform subscription and partner services can still be split.
US-M4-15Raise a finding as a claim about an assetIntegrity EngineerFR-DMR-01Alpha (prototype)
As an Integrity Engineer, I want a finding to be a claim that a named asset exhibits a condition, grounded in one condition class or one API 571 mechanism and citing the images that show it, so that what I decide on is an engineering statement and not a tagged photo.
Given an asset and a set of images that depict it, when I raise a finding, then it records exactly one grounding (condition class or mechanism), the cited images, and starts as "possible".
Given a finding with no grounding or two groundings, when it is written, then the platform rejects it rather than storing an ambiguous claim.
US-M4-16Screen an asset for credible mechanismsIntegrity EngineerFR-DMR-02Alpha (prototype)
As an Integrity Engineer, I want the platform to propose the API 571 mechanisms credible for an asset's material, service and unit so that my review starts from the catalogue rather than from memory.
Given an asset profile, when screening runs, then each credible mechanism is proposed as a Tier-2 "possible" finding with its screening rationale.
Given a screening proposal, when I read it, then it is labelled as screening, never as evidence, and asks me to confirm or dismiss.
US-M4-17Correct a finding and curate its evidenceIntegrity EngineerFR-DMR-03Alpha (prototype)
As an Integrity Engineer, I want to re-ground a finding, replace its notes and add or remove the images it cites, so that a finding improves as evidence improves without losing its history.
Given an existing finding, when I change its grounding, notes or cited images, then the current row reflects the change and an edit-log row records what changed, by whom, and when.
Given an edited finding, when I open its history, then every edit is listed in order and none has been overwritten.
US-M4-18Decide a finding, with reasoning, in an append-only logIntegrity EngineerFR-OPS-02Alpha (prototype)
As an Integrity Engineer, I want to confirm, dismiss or reinstate a finding with my reasoning recorded verbatim against my name, so that a decision is an auditable judgement and a dismissal can be reversed without erasing it.
Given a "possible" finding, when I confirm or dismiss it with reasoning, then a decision row is appended naming me, the verdict, the reasoning and the time, and the finding's credibility reflects the latest decision.
Given a dismissed finding, when I reinstate it, then the dismissal is superseded (both rows remain) and the finding returns to "possible", not to confirmed.
US-M4-19Work one campaign queueIntegrity EngineerFR-OPS-03Alpha (prototype)
As an Integrity Engineer, I want one worklist per campaign over its findings and notes, ordered by what changed most recently, so that I can answer "what have I not looked at" without leaving the queue.
Given a campaign with findings and notes, when I open its worklist, then both appear in one feed ordered by last change, each showing its kind from the organisation's note vocabulary.
Given a note that describes a condition on an asset, when I promote it, then a finding is created from it and the note records the promotion; the reverse is not offered.
US-M4-20Record an action that names its findingIntegrity EngineerFR-OPS-04Alpha (prototype)
As an Integrity Engineer, I want a work order or observation to carry the finding it follows from, so that the reason for the work travels with it into the export and the hand-off.
Given a decided finding, when I record an action from it, then the action stores the finding reference, a status of open, and my summary.
Given an open action, when work is scheduled or completed, then its status moves open → scheduled → complete and the change is attributed.
US-M4-21Hand a campaign over, or send it backData ExplorerFR-OPS-05Alpha (prototype)
As a Data Explorer, I want to send a campaign for review when the evidence is ready, and as an Integrity Engineer I want to return it with a reason when it is not, so that the relay between the two workspaces is a state we can both see.
Given a campaign in indexing, when the Data Explorer sends it for review, then its lifecycle becomes ready for review and it appears under "awaiting an engineer" in the Integrity workspace.
Given a campaign awaiting review, when the Integrity Engineer returns it, then a reason is required, the lifecycle returns to indexing, and the campaign history shows who moved it, when, and why.
US-M4-22Attribute an image to an asset by the tag it depictsData ExplorerFR-CAD-09Alpha (prototype)
As a Data Explorer, I want an image attributed to an asset only when the equipment tag in the frame has been read and I have confirmed it, so that a photo of a cable beside a vessel never becomes evidence about the vessel.
Given an image whose OCR pass read an equipment tag, when I open the asset, then the image is offered as a suggested attribution that I confirm or reject, and my decision propagates to its RGB/thermal pair.
Given images near an asset that carry no tag, when I open the asset, then they are listed as nearby (or framed, where camera pose is known) context and are never cited as evidence unless an engineer curates them onto a finding.
US-M4-23Drive the loop through an AI agent, as myselfIntegrity EngineerFR-APP-22Alpha (prototype)
As an Integrity Engineer, I want an AI agent connected through the MCP surface to raise, correct, decide and curate findings and record actions as me — under my permissions and with my name on every write — so that the assistant can do the legwork while the judgement stays mine.
Given an agent acting under my session, when it reads or writes, then row-level security applies as it would in the app, and every write is attributed to me.
Given an agent asked to decide a finding, when I have not stated my conclusion, then it reports what it sees and asks, rather than deciding on my behalf; a decision write requires explicit confirmation.
US-M4-24Carry on a conversation, and keep conversations apartIntegrity EngineerFR-APP-03Q3 Target
As an Integrity Engineer, I want the assistant to remember what we were just talking about — and to offer me prompts about what I am looking at — so that a follow-up question is a short sentence, not a restatement of the whole context, and so that two conversations I have open never bleed into each other.
Given I asked about a campaign's images, when I follow up with "only the thermal ones", then the answer narrows the previous one — the images returned are a subset of the first answer, from the same campaign, without my naming it again.
Given two conversations open on different threads, when I continue one of them, then nothing from the other appears in the answer, and every event in the stream carries the thread I asked on — an engine never answers on a thread of its own choosing, and a blank thread id never lands me in someone else's session.
Given a selected campaign, module, candidate or asset, when I open the assistant, then the prompts it offers are about that selection — the campaign by name, the asset by tag — and change when the selection does.
Given a conversation several turns long, when I ask "which dataset did I start with", then the answer names the first dataset I asked about, not the most recent.
As an Integrity Engineer, I want read-only SCADA / OPC UA signals aligned to assets so that IOW exceedances can be reasoned about alongside physical evidence — without Kav AI ever writing to the control system.
Given an OPC UA endpoint and a tag map, when the connector runs, then it reads values read-only and associates each to its asset in the world model.
Given an IOW threshold is exceeded, when the value is ingested, then it is surfaced as a time-series signal on the asset, not as an autonomous action.
Given the connector loses the endpoint, when a read fails, then it degrades to last-known with a staleness flag rather than blocking other evidence.
US-M5-023D CAD model overlayData ExplorerFR-VIS-01Q4 Target
As a Data Explorer, I want the engineering CAD model overlaid on the photorealistic scene so that I can read findings against as-designed geometry.
Given an IFC4 model and the captured scene, when I open the unit, then the CAD overlay registers to the same coordinate frame as the assets and imagery.
Given the overlay is on, when I inspect an asset, then its design geometry and field evidence are visible together.
As a Design Engineer, I want CAD versions tracked and compared to as-built so that engineering changes and field deviations are visible, across more than one CAD format.
Given two CAD versions, when I diff them, then added/removed/changed elements are listed and highlighted in 3D.
Given an as-built capture vs the as-designed model, when compared, then deviations beyond tolerance are flagged on the affected assets.
Given an engineering change, when a new revision lands, then affected assets are notified for re-review.
Given an RVT or DGN source, when ingested, then it resolves to the same asset schema as the IFC4 path (no format-specific dead end).
As an Integrity Engineer, I want a direct read from the operator's P&ID SQL database so that logical process structure is available even where a DEXPI export is not.
Given P&ID database credentials, when the connector reads, then equipment, lines, and connectivity resolve to the same world-model assets as the DEXPI path.
Given a tag present in SQL but not in CAD, when reconciled, then it is surfaced for dual-tagging rather than dropped.
US-M5-14Chat grounded on the 3D mapData ExplorerFR-APP-04Q4 TargetFR-APP-05Q3 Target
As a Data Explorer, I want chat answers to highlight the relevant assets and overlays in the 3D model so that I can see where a finding is, not just read about it. (Pushed from M4: viewport grounding builds on the M4 asset-focus work and the CAD/P&ID-anchored world model.)
Given an answer that references specific assets, when it is returned, then those assets are selectable/highlighted in the 3D view.
Given an interactive overlay (thermal, gas, OGI) is available for the asset, when I open the result, then the relevant overlay can be toggled on in context.
US-M5-15Analyze an image on demandIntegrity EngineerFR-AI-10Q4 Target
As an Integrity Engineer, I want to ask for an image to be analyzed from the viewer so that I get findings when I want them, without composing a chat message and hoping the assistant chooses the right tool.
Given an image I am permitted to see, when I request analysis, then the findings returned correspond to that image and no other, and each is labelled with a location on the image.
Given an image the analysis cannot read, when the attempt completes, then I am told it failed and why, rather than shown an empty result that looks like a clean surface.
Given I lack permission for an image, when I request analysis, then the request is refused, and refusal is indistinguishable from the image not existing.
US-M5-16Review what the model proposedIntegrity EngineerFR-AI-11Q4 Target
As an Integrity Engineer, I want proposed annotations shown as suggestions I can accept, reject, or recategorize so that model output never enters the record as though a person had drawn it.
Given an analyzed image, when I open it, then suggestions are visually distinct from stored annotations and are not counted as annotations anywhere in the product.
Given a suggestion I agree with, when I accept it, then it becomes an annotation attributed to me as the accepting reviewer, and the suggestion records that it was accepted.
Given a suggestion with the wrong label, when I change its category and accept it, then the stored annotation carries my category and the record keeps what the model originally proposed.
Given I accept an annotation, when the integrity workflow is consulted, then nothing has been confirmed as an operational finding — accepting an annotation is quality assurance, not integrity confirmation.
US-M5-17Analyze a collection without waiting on the pageIntegrity EngineerFR-AI-12Q4 Target
As an Integrity Engineer, I want to analyze a set of images as one run I can leave and come back to so that a large campaign is a background job rather than an afternoon spent holding a browser tab open.
Given a selection of images, when I ask to analyze them, then I am shown how many are eligible, how many are already analyzed and will be skipped, and how many previously failed and will be retried — before anything runs or is billed.
Given a run in progress, when I reload the page or open it elsewhere, then progress is accurate, because it is read from the run rather than from my connection.
Given a run in progress, when I cancel it, then queued work stops and I am told what had already completed.
Given a run where some images could not be read, when it finishes, then it reports how many succeeded and how many failed with reasons, and is not presented as a plain success.
Given images already analyzed by the same skill version, when I start a run over them, then they are skipped by default and re-analyzed only if I explicitly ask for it.
US-M5-18Know what produced a finding, and who decided about itIntegrity EngineerFR-AI-13Q4 Target
As an Integrity Engineer, I want every finding to carry what produced it and every decision to be kept so that a result can be reproduced, questioned, and audited a year later.
Given any finding, when I inspect it, then I can see which backend, model version, and skill version produced it, and when.
Given a reviewer who rejects a suggestion and later accepts it, when the history is read, then both decisions are present with their actors and times — the later one does not overwrite the earlier.
Given two backends configured for the same skill, when their findings are compared, then each is attributable to the backend that produced it rather than to the skill alone.
As an On-call Integrity Engineer, I want the cross-source correlation engine to report multi-source confirmed performance so that the closed loop has a measured accuracy bar.
Given the multi-source confirmed set, when measured, then TPR / FPR are reported against the calibration curve (target TPR > 98% / FPR < 2%).
Given a correlated finding, when surfaced, then the contributing sources and match basis are shown so the confirmation is auditable.
US-M5-06Physical-AI reasoning to remediationIntegrity EngineerFR-ANO-02Q4 Target (research-gated)
As an Integrity Engineer, I want a finding reasoned from anomaly → damage mechanism → remediation so that I get a defensible recommendation, with a safe fallback if the capability is not yet validated.
Given a confirmed finding, when reasoning runs, then the suggested mechanism and remediation cite API 571/581/584 grounding.
Given the research bar (> 90% agreement with the IE panel) is not met, when the feature ships, then it falls back to descriptive reporting only — no remediation advice.
Given any suggestion, when surfaced, then it passes the Filter Skill and confidence gate first (priority order preserved).
US-M5-07Synthetic data for rare defectsData ExplorerFR-MDA-02Q4 Target (research-gated)
As a Data Explorer, I want synthetic data generation for rare defect classes so that detection models improve where real examples are scarce.
Given a rare defect class, when synthetic examples are generated, then they are labelled synthetic and never mixed untracked into evaluation sets.
Given the spike does not meet its go criterion, when assessed, then the fallback (no synthetic augmentation) is recorded and detection proceeds on real data only.
US-M5-10Calibrated, gated AI outputsOn-call Integrity EngineerFR-AI-02Q4 TargetFR-AI-03Q4 TargetFR-AI-04Q4 Target
As an On-call Integrity Engineer, I want confidence calibration, a chain-level consistency gate, and an out-of-distribution (OOD) detector so that I can trust what reaches my dashboard.
Given a stated confidence, when compared to empirical accuracy per bucket, then calibration error is within target and reported per campaign.
Given a chain of inferences, when Stage 3.5 runs, then internally inconsistent conclusions are withheld.
Given an out-of-distribution input, when detected, then the output is marked "UNCERTAIN — REVIEW REQUIRED" and the OOD detector's update cadence is honoured.
US-M5-19A new detector must earn its placeIntegrity EngineerFR-AI-14Q4 Target
As an Integrity Engineer, I want a new model or skill revision measured before it becomes the default so that "newer" is never mistaken for "better" in a product whose output is inspection evidence.
Given a candidate backend or skill revision, when adoption is proposed, then it has been measured on a frozen, independently annotated evaluation set using the agreed protocol, and the numbers are recorded against that revision.
Given a candidate that is worse than the incumbent on the agreed metric, when adoption is proposed, then it does not become the production default regardless of cost or speed advantages.
Given review history from accepted and rejected suggestions, when quality is assessed, then that history is used as training evidence and not as the evaluation set — it only observes what the current model proposed, so it cannot show what the model missed.
As an On-call Integrity Engineer, I want findings correlated across modalities (tag / match / score / surface) so that multi-source agreement raises confidence and contradictions are flagged rather than averaged away.
Given anomalies from ≥ 2 sources on one asset, when correlated, then they are merged into a single finding with a calibrated confidence.
Given sources that disagree, when correlated, then the contradiction is surfaced for review, not silently averaged.
Given the cross-modal anomaly spike does not meet its bar, when it ships, then it falls back to manual triage rather than autonomous flagging.
As an On-call Integrity Engineer, I want AI-suggested damage mechanisms checked against a deterministic susceptibility filter so that implausible suggestions never reach my action dashboard.
Given a suggested mechanism inconsistent with the asset's material/process, when it is generated, then the Filter Skill rejects it before surfacing.
Given an output below the confidence threshold, when it is produced, then it is marked "UNCERTAIN — REVIEW REQUIRED" and withheld from the Critical Action dashboard.
Given cross-source corroboration exists, when it is applied, then it can raise priority but cannot override a Filter Skill rejection (priority order: Filter Skill > consistency gate > cross-source uplift).
US-M5-13RBI inspection intervals and scopeIntegrity EngineerFR-RBI-01Q4 TargetFR-RBI-02Q4 TargetFR-MDA-01Q4 Target
As an Integrity Engineer, I want API 581 risk and inspection-interval calculation within a clear automation boundary, benchmarked against industry RAM data, so that recommended intervals are defensible.
Given asset, damage-mechanism, and consequence inputs, when computed, then the inspection interval follows API 581 with the inputs shown.
Given an equipment class outside the automation boundary, when requested, then it is clearly marked out-of-scope rather than silently computed.
Given Solomon (or sector-equivalent) RAM data, when benchmarked, then results are expressed relative to the peer profile.
As an Integrity Engineer, I want a certified SAP PM connector so that approved findings become planned maintenance in the system of record without re-keying.
Given an engineer-approved finding, when I hand it off, then an SAP PM work order is created with asset, evidence, and recommended action.
Given the connector is certified, when it writes, then it conforms to the operator's SAP PM integration requirements and the canonical HITL sign-off rule.
As an Operations Supervisor, I want analysis output and anything derived from our imagery to stay inside my tenant so that using the product does not quietly contribute our site's data to someone else's model.
Given suggestions and review decisions generated from our images, when any tenant boundary is applied, then they are scoped to our tenant like the images themselves.
Given a pilot that ends without continuation, when deletion is requested, then suggestions, decisions, and derived datasets or exports are covered by the same deletion workflow as the imagery.
Given a proposal to train a shared model, when our data would be included, then it is excluded unless we have authorized it in writing, and any derived artefact remains traceable to its source tenant so the obligation can be assessed.
As an Operations Supervisor, I want SOC 2 Type II, compliance management, and a customer cloud-tenant / air-gapped deployment option so that the platform clears procurement and OT security review.
Given a procurement review, when SOC 2 Type II evidence is requested, then the report and compliance package are available.
Given a high-security site, when deployed, then a customer-tenant or air-gapped profile runs the same capabilities within the operator's perimeter.
Given a high-hazard site, when the dose-aware workflow is scoped (H1 2027), then it is labelled roadmap, not a v1 commitment.
As an Operations Supervisor, I want robot/KRSI captures ingested and orchestrated for coverage so that the facility is inspected on a schedule without manual flight planning.
Given a KRSI capture set, when it is ingested, then each frame is normalized, provenance-tagged (carrier + modality), and spatially anchored.
Given a coverage plan, when patrols run, then covered vs missed assets are reported so gaps are visible.
Given repeated patrols, when fleet analytics run, then coverage and capture quality trends are available per unit.
As an Operations Supervisor, I want fixed navigation beacons and a communication backbone so that autonomous capture is reliable where onboard SLAM alone is not.
Given installed beacons, when a robot navigates, then localization meets the registration-accuracy target for as-built comparison.
Given a backbone outage, when connectivity drops, then captures buffer locally and sync on reconnect rather than being lost.
As a Data Explorer, I want full spatial navigation through the unit so that I can move through confined and hard-to-reach space and follow a repeatable robot path.
Given the anchored 3D scene, when I navigate, then confined-space fly-through and spatial asset search resolve to the same assets as the 2D and map views.
Given a robot patrol path, when replayed, then the navigation follows the beacon-anchored route for repeatable coverage.
The story files under docs/portfolio/product/user-stories/ are the source; the handbook page and the milestone review are regenerated from the same text.
What a pilot runsThe inspection-evidence integrity loop, end to end, on the operator's own imagery — process telemetry only if they connect it.
A 90-day pilot starts from inspection imagery the operator already captures: candidates are screened, findings are raised as evidence-grounded claims on the mechanism catalogue, an engineer confirms or dismisses them with reasoning, and actions and evidence-bound reports follow. No LiDAR, no engineering CAD and no SCADA connection are required to start; each is an entry the same chain gains later. The subscription tiers carry the loop in every tier, with the process-telemetry entry — operating-window checks and quantified risk — from Q4 2026 in Professional and above.
What this needs from you
Two dates to hold. The M4 close-out checkpoint at the end of September, where every M4 requirement is flipped or moved with evidence — the rule since v3.9, and the reason these tables can be trusted. And the first campaign an engineer actually works through the loop, which is what turns the alpha column from built into used.
One input: which evidence modality's data to secure next. Survey readings are already in hand and unlock the second source; process telemetry needs pilot access that has slipped twice.
The architecture, and how we work in it
Sections are collapsed — the summaries are the overview.
Where things liveTwo runtimes, one API, one place where access is decided.
Two runtimes, one API, one place where access is decided. The analysis service has no private route to the database — it comes back through the same API carrying the caller's identity, which is why an assistant inherits exactly the permissions of the person who asked.
What holds it togetherEvery boundary that two things must agree on is a defined object.
Every boundary that two things must agree on is a defined object. The loop, not the boxes, is the mechanism: a surface that no longer matches its definition cannot merge.
How a change movesNobody is asked to remember to update the API spec, the client, the command line or the connector.
Nobody is asked to remember to update the API spec, the client, the command line or the connector. They are outputs, and the build is what notices when they stop matching. A contract change begins one step earlier, with a test written to fail.
“Why do we need a command line at all?”The scepticism is right. It is aimed at the wrong command line.
A hand-written CLI is a liability — it drifts from the API, rots quietly, and turns into a support burden nobody owns. This one is emitted from the same definition as the API itself. Nobody writes it, and it cannot drift. So the question is not whether it is worth building; it is whether it is worth deleting something that already costs nothing to keep. Four things it does that nothing else does as cheaply:
It is the cheapest proof the API is honest.Because it is generated from the specification, it exercises the specification. A parameter the spec invents, an authentication claim that is not actually true — the command line fails on it immediately and loudly. We can probe every operation and report exactly where the spec’s security claims do not match what the endpoint does. That is contract verification, not developer convenience.
It is how we debug the assistant.An agent takes an action and something looks wrong. With a command line, an engineer reproduces that exact call by hand, under the same identity and the same gates, and sees the same result. Without one, the same investigation is log-reading and inference. The more we let the assistant do, the more this is the difference between a ten-minute answer and a lost afternoon.
It is a customer channel we have already shipped.A customer’s operations team wanting a nightly export, a bulk import, or a hook into their own scheduler is currently a request to build an integration. Generated, it is already there — one projection, rather than one project per customer.
It is how production gets diagnosed at two in the morning.Making the same call the application makes, with the same identity, without opening a browser or hand-assembling a token. Every minute spent reconstructing a request is downtime.
Assistant apps as a fifth surfaceThe pieces exist. One of them is the least protected thing we own.
Hosts that render a tool’s result as interactive UI inside the conversation — rather than as text — are a natural fifth projection, and the field that declares what a tool can render is already the hook for it. Two of the three pieces are built: the connector is live, and we already have a declarative agent-to-UI protocol whose stated premise is components “safely rendered by clients without executing arbitrary code.” That is the same premise these hosts are built on, written down before we needed it — twenty-four components, eighteen generic and six of our own domain.
Gate the render contract before adding a consumer.It is the one defined object with no schema, no generated output, and no test or CI reference anywhere. Six of its components duplicate payloads the event contract already carries, so we have two rendering paths for the same data and only one of them is verified. A third consumer on an ungated contract is how drift starts.
Never let a tool emit host-specific UI.Hosts differ in widget vocabulary and capability, and they change under us. A tool should declare what it means — a gallery of these images, with these fields — and each host projects that into its own rendering. Emitting one host’s components directly would couple our tool definitions to someone else’s product roadmap.
The exposure rules already cover it.A new surface inherits the same defaults: denied unless declared, no elevated-authority tool reachable, confirmation required before a model takes a consequential action. Adding a host is a declaration, not a new security review.
“Why not just give the assistant its own database access?”Because then you implement the security model twice — and only one copy gets tested.
Authorization stops being written twice.Every access rule already lives in the database. A private path means reimplementing organisation scoping, role checks and row filters inside the assistant — a second copy of the security model that will drift from the first, and that only the assistant’s tests cover. Acting as the caller means a new tool ships with zero new authorization code.
A prompt injection is bounded by that user’s own permissions.Untrusted text reaches the model constantly — a note field, a document, a caption on an uploaded image. Acting as the person, the worst a malicious instruction achieves is what that person could already do themselves. Holding a service credential, the same injection reads every organisation on the platform. That is the difference between an incident and a breach, and it is decided by this one design choice.
There is no long-lived key to leak.A service credential sitting in the assistant’s environment can be exfiltrated, written to a log, or coaxed into a response. A short-lived token scoped to one person expires on its own and is worth little if it escapes.
The audit trail is already correct.Writes carry a user identity because they always did. There is no second “the assistant did this” trail to build, and no reconciliation between two records of who changed a finding.
Every existing gate applies for free.Validation, rate limits, confirmation requirements, business rules — the assistant inherits all of them by coming through the same door. A private path would need each one re-implemented and re-tested on a second route.
The two fair objections, answered honestly. The extra hop is real, but it is the same hop the browser makes — if it is too slow for the assistant it was already too slow for the application, and it gets fixed once for both. And genuinely user-less work, like scheduled analysis, does need a service identity: the answer is a named service principal with its own narrow permissions, declared in the tool’s authority field and reviewed like any other, rather than the assistant borrowing broad access and giving it back.
The evidence, from one releaseFour copies drifted from what they described. The build caught none of them, because the build was not looking.
Everything on this page argues that a definition should have one home and every surface should be generated from it. Cutting the 3.10 product-document release turned that argument into a list. Each of these was a hand-maintained copy of something that already had a source of truth:
A guard holding a copy of the data it guards.The consistency linter carried a hard-coded list of the milestones it exists to protect, so adding one meant editing the check. Replaced with a rule about shape — the identifiers must be contiguous and in order — which still catches a drop, a reorder or a typo without duplicating what it validates.
A dated artifact checked against a live source.A planning deck cut for one quarter was required to name every milestone in the current ledger. Committing anything for 2027 would have forced us either to falsify a historical document or to block the roadmap. The check now derives the deck's own cycle and asks only for what existed by then.
Twenty-six sentences that changed meaning without being edited.One label was carrying a date and a commitment level. When the first half of 2027 became committed, prose written when it meant "directional" silently began reading as a promise — including customer-facing availability claims. Found and swept by hand; five historical entries deliberately left alone.
A release recipe contradicting its own decision.The instructions said to freeze copies of each derived document into the release folder. The recorded decision said the opposite, and the evidence agreed with the decision: no release has ever had that folder. The recipe had simply never been corrected.
The pattern is the point. None of these was carelessness, and none was caught by review — three were caught only because a build refused to proceed, and the fourth because someone went looking for prior art that turned out not to exist. That is the difference between a convention and a guarantee, and it is the whole argument for generating a surface rather than maintaining it.
One level deeper — what each capability is made ofEach capability expands into the themes it owns — the CTO altitude.
World ModelAsset identity, spatial registration, tag reconciliation, engineering context, photorealistic 3D scene (3DGS), AI assistant brain (persistent memory and knowledge)0 / 11+1 alpha+3 in progress
Damage mechanism review on inspection evidence0 / 3
Operator HandoffVerification queue, recommendation packets, inspection plan handoff0 / 10+6 alpha+0 in progress
Reports and recommendation packets0 / 3
IDMS and work-management integration0 / 2
Partner-integrated delivery0 / 1
Proactive coverage prompts0 / 1
Verification queue and campaign hand-off0 / 3
Each capability expands into the themes it owns — the CTO altitude. Counts are delivered requirements over filed. A zero is not always an empty theme — it can mean the pieces exist but nothing wires them together yet, and 0 / 0 means the capability claims the theme but nothing is filed under it. World Model's CAD and P&ID requirements are the concrete case: a 135,277-row extraction is real and queryable, but no repeatable ingestion path writes to it, and a reconciliation table and its resolver function both exist correctly and have zero callers between them. See docs/reports/20260903_cad_pid_data_vs_pipeline.md.
The rule, and what it buysDesign the contract first where two runtimes must agree; let structure emerge everywhere else — and where the model is not yet complete.
One rule decides where a contract is designed up front
Design the object first wherever two runtimes must agree and neither can read the other’s source — that is the case the compiler cannot help with. Everywhere a single team owns both sides and can refactor cheaply, let the structure emerge and keep the definition generated from the code. That test, rather than a preference for specifications or for agility, sorts every boundary we have.
The property this buys is narrow, and worth stating exactly
A specification generated from the code cannot lie about what runs. A contract generated into both runtimes cannot drift between them. Neither claim rests on review discipline — which is why they survive a team growing, a new hire, or a quarter getting busy.
Where the model is not yet complete
Tools are the gap. The schema a tool declares is not yet proven to be the schema the model receives, and more than one vocabulary describes what a tool does to the world. The design closes both, and the first step is a test that fails today — the honest way to know a gap is real rather than theoretical.
What this needs from you
Agreement on the rule itself, since we will apply it to the next boundary as well as this one. And a view on whether the tool object belongs in the shared contract package beside the event types — that placement is what makes it generate into both runtimes rather than one.
The model, four lines
Define the object at the boundary.
Author once, in the implementer’s language.
Generate every surface from it.
Verify by regenerating in CI.
Design first when…
Two runtimes must agree.
Neither can read the other’s source.
The mistake is a migration, not a refactor.
Let it emerge when…
One team owns both sides.
The definition can be generated from the code.
Reversing costs an afternoon.
What does this change about how we deliver — and how do I add one?
Sections are collapsed — the summaries are the overview.
How delivery changesA feature has a known shape, a new surface is a projection, correctness is structural — and what is left.
A feature has a known shape
Define the object, implement it once, and the surfaces come out. Estimates become repeatable, and a feature is reviewable as one vertical slice rather than four loosely related changes landing over three weeks.
Adding a surface is a projection, not a rewrite
The command line is already generated from the API definition rather than written against it. Bringing a new surface online means writing the projection once, for every object — not an adapter per capability, forever.
Correctness is structural, not vigilance
The failures that would otherwise depend on someone noticing in review — a specification drifting from behaviour, two runtimes disagreeing, a generated file edited by hand — are build failures. That is the part that keeps working when the team is busy, which is exactly when it matters.
What is left, and what it costs
The tool object and the three tests that hold it. Durable background work only when a real workload needs it — we already have progress and streaming, so what we would be adding is only durability. Sequencing starts with one question about the agent framework that takes about half a day and prices the rest.
Two levels deeper — every requirement, with its milestone and statusEach theme opens into the requirements filed under it — the engineering altitude.
How to add oneAuthor in your own runtime, declare what it does, exposure defaults to deny, write the failing test first.
Author the object in your own runtime
Schemas are authored in the language of whoever implements the thing, and emitted to a portable schema as the interchange. Nobody writes in another runtime's idiom to participate. What crosses the boundary is the emitted schema, not the authoring format.
Declare what it does, not just what it takes
Identity and version; the human title and the model-facing description; input and output; what it can render; what it does to the world; whose authority it acts under; and a binding that is either an existing API operation or a callable. The binding is what lets a capability that is not an HTTP endpoint be a first-class tool anyway.
Exposure defaults to deny
A capability reaches a surface because someone declared that it should, visible in a diff. The invariants that follow are assertions rather than conventions — among them, a tool running under elevated authority is never reachable through the assistant connector, and a tool that changes something must require confirmation before a model may call it.
Write the failing test first
Assert that the schema a model receives equals the one that was declared. It fails today. That output is the specification for what finishing means — and it stays afterwards as the gate that would have caught it the first time.
Long work is an execution mode, not a second object
Duplicating input, output, effect and authority into a parallel definition would recreate the drift we are removing. The rule: durable state is warranted when the answer must outlive the conversation that asked for it. Everything shorter rides the stream we already have.
Skills sit one layer above tools — they compose them
A skill is a versioned prompt that uses existing tools to solve one specific problem. Two exist today, as markdown with a version and a description that says when to load it: kap-cad-pid composes kap_query and kap_execute_sql to answer questions about the digital twin, P&ID and equipment tags; kap-dataset-search composes five read tools to answer questions about a caller's own campaigns, images and readings. Its shape is identity and version, the problem it solves, the tools it is allowed to use, the instructions, and any reference files. What it does not declare — and should not — is effect or authority: a skill can do nothing its tools cannot, so its effect is the strongest effect among the tools it composes, and it acts under whoever invoked it. Exposure and confirmation therefore fall out of the tool list rather than being declared twice. The thing the code calls ImageSkill is not a skill under this definition: it is a prompt bound to a model that produces a result — a tool with a model-call binding — and it belongs in the Tool object, where its missing effect and authority fields are the Tool gap already described above.
What this needs from you
Someone to answer the framework question — whether the agent framework accepts an explicit schema, or whether the callable's signature must be generated from it — then write the failing test. Neither needs design work first.
And a decision on whether new capabilities are expected to ship with their projections from day one. That convention is cheap to hold now and expensive to retrofit.
Skill does not fold into Tool: they are different layers, and a skill's whole value is that it is made of tools. What folds into Tool is ImageSkill. The open question is narrower — whether a skill's uses list should be a checked field rather than prose, so that "which tools may this skill call" is answerable by the build, the way "which tools may this engine call" will be once the Tool object lands.
Where does the catalogue stand — every requirement, one picture?
The 84 requirements as 84 blocks, grouped by milestone or by capability, filled where the work has landed: 15 delivered, 12 alpha (prototype), 4 in progress; the other 53 are targets. The next gate is M4 · Persistent Sensing (Q3 2026). Nothing here is typed: the blocks, their fills and every count come from the catalogue.
Group by
One block per requirement; within a band the landed work sits first, so the filled front of each row is the progress. Filled means the requirement has reached a bar — Delivered on the production trunk, Alpha (prototype) a functional prototype on the alpha lineage, counted beside Delivered and never in it, In progress under way. Outlined blocks are dated targets. By milestone, the bands run in ledger order and the last band is the horizon, directional rather than committed; by capability, the five deep modules come first, then the cross-cutting concerns that own something. The count at the end of a band is delivered over filed. Hover a block for the requirement; click it to open it in the Requirements tab.
Where does one requirement stand — and what else is filed beside it?
The 84 requirements of the product document as one flat list: search by identifier or words, narrow by status, milestone, capability or area, sort by any column. Every cell is the catalogue's own — feature, status and milestone come through the same variables as every other tab, and the filter keys are generated from the same file — so a re-filing in requirements.yaml changes this list and nothing else has to be edited.
84 of 84
Feature
Scenarios
FR-APP-07US-M0-02
Web-based 3D inspection viewer (Cesium geospatial scene)
Application Surface3D viewer and overlays
App
Critical
M0
Q3 2025
Delivered
0 · 0 verified
FR-APP-08US-M0-03
Operator dashboard
Application SurfaceDashboard and defect gallery
App
High
M0
Q3 2025
Delivered
0 · 0 verified
FR-APP-09US-M0-04
Visual defect gallery
Application SurfaceDashboard and defect gallery
App
High
M0
Q3 2025
Delivered
0 · 0 verified
FR-EVI-01US-M0-01
RGB inspection imagery ingestion
Evidence IntakeSensor ingestion
Platform
High
M0
Q3 2025
Delivered
0 · 0 verified
FR-SEC-04US-M0-05
Multi-tenant authentication & workspace access control
Security & ComplianceAccess control and tenant governance
Platform
Critical
M0
Q3 2025
Delivered
0 · 0 verified
FR-AI-05US-M1-02
First-generation AI chat assistant
Application SurfaceContextual data chat
AI Assistant
High
M1
Q4 2025
Delivered
0 · 0 verified
FR-APP-10US-M1-03
Automated agentic task coordination
Application SurfaceAgentic task coordination
AI Assistant
High
M1
Q4 2025
Delivered
0 · 0 verified
FR-APP-12US-M1-04
Gas measurement visualization
Application Surface3D viewer and overlays
App
Medium
M1
Q4 2025
Delivered
0 · 0 verified
FR-EVI-02US-M1-01
Multimodal evidence pipeline foundation
Evidence IntakeSensor ingestion
AI Assistant
High
M1
Q4 2025
Delivered
0 · 0 verified
FR-AI-06US-M2-01
Structured dataset-list responses
Application SurfaceStructured assistant responses
AI Assistant
High
M2
Q1 2026
Delivered
0 · 0 verified
FR-AI-07US-M2-02
Structured image-gallery responses
Application SurfaceStructured assistant responses
AI Assistant
High
M2
Q1 2026
Delivered
0 · 0 verified
FR-APP-11
Provenance-cited chat answers
Application SurfaceContextual data chat
App
Medium
M2
Q2 2026
In Progress (Q2)
0 · 0 verified
FR-OPS-01US-M2-03
Work-order / recommendation export
Operator HandoffReports and recommendation packets
App
Medium
M2
Q2 2026
Alpha (prototype)
0 · 0 verified
FR-APP-02US-M3-01
Contextual data chat Ph.0
Application SurfaceContextual data chat
AI Assistant
Critical
M3
Q2 2026
Delivered
0 · 0 verified
FR-APP-13US-M3-08
Agent failure recovery
Application SurfaceAssistant reliability and safety
AI Assistant
Critical
M3
Q2 2026
Delivered
0 · 0 verified
FR-APP-14US-M3-09
Contextual chat agent evaluation gate
Application SurfaceAssistant reliability and safety
Evidence IntakeAutonomous patrol and site infrastructure
Platform
High
M6
Q1 2027
Q1 2027 Target
0 · 0 verified
FR-APP-21
Proactive coverage prompts
Operator HandoffProactive coverage prompts
AI Assistant
Low
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-NAV-01US-H1-03
Full spatial navigation
Evidence IntakeAutonomous patrol and site infrastructure
App
Medium
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-ROB-02US-H1-02
Fixed infrastructure: navigation beacons
Evidence IntakeAutonomous patrol and site infrastructure
Infrastructure
High
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-ROB-03US-H1-02
Fixed infrastructure: communication backbone
Evidence IntakeAutonomous patrol and site infrastructure
Infrastructure
High
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-ROB-04US-H1-01
Coverage orchestration
Evidence IntakeAutonomous patrol and site infrastructure
Platform
Medium
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-ROB-05US-H1-01
Fleet intelligence analytics
Evidence IntakeAutonomous patrol and site infrastructure
App / AI Assistant
High
M7
Q2 2027
Q2 2027 Target
0 · 0 verified
FR-APP-20
The assistant asks
Application SurfaceProactive assistant
AI Assistant
Medium
H2-2027
H2 2027
H2 2027 Roadmap
0 · 0 verified
FR-NUC-01US-M5-09
Dose-aware inspection workflow
Deployment ProfileHigh-hazard and regulated sites
Platform
Medium
H2-2027
H2 2027
H2 2027 Roadmap
0 · 0 verified
Default order is milestone, then identifier. Hover a row for the requirement's description. Alpha (prototype) is a functional prototype on the alpha lineage; Delivered means the production trunk. The Product Manager tab gives the same list as buckets, the Engineering tab as a drill-down by capability and theme.
What is built, and what is next
Most of the model is running and enforced on every change. This is the honest status.
Written at the 3.10 cut (2026-09-02); the counts on this page are live at 3.12 (2026-09-03). The proactive assistant moved from a design to four dated requirements, and the first half of 2027 stopped being a horizon: it is now two committed milestones — Repeat Visit (Q1 2027) and Continuous Coverage (Q2 2027). Everything below is generated from one source and checked on every change, which is why the release propagated to every derived document without anyone updating them by hand.
By technical boundary object
Object
What it powers
Status
Data
Typed access to every record, generated from reviewed migrations
Built
Access rule
Who can see and change what — enforced in the database itself
Built
Error
A stable code and a stated action, so a caller can respond rather than guess
Built
Operation
The published API, the typed client, the command line, the assistant connector
Built
Agent event
Everything the assistant streams back into the interface
Built
Tool
What the assistant can do, on all four surfaces at once — the one piece still missing
Next
Job
Work that outlives the conversation that asked for it
On demand
Skill
A versioned prompt that composes existing tools to solve one problem — effect and authority derive from the tools it uses