Plane 03

Assurance.
Prove it held.

Governed execution should leave a trail worth reading. Assurance turns everything the Runtime does into a reconstructable account, a continuous test, an accountable cost, and evidence an examiner can query.

01Assurance · Beyond uptime

Can you reconstruct why an agent did what it did, or only confirm that it ran?

The trail that debugs an agent is the trail that answers a regulator.

Decision observability is a reconstructable account of an agent's reasoning and tool calls, not a log that it executed. It is also, not by coincidence, the examinable record.

Layer one · infrastructure

Latency, throughput, cost, errors. Necessary, well understood, and where most tools stop.

Layer two · quality

Drift, hallucination, faithfulness, grounding. Is the output still correct, not merely produced.

Layer three · behavioral and decision

Did the agent stay in mandate. Did tool calls stay in scope. Why did it choose this path.

The failure if you stop at layer one

You can report that an agent was fast and cheap while it was quietly acting outside its mandate. The layers are not substitutes. Behavioral telemetry is where autonomy is actually governed.

Replaces

AI observability platformAgent tracingRisk operations dashboard

Decision trace · one flat log line, expanded

rebalance_agent · action proposed

read: portfolio_positions (entitled)

reasoning: drift exceeds mandate band

tool call: order_mgmt.place (rung 3)

mandate check · in scope

authority check · cleared by MR-014

committed · dossier written

02Assurance · The testing surface

Is assurance a gate you pass once, or a discipline that runs as long as the agent does?

A launch gate certifies a snapshot of a system that moves.

A model that passed at deployment can fail next month as inputs shift and the world changes. Assurance has to run continuously, and it has to test for more than function.

Faithfulness and grounding

Does the output follow from the evidence, or is it plausible invention.

Bias and fairness

Across the populations a fiduciary serves. Tested, not assumed, and traced to adverse action explainability.

Adversarial resistance

Prompt injection, tool poisoning, and attempts to escalate authority, run as a standing red team.

Trajectory correctness

For a multi step agent, was the path sound, not only the final answer.

The external trust layer

Certifiable standards make assurance legible to third parties: ISO 42001 for the management system, agent level standards for the autonomous systems. Internal testing becomes an external credential.

Replaces

LLM and agent evaluationRed teaming and adversarialCertification and attestation

Continuous assurance · a score that does not hold still

score dips as inputs shift · a re-test fires automatically · assurance recovers

03Assurance · Spend is a control

Is every autonomous decision attributable to a budget, an owner, and a unit cost?

A consumption spike is a behavioral alarm.

The same chokepoint that routes model calls and applies policy also meters and caps spend. Instrument once, and cost governance and behavioral governance share the same seat.

Budgets as blast radius limits

Per agent and per workflow ceilings a process cannot exceed. The bound contains a runaway loop before it becomes an incident.

Spend anomaly as early warning

A spike fires before quality metrics move, because a changed agent burns resources before it produces a visibly wrong result.

Chargeback makes autonomy accountable

Every agent's cost lands on a named owner. No one accountable for the bill means no one accountable for the behavior.

Unit economics prove the case

Cost per governed decision against the cost it replaces. The number that tells you where autonomy belongs and where it is theater.

Replaces

AI gateway and routingFinOps for AIMetering and chargeback

Spend anomaly · the earliest behavioral alarm

spike fires before quality metrics move · loop halted at the ceiling · cost lands on a named owner

04Assurance · The payoff

When the examiner arrives, is it a query against a system, or a project staffed for a quarter?

If evidence is a byproduct of governed execution, the exam is a query.

This surface converts a defensible operating model into a defensible regulatory posture. Every governed action becomes Living Evidence, immutable and mapped to the control it satisfies, packaged the way an examiner wants to read it.

Immutable action logs

Every consequential action, with its authority basis, written once and unalterable.

Decision dossiers

The reconstructable account of why an agent acted, assembled on demand from the telemetry.

Control attestation

Continuous proof that each control was in force, not a point in time assertion signed once a year.

Map once, prove everywhere

One control test satisfies every framework that asks for the same thing. Coverage stops being a per framework project.

The exam becomes a query

Ask the system a question a regulator would ask, and the record answers with the dossier and the evidence attached. The examination turns into a search box.

Replaces

Evidence and audit trailControl mapping engineContinuous compliance

Proof on demand · the exam becomes a query

?

show the basis for the rebalance on 14:32:07

Decision dossier · rebalance_agent, on behalf of advisor A. Chen, rung 3, cleared by control MR-014.

Immutable log · authority basis, data touched, and reasoning, written once and unaltered.

Satisfies · SR 11-7, NIST AI RMF, Reg BI. Evidence attached.

See it on your own estate.

A 30 minute working session against a control, an agent, or a rule you already answer to.

Book a demo