Mark HollandSenior AI Solutions Engineer

Observability & Governance Engine

Built, stage oneDesigned at work for a national healthcare data company

An engine that sits between an organization's AI applications and the model providers. It records every AI call with the exact data versions behind it, keeps those records sealed, reports four numbers from them, and in later stages runs the organization's own checks on answers before a user sees them. Datadog stays for infrastructure health. The engine feeds it numbers, never text.

The first version of this page described a layer that sat beside Datadog and enriched what it already collected. Building it changed the design. A monitoring tool only sees what already happened. To prove what went into an answer, and later to act on one, the engine has to sit in the path of the call itself. Stage one, recording only, is built and tested. The stages after it are gated on real traffic.

Stage one, built: every AI call goes through the gateway, which writes one sealed record per answer. Check, Hold, and Meter are later stages.

What it does

One sealed record per answer

  • Every AI call passes through a gateway on its way to the model and back. The gateway is an adopted open source proxy, not one I wrote, and it speaks the model provider's own API, so the application's request goes through unchanged.
  • The gateway writes one record per answer: a pseudonymous user reference, the application and its version, the prompt template and its version, the data versions and source references the answer drew on, what came back, what was done with it, and how much delay the engine added.
  • Records are insert only, chained by hash, anchored on a schedule, and stored under object lock so they cannot be altered or deleted before the retention period ends. Each tenant's records are encrypted under that tenant's own key.
  • A record is opened only under an approved, time limited grant, and every read is logged. Reviewers see what they were granted and nothing else.
  • Datadog receives counts and timings only, through a collector that strips everything else. A test captures everything the collector would send and fails if any prompt text, answer text or user identifier is in it.

Why this and not a monitoring tool

A different question

CloudWatch and Datadog answer whether the system is healthy: up, fast, error free, trending which way. They are built for operators. The engine answers a question they cannot: for this specific answer that went to this specific person, what went into it, and can we prove it.

Three things only the engine gives you.

  • Proof per answer

    A sealed record that can be rebuilt months later with the exact source versions, for an auditor, a client or a complaint.

  • Control in the path

    Because the engine sits in the path of every AI call, it is built to check an answer and hold it for a person before the user sees it; holding arrives in a later stage. CloudWatch and Datadog only see what already happened, so the best they can do is tell someone after the fact.

  • Protections built in rather than bolted on

    Prompt and answer text is stored only in the engine's locked record store and never reaches a log or a dashboard. Each client's records are encrypted under their own key, so if a client asks for their data to be deleted, destroying that key makes their records unreadable for good with no effect on anyone else. And no record is opened without an approved grant, with every read logged.

The four numbers

What stage one reports

The four numbers stage one reports, and the condition that stops the next stage
NumberStop condition
Added delay at the 95th percentile per runOver 50 milliseconds
Share of runs with a complete record, checked against the application's own count of runsUnder 99 percent
What the checks would have caught, with a labeled sample of false alarmsNo stop condition; this shapes stage two
Reviewer hours the flag rate impliesMore than the team can staff
Numbers 3 and 4: in stage one, the checks run in the background against the recorded answers and never change what a user sees.

Why the gate

Two weeks of real traffic

Stage one cannot change, slow or block an answer, so it is safe by design. Stage two is the first time the engine can flag an answer or stop one, and at that point a mistake in the engine is a mistake a user sees. So stage two does not start until the four numbers have been measured over at least two weeks of production traffic and signed off. Two weeks rather than two days because a short window only sees one kind of day; two weeks covers weekdays, weekends, at least one release and the busy stretches, with enough runs for the numbers to mean something.

Rules that never change

Nine rules the code is tested against

  1. Checks in the request path fit in 50 milliseconds at the 95th percentile.
  2. Checks that use a model to judge never block; only deterministic checks can.
  3. Every check starts in watch mode and is promoted on measured false alarm rates.
  4. The tenant comes from the credential, never from anything the caller sends.
  5. One encryption key per tenant.
  6. Sources are recorded by reference and hash, never by copy.
  7. Full text never reaches Datadog or a log line. A test scans every package for it.
  8. The emergency bypass still records; it never skips recording.
  9. Stage one changes no answer.

Stages

What comes after recording

  1. RecordDone

    Built. The gateway, the record store, the access grants, the four numbers report, and the reviewer console, where approved grants open records.

  2. Check (not started)

    The organization's own checks run on each answer in the request path, in watch mode first, then enforcing once their false alarm rate is known and accepted.

  3. Hold (not started)

    A flagged answer waits for a person, with a review queue sized to the measured flag rate.

  4. Meter (not started)

    Usage counted against customer contracts, and a second application onboarded.

How it was built

Specs first, then prompts

I wrote the specifications first: an overview with the rules above, the architecture, the record schema, storage and keys, security, telemetry, a build order with a gate after each stage, and a decisions file that lists every default chosen so the build never re-decides them. Then a set of build prompts, run in order with Claude Code, each naming the tests that must pass before the next one starts. Anything the specs did not settle went into a decisions pending file with the safer default chosen in the meantime, rather than being assumed in code. Two independent reviews of the finished stage one found no rule broken; they did find four problems in the integration patch for the host application, all fixed before it was handed over.

  • Over 600 tests, including a security suite that scans for content in logs and metrics, cross tenant reads, grant misuse and key destruction.
  • A synthetic evidence run of 5,000 calls on a laptop: median added delay 5.3 ms, 95th percentile 26.5 ms, every record complete, full reconstruction of a record with its six sources in 2.5 seconds.
  • Built and tested on a machine with no Docker, with a native services script standing in for the container stack.

Stack

  • Python 3.12
  • FastAPI
  • LiteLLM proxy
  • PostgreSQL 16
  • S3 with Object Lock (MinIO locally)
  • KMS
  • Redis
  • SQS
  • OpenTelemetry collector
  • Datadog
  • Terraform on AWS ECS Fargate
  • pytest
  • Claude Code