Observability & Governance Engine
Built, stage oneDesigned at work for a national healthcare data company
An engine that sits between an organization's AI applications and the model providers. It records every AI call with the exact data versions behind it, keeps those records sealed, reports four numbers from them, and in later stages runs the organization's own checks on answers before a user sees them. Datadog stays for infrastructure health. The engine feeds it numbers, never text.
The first version of this page described a layer that sat beside Datadog and enriched what it already collected. Building it changed the design. A monitoring tool only sees what already happened. To prove what went into an answer, and later to act on one, the engine has to sit in the path of the call itself. Stage one, recording only, is built and tested. The stages after it are gated on real traffic.
What it does
One sealed record per answer
- Every AI call passes through a gateway on its way to the model and back. The gateway is an adopted open source proxy, not one I wrote, and it speaks the model provider's own API, so the application's request goes through unchanged.
- The gateway writes one record per answer: a pseudonymous user reference, the application and its version, the prompt template and its version, the data versions and source references the answer drew on, what came back, what was done with it, and how much delay the engine added.
- Records are insert only, chained by hash, anchored on a schedule, and stored under object lock so they cannot be altered or deleted before the retention period ends. Each tenant's records are encrypted under that tenant's own key.
- A record is opened only under an approved, time limited grant, and every read is logged. Reviewers see what they were granted and nothing else.
- Datadog receives counts and timings only, through a collector that strips everything else. A test captures everything the collector would send and fails if any prompt text, answer text or user identifier is in it.
Why this and not a monitoring tool
A different question
CloudWatch and Datadog answer whether the system is healthy: up, fast, error free, trending which way. They are built for operators. The engine answers a question they cannot: for this specific answer that went to this specific person, what went into it, and can we prove it.
Three things only the engine gives you.
Proof per answer
A sealed record that can be rebuilt months later with the exact source versions, for an auditor, a client or a complaint.
Control in the path
Because the engine sits in the path of every AI call, it is built to check an answer and hold it for a person before the user sees it; holding arrives in a later stage. CloudWatch and Datadog only see what already happened, so the best they can do is tell someone after the fact.
Protections built in rather than bolted on
Prompt and answer text is stored only in the engine's locked record store and never reaches a log or a dashboard. Each client's records are encrypted under their own key, so if a client asks for their data to be deleted, destroying that key makes their records unreadable for good with no effect on anyone else. And no record is opened without an approved grant, with every read logged.
The four numbers
What stage one reports
| Number | Stop condition |
|---|---|
| Added delay at the 95th percentile per run | Over 50 milliseconds |
| Share of runs with a complete record, checked against the application's own count of runs | Under 99 percent |
| What the checks would have caught, with a labeled sample of false alarms | No stop condition; this shapes stage two |
| Reviewer hours the flag rate implies | More than the team can staff |
Why the gate
Two weeks of real traffic
Stage one cannot change, slow or block an answer, so it is safe by design. Stage two is the first time the engine can flag an answer or stop one, and at that point a mistake in the engine is a mistake a user sees. So stage two does not start until the four numbers have been measured over at least two weeks of production traffic and signed off. Two weeks rather than two days because a short window only sees one kind of day; two weeks covers weekdays, weekends, at least one release and the busy stretches, with enough runs for the numbers to mean something.
Rules that never change
Nine rules the code is tested against
- Checks in the request path fit in 50 milliseconds at the 95th percentile.
- Checks that use a model to judge never block; only deterministic checks can.
- Every check starts in watch mode and is promoted on measured false alarm rates.
- The tenant comes from the credential, never from anything the caller sends.
- One encryption key per tenant.
- Sources are recorded by reference and hash, never by copy.
- Full text never reaches Datadog or a log line. A test scans every package for it.
- The emergency bypass still records; it never skips recording.
- Stage one changes no answer.
Stages
What comes after recording
RecordDone
Built. The gateway, the record store, the access grants, the four numbers report, and the reviewer console, where approved grants open records.
Check (not started)
The organization's own checks run on each answer in the request path, in watch mode first, then enforcing once their false alarm rate is known and accepted.
Hold (not started)
A flagged answer waits for a person, with a review queue sized to the measured flag rate.
Meter (not started)
Usage counted against customer contracts, and a second application onboarded.
How it was built
Specs first, then prompts
I wrote the specifications first: an overview with the rules above, the architecture, the record schema, storage and keys, security, telemetry, a build order with a gate after each stage, and a decisions file that lists every default chosen so the build never re-decides them. Then a set of build prompts, run in order with Claude Code, each naming the tests that must pass before the next one starts. Anything the specs did not settle went into a decisions pending file with the safer default chosen in the meantime, rather than being assumed in code. Two independent reviews of the finished stage one found no rule broken; they did find four problems in the integration patch for the host application, all fixed before it was handed over.
- Over 600 tests, including a security suite that scans for content in logs and metrics, cross tenant reads, grant misuse and key destruction.
- A synthetic evidence run of 5,000 calls on a laptop: median added delay 5.3 ms, 95th percentile 26.5 ms, every record complete, full reconstruction of a record with its six sources in 2.5 seconds.
- Built and tested on a machine with no Docker, with a native services script standing in for the container stack.
Stack
- Python 3.12
- FastAPI
- LiteLLM proxy
- PostgreSQL 16
- S3 with Object Lock (MinIO locally)
- KMS
- Redis
- SQS
- OpenTelemetry collector
- Datadog
- Terraform on AWS ECS Fargate
- pytest
- Claude Code