From AI Idea to ProductionPart 6 of 20
Building Beyond the Demo
A demo proves an AI idea can work. Production proves it keeps working. The six layers I add before I call an AI system production-ready.

A demo proves the idea can work. Production proves the system keeps working after real users, real data, real integrations, and real failures show up.
In Part 5 I wrote about creating the spec before writing code. This part covers the next problem. Once you have a spec and a prototype that works, how do you build the real thing so it survives contact with the real environment?
The most common mistake I see is taking a prototype that impressed everyone in a meeting and simply adding features to it. It feels like progress. In practice it piles up technical debt faster than business value, because everything the prototype skipped on purpose is still missing. It just has more weight sitting on top of it now.
This article walks through the six layers I add before I call an AI system production-ready. For each one you get what goes wrong without it, how to set it up, and something you can copy into your own project.

A demo that works perfectly
Let me use a made-up example, because it is a pattern almost everyone building with AI will recognize.
Picture a customer support assistant built for an online retailer. In the demo, a manager types "Where is my order?" The assistant looks up the order, explains the delay, and offers a refund on the shipping charge. It is fast, polite, and accurate. The room is sold.
Here is what that demo is running on.
- One environment, which is the developer's laptop
- One API key, pasted into a config file
- One model, called directly from the code
- One prompt, edited by hand whenever something looks off
- Logs that print to the screen and disappear
- A fake order system that always returns a clean answer
- Twenty test questions the builder wrote
That is exactly the right setup for a demo. Its whole job is to answer one question, which is whether the idea is worth building. It is the wrong setup for something a business will depend on.
What breaks first
Imagine that assistant goes live the way it was demoed. None of what follows is exotic. These are the ordinary ways a prototype fails in the real world, and each one is fixed by a specific layer.
| What goes wrong | What the customer or business sees | The layer that prevents it |
|---|---|---|
| Security rotates the API key pasted in the config file | The assistant goes down for everyone at once | Secrets |
| A developer tweaks the prompt to sound friendlier | It starts offering refunds it used to escalate, and nobody can say which prompt was running | Versioning |
| The real order system times out at busy hours | The assistant gives wrong or empty answers | Modularity and environments |
| Someone copies a real customer conversation into the test set | Personal data now lives somewhere it shouldn't | Environments |
| A customer says a promised refund never arrived | Nobody can reconstruct what the assistant read, called, or decided | Logging |
| Automatic refunds misbehave on day one | The only way to stop it is to take the whole assistant offline | Feature flags |
None of these are solved by a better prompt. They are solved by structure.
Layer 1. Environments
Development, test, staging, and production should be separate, each with its own data, credentials, and settings. The rule I hold to is simple. Production data and production credentials never become convenient test fixtures.
| Development | Test | Staging | Production | |
|---|---|---|---|---|
| Purpose | Build and experiment | Automated tests on every change | Rehearse the real release | Real customers |
| Data | Synthetic only | Synthetic and fixed test sets | Representative, de-identified | Real |
| Credentials | Dev keys, low limits | Test keys | Staging keys | Production keys, tightest access |
| Integrations | Mocked | Mocked plus contract tests | Real systems in sandbox mode | Real systems |
| Who has access | Builders | Pipeline | Builders, QA, business owners | Operations, on-call |
| Logging detail | Everything | Everything | Production level | Production level, personal data masked |
How to set it up
- Give each environment its own configuration, its own credentials, and its own data store.
- Make the code read which environment it is running in. Never hard-code a URL, key, or database name.
- Build a synthetic data set that has the same shape as real data, including the messy cases like missing fields, very long messages, and unusual names.
- Block production credentials from working anywhere outside production.
Tip
If you need realistic data to reproduce a bug, generate a synthetic copy of the problem. Never copy the real record.
Layer 2. Modularity
The prototype is usually one long piece of code where retrieval, model calls, business rules, permissions, and logging are all tangled together. Production needs them to be separate parts with clear boundaries.
| Component | Its one job | What it must never do |
|---|---|---|
| Retrieval | Find the right documents and records | Decide what the customer is allowed |
| Model calls | Talk to the AI model through one shared doorway | Hold credentials or call systems directly |
| Business rules | Apply limits like "refunds over $50 need a person" | Live inside a prompt |
| Permissions | Decide what this user and this assistant may do | Be bypassed by clever wording |
| Tools | Perform the action, such as issuing a refund | Run without a permission check |
| Logging | Record what happened at every step | Store personal data it doesn't need |
The test is whether you can swap one part without rewriting the others. If changing models means rewriting the whole workflow, it is too tangled. If a refund limit lives inside a prompt instead of in plain code, someone can talk their way around it.
That last point is one I build into everything. In my own builds, the model can suggest an action, but plain code decides whether the action is allowed.
How to set it up
- Write down the inputs and outputs for each component before you split them. This is the contract between them.
- Put every model call behind one function, so changing models touches one place.
- Move every limit and permission out of the prompt and into code that the model cannot edit.
- Write a test for each contract, so a change in one component that breaks another fails immediately.
Layer 3. Versioning
Everyone versions their code. With AI systems, code is only part of what changes behavior. You need to be able to answer "what exactly was running at 2:14 on Tuesday?" for all of it.
| What to version | Why it changes behavior |
|---|---|
| Prompts | One sentence can change what the assistant will do |
| Model and its settings | A new model or temperature changes answers |
| Retrieval settings | Which index, how many results, and which filters change what the model sees |
| Tool definitions | Changing a tool's inputs changes what actions are possible |
| Business rules and permissions | Changes what is allowed |
| Evaluation set | Changes what "passing" means |
How to set it up
Keep prompts and settings in version control next to the code, not pasted into a dashboard. Give every release a short manifest, and stamp the release ID into every log entry. When behavior changes, you can line it up with the exact release that caused it.
A release manifest can be this simple.
release: support-assistant-2026.10.07-3
code: 4f2a91c
prompt: support-system-prompt v12
model: provider-model-name, temperature 0.2
retrieval: help-center-index v5, top 6 results, policy filter on
tools: order_lookup v3, refund_request v2
rules: refund-limits v4 (person required over $50)
eval_set: support-golden-set v9 (412 cases, 97% pass)
approved_by: release owner name
Layer 4. Secrets
API keys, database passwords, and tokens belong in a secret manager that is separate for each environment. They never belong in source code, config files, shared documents, or prompts.
How to set it up
- Move every credential into a secret manager, with separate credentials for each environment.
- Scan your code history for keys that were ever committed, and rotate any you find. Deleting the line is not enough, because the key is still in the history.
- Let tools fetch credentials at the moment they need them. The model never sees one.
- Practice a rotation. Change a key in staging and confirm nothing breaks and nothing needs a code change.
In a system I'm building for myself, credentials sit behind a proxy the agents never see. The agent can ask for an action, but it can't read, copy, or leak the key that performs it.

Layer 5. Logging
Logging that prints to a screen is for debugging. Production logging has to let you reconstruct exactly what happened in a single request, weeks later, for someone who wasn't there.
| What to record | Why it matters |
|---|---|
| A request ID that follows the request everywhere | Ties every step of one conversation together |
| Who asked, plus the environment and release ID | Explains context and links behavior to a release |
| What was retrieved | Shows whether the answer used the right information |
| Every tool call, with inputs and result | Shows what the system did, not just what it said |
| The decision and the reason | Shows why it acted, escalated, or refused |
| Errors, retries, and timeouts | Shows where it struggled |
| Time taken and cost | Shows whether it is fast enough and affordable |
| The final business outcome | Shows whether the result actually happened |
Here is what one masked log entry might look like.
{
"request_id": "req-7f3c",
"release": "support-assistant-2026.10.07-3",
"env": "production",
"user": "customer-****8812",
"retrieved": ["help/shipping-delays", "help/refund-policy"],
"proposed_action": "refund_shipping",
"rule_check": "passed, amount 7.99 under limit 50",
"tool_call": {"name": "refund_request", "status": "confirmed", "ms": 640},
"outcome": "refund_issued",
"latency_ms": 1710,
"cost_usd": 0.004
}
Watch out
Mask what you don't need, and decide up front how long logs are kept. A log full of personal data is a liability, not an asset.
Layer 6. Feature flags
A feature flag is a switch that turns one capability on or off without a new deployment. It sounds small. It is one of the most valuable things you can add to an AI system that takes actions.

| Stage | Who gets it | Move to the next stage when | Roll back when |
|---|---|---|---|
| 1 | Internal staff | A week with no wrong refunds | Any refund that breaks a rule |
| 2 | 5% of customers | Error rate and complaints match the old process | Complaints rise or a rule is broken |
| 3 | 25% of customers | Cost per refund is stable and reviewers agree with its decisions | Cost or error rate spikes |
| 4 | Everyone | Two weeks stable at 25% | Same triggers, flag stays in place |
Go back to the automatic refund example. Without a flag, a bad day means taking the whole assistant offline. With a flag, you switch off automatic refunds, the assistant goes back to sending refund requests to a person, and customers keep getting help with everything else while you find the problem.
A simple rollback plan, written before launch
- Who can flip the switch, and how to reach them after hours.
- The exact flag to turn off, and what the system does when it is off.
- How to confirm the "off" path is working.
- Who tells the business and customers, and what they say.
- How to find every action taken while the problem was live, using the logs.
Common pitfalls
- Running the demo in production and calling it a pilot
- Hard-coded keys anywhere in the code or config
- Skipping staging because "it worked in test"
- Logging that can't reconstruct a single request
- Integrations that assume the other system always answers quickly and correctly
- All-or-nothing releases with no way to turn off one feature
- No named owner for the system after launch
Why this is worth the effort
Production engineering can look like overhead, because none of it shows up in a demo. It shows up in the results.
| Green-dollar results (show up in the budget) | Blue-dollar results (operational value you measure on purpose) |
|---|---|
| Fewer incidents to respond to | Higher reliability |
| Less rework after bad releases | Faster, safer delivery of new features |
| Faster troubleshooting | More team capacity, less firefighting |
| Lower cost when something does go wrong | Confidence to scale to more users |
The demo proves what's possible. The architecture decides whether that possibility turns into dependable business value.
Production readiness scorecard
Score each layer 0, 1, or 2. Anything under 2 tells you where to work next.
0 of 12
0 of 6 scored.
Checklist before you call it production-ready
Tick each box you can say yes to.
0 of 14 done
Next in the series
Part 7 is about testing AI systems properly. That means going beyond "the answer looked right" to test retrieval quality, tool use, permissions, recovery, speed, cost, and whether the business outcome actually happened.
What's one thing you've seen work well in a prototype that became fragile once it reached a real environment? I'd like to hear it.
Mark Holland builds production AI systems for regulated healthcare, with a focus on testing, governance, and making AI dependable after go-live.