Mark HollandSenior AI Solutions Engineer

From AI Idea to ProductionPart 6 of 20

Building Beyond the Demo

A demo proves an AI idea can work. Production proves it keeps working. The six layers I add before I call an AI system production-ready.

12 minute read

A demo proves the idea can work. Production proves the system keeps working after real users, real data, real integrations, and real failures show up.

In Part 5 I wrote about creating the spec before writing code. This part covers the next problem. Once you have a spec and a prototype that works, how do you build the real thing so it survives contact with the real environment?

The most common mistake I see is taking a prototype that impressed everyone in a meeting and simply adding features to it. It feels like progress. In practice it piles up technical debt faster than business value, because everything the prototype skipped on purpose is still missing. It just has more weight sitting on top of it now.

This article walks through the six layers I add before I call an AI system production-ready. For each one you get what goes wrong without it, how to set it up, and something you can copy into your own project.

Infographic for Part 6, Building Beyond the Demo: five stages from develop to monitor, each with its practices, plus key practices, common pitfalls, the payoff, and green-dollar and blue-dollar impact.
The path from a working idea to a system built to last.

A demo that works perfectly

Let me use a made-up example, because it is a pattern almost everyone building with AI will recognize.

Picture a customer support assistant built for an online retailer. In the demo, a manager types "Where is my order?" The assistant looks up the order, explains the delay, and offers a refund on the shipping charge. It is fast, polite, and accurate. The room is sold.

Here is what that demo is running on.

  • One environment, which is the developer's laptop
  • One API key, pasted into a config file
  • One model, called directly from the code
  • One prompt, edited by hand whenever something looks off
  • Logs that print to the screen and disappear
  • A fake order system that always returns a clean answer
  • Twenty test questions the builder wrote

That is exactly the right setup for a demo. Its whole job is to answer one question, which is whether the idea is worth building. It is the wrong setup for something a business will depend on.

The demoPromptModelAPI keyFake order systemScreen logsOne laptopThe production systemEnvironmentsModulesVersionsSecretsLoggingFeature flagsThe demoPromptModelAPI keyFake order systemScreen logsOne laptopThe production systemEnvironmentsModulesVersionsSecretsLoggingFeature flags
Same idea. Different level of discipline.

What breaks first

Imagine that assistant goes live the way it was demoed. None of what follows is exotic. These are the ordinary ways a prototype fails in the real world, and each one is fixed by a specific layer.

Six ordinary failures, and the layer that prevents each one
What goes wrongWhat the customer or business seesThe layer that prevents it
Security rotates the API key pasted in the config fileThe assistant goes down for everyone at onceSecrets
A developer tweaks the prompt to sound friendlierIt starts offering refunds it used to escalate, and nobody can say which prompt was runningVersioning
The real order system times out at busy hoursThe assistant gives wrong or empty answersModularity and environments
Someone copies a real customer conversation into the test setPersonal data now lives somewhere it shouldn'tEnvironments
A customer says a promised refund never arrivedNobody can reconstruct what the assistant read, called, or decidedLogging
Automatic refunds misbehave on day oneThe only way to stop it is to take the whole assistant offlineFeature flags

None of these are solved by a better prompt. They are solved by structure.

Layer 1. Environments

Development, test, staging, and production should be separate, each with its own data, credentials, and settings. The rule I hold to is simple. Production data and production credentials never become convenient test fixtures.

The four environments side by side
DevelopmentTestStagingProduction
PurposeBuild and experimentAutomated tests on every changeRehearse the real releaseReal customers
DataSynthetic onlySynthetic and fixed test setsRepresentative, de-identifiedReal
CredentialsDev keys, low limitsTest keysStaging keysProduction keys, tightest access
IntegrationsMockedMocked plus contract testsReal systems in sandbox modeReal systems
Who has accessBuildersPipelineBuilders, QA, business ownersOperations, on-call
Logging detailEverythingEverythingProduction levelProduction level, personal data masked
DevelopTestStageReleaseMonitortests passend-to-end scenariosand sign-offflag on for a smallgroupmetrics healthywhat we learnedDeveloptests passTestend-to-end scenarios andsign-offStageflag on for a small groupReleasemetrics healthyMonitorwhat we learned
The release path: every change passes a check before the next stage, and what you learn in production goes back into development.

How to set it up

  1. Give each environment its own configuration, its own credentials, and its own data store.
  2. Make the code read which environment it is running in. Never hard-code a URL, key, or database name.
  3. Build a synthetic data set that has the same shape as real data, including the messy cases like missing fields, very long messages, and unusual names.
  4. Block production credentials from working anywhere outside production.

Tip

If you need realistic data to reproduce a bug, generate a synthetic copy of the problem. Never copy the real record.

Layer 2. Modularity

The prototype is usually one long piece of code where retrieval, model calls, business rules, permissions, and logging are all tangled together. Production needs them to be separate parts with clear boundaries.

LoggingRecords everystepUser requestRetrievalModelproposed actionplain code decidesBusiness rulesPermissionsapprovedToolsOrder systemRefund systemREJECTEDEscalate to apersonUser requestRetrievalModelproposed actionplain code decidesBusiness rulesPermissionsapprovedToolsOrder systemRefund systemREJECTEDEscalate to a personLoggingRecords every step
Model proposes, code decides.
Each component's one job, and what it must never do
ComponentIts one jobWhat it must never do
RetrievalFind the right documents and recordsDecide what the customer is allowed
Model callsTalk to the AI model through one shared doorwayHold credentials or call systems directly
Business rulesApply limits like "refunds over $50 need a person"Live inside a prompt
PermissionsDecide what this user and this assistant may doBe bypassed by clever wording
ToolsPerform the action, such as issuing a refundRun without a permission check
LoggingRecord what happened at every stepStore personal data it doesn't need

The test is whether you can swap one part without rewriting the others. If changing models means rewriting the whole workflow, it is too tangled. If a refund limit lives inside a prompt instead of in plain code, someone can talk their way around it.

That last point is one I build into everything. In my own builds, the model can suggest an action, but plain code decides whether the action is allowed.

How to set it up

  1. Write down the inputs and outputs for each component before you split them. This is the contract between them.
  2. Put every model call behind one function, so changing models touches one place.
  3. Move every limit and permission out of the prompt and into code that the model cannot edit.
  4. Write a test for each contract, so a change in one component that breaks another fails immediately.

Layer 3. Versioning

Everyone versions their code. With AI systems, code is only part of what changes behavior. You need to be able to answer "what exactly was running at 2:14 on Tuesday?" for all of it.

What to version, and why it changes behavior
What to versionWhy it changes behavior
PromptsOne sentence can change what the assistant will do
Model and its settingsA new model or temperature changes answers
Retrieval settingsWhich index, how many results, and which filters change what the model sees
Tool definitionsChanging a tool's inputs changes what actions are possible
Business rules and permissionsChanges what is allowed
Evaluation setChanges what "passing" means

How to set it up

Keep prompts and settings in version control next to the code, not pasted into a dashboard. Give every release a short manifest, and stamp the release ID into every log entry. When behavior changes, you can line it up with the exact release that caused it.

A release manifest can be this simple.

YAML
release: support-assistant-2026.10.07-3
code: 4f2a91c
prompt: support-system-prompt v12
model: provider-model-name, temperature 0.2
retrieval: help-center-index v5, top 6 results, policy filter on
tools: order_lookup v3, refund_request v2
rules: refund-limits v4 (person required over $50)
eval_set: support-golden-set v9 (412 cases, 97% pass)
approved_by: release owner name

Layer 4. Secrets

API keys, database passwords, and tokens belong in a secret manager that is separate for each environment. They never belong in source code, config files, shared documents, or prompts.

neverModelrequest: refundorder A-101Tool servicefetches the keySecret managercallsRefund systemModelrequest: refundorder A-101Tool servicecallsRefund systemSecret managerfetches the keynever
The model asks. The tool acts. The key stays locked away.

How to set it up

  1. Move every credential into a secret manager, with separate credentials for each environment.
  2. Scan your code history for keys that were ever committed, and rotate any you find. Deleting the line is not enough, because the key is still in the history.
  3. Let tools fetch credentials at the moment they need them. The model never sees one.
  4. Practice a rotation. Change a key in staging and confirm nothing breaks and nothing needs a code change.

In a system I'm building for myself, credentials sit behind a proxy the agents never see. The agent can ask for an action, but it can't read, copy, or leak the key that performs it.

The agent can ask for an action, but it can't read, copy, or leak the key that performs it.

Layer 5. Logging

Logging that prints to a screen is for debugging. Production logging has to let you reconstruct exactly what happened in a single request, weeks later, for someone who wasn't there.

request req-7f3c0 msreceived120 msretrieved 2help articles900 msmodel proposedrefund910 msrule checkpassed under$501.4 srefund toolcalled1.6 srefundconfirmed1.7 sreply sentrequest req-7f3c0 msreceived120 msretrieved 2 help articles900 msmodel proposed refund910 msrule check passed under $501.4 srefund tool called1.6 srefund confirmed1.7 sreply sent
If you can rebuild this from your logs, you have observability.
What to record for every request
What to recordWhy it matters
A request ID that follows the request everywhereTies every step of one conversation together
Who asked, plus the environment and release IDExplains context and links behavior to a release
What was retrievedShows whether the answer used the right information
Every tool call, with inputs and resultShows what the system did, not just what it said
The decision and the reasonShows why it acted, escalated, or refused
Errors, retries, and timeoutsShows where it struggled
Time taken and costShows whether it is fast enough and affordable
The final business outcomeShows whether the result actually happened

Here is what one masked log entry might look like.

JSON
{
  "request_id": "req-7f3c",
  "release": "support-assistant-2026.10.07-3",
  "env": "production",
  "user": "customer-****8812",
  "retrieved": ["help/shipping-delays", "help/refund-policy"],
  "proposed_action": "refund_shipping",
  "rule_check": "passed, amount 7.99 under limit 50",
  "tool_call": {"name": "refund_request", "status": "confirmed", "ms": 640},
  "outcome": "refund_issued",
  "latency_ms": 1710,
  "cost_usd": 0.004
}

Watch out

Mask what you don't need, and decide up front how long logs are kept. A log full of personal data is a liability, not an asset.

Layer 6. Feature flags

A feature flag is a switch that turns one capability on or off without a new deployment. It sounds small. It is one of the most valuable things you can add to an AI system that takes actions.

1Internal staffA week with nowrong refunds25 percent ofcustomersError rate andcomplaints matchthe old process325 percentCost per refund isstable andreviewers agreewith its decisions4EveryoneTwo weeks stable at25%Kill switchOff in seconds,assistant keeps helping4. EveryoneTwo weeks stable at 25%3. 25 percentCost per refund is stable andreviewers agree with its decisions2. 5 percent of customersError rate and complaints matchthe old process1. Internal staffA week with no wrong refundsKill switchOff in seconds, assistant keeps helping
Earn the next step, and keep a kill switch beside every one.
A feature flag is a switch that turns one capability on or off without a new deployment.
Moving through the rollout, and when to roll back
StageWho gets itMove to the next stage whenRoll back when
1Internal staffA week with no wrong refundsAny refund that breaks a rule
25% of customersError rate and complaints match the old processComplaints rise or a rule is broken
325% of customersCost per refund is stable and reviewers agree with its decisionsCost or error rate spikes
4EveryoneTwo weeks stable at 25%Same triggers, flag stays in place

Go back to the automatic refund example. Without a flag, a bad day means taking the whole assistant offline. With a flag, you switch off automatic refunds, the assistant goes back to sending refund requests to a person, and customers keep getting help with everything else while you find the problem.

A simple rollback plan, written before launch

  1. Who can flip the switch, and how to reach them after hours.
  2. The exact flag to turn off, and what the system does when it is off.
  3. How to confirm the "off" path is working.
  4. Who tells the business and customers, and what they say.
  5. How to find every action taken while the problem was live, using the logs.

Common pitfalls

  • Running the demo in production and calling it a pilot
  • Hard-coded keys anywhere in the code or config
  • Skipping staging because "it worked in test"
  • Logging that can't reconstruct a single request
  • Integrations that assume the other system always answers quickly and correctly
  • All-or-nothing releases with no way to turn off one feature
  • No named owner for the system after launch

Why this is worth the effort

Production engineering can look like overhead, because none of it shows up in a demo. It shows up in the results.

Green-dollar and blue-dollar results
Green-dollar results (show up in the budget)Blue-dollar results (operational value you measure on purpose)
Fewer incidents to respond toHigher reliability
Less rework after bad releasesFaster, safer delivery of new features
Faster troubleshootingMore team capacity, less firefighting
Lower cost when something does go wrongConfidence to scale to more users

The demo proves what's possible. The architecture decides whether that possibility turns into dependable business value.

Production readiness scorecard

Score each layer 0, 1, or 2. Anything under 2 tells you where to work next.

0 of 12

0 of 6 scored.

Environments
Modularity
Versioning
Secrets
Logging
Feature flags

Checklist before you call it production-ready

Tick each box you can say yes to.

0 of 14 done

Next in the series

Part 7 is about testing AI systems properly. That means going beyond "the answer looked right" to test retrieval quality, tool use, permissions, recovery, speed, cost, and whether the business outcome actually happened.

What's one thing you've seen work well in a prototype that became fragile once it reached a real environment? I'd like to hear it.

Mark Holland builds production AI systems for regulated healthcare, with a focus on testing, governance, and making AI dependable after go-live.