Mark HollandSenior AI Solutions Engineer

From AI Idea to ProductionPart 5 of 20

Spec Before Code: Writing Down What the AI May and May Not Do

Why an AI project needs a written spec before any code, and what it should cover: data, limits, human review, testing, and cost. Part 5 of 20.

8 minute read

AI projects that go wrong usually go wrong in the first week, long before anyone writes a prompt. Someone sees a demo, and a developer opens an editor. A few weeks later there is a working prototype, and nobody can say whether it is finished, because nobody wrote down what finished means.

This is Part 5 of From AI Idea to Production: the document that comes before the code. For each approved app on my team, I run discovery calls with managers, directors, and business-unit teams. We map the workflow, the pain points, the data sources, and the success metrics. I turn what we learn into requirements, PRDs, and specs. Then the build starts.

An idea or ademoSomeone sees whatAI could do.DiscoveryCalls withmanagers,directors, andbusiness-unitteams.The specRequirements,PRDs, and specs,written down.The buildStarts once thespec is written.An idea or a demoSomeone sees what AI could do.DiscoveryCalls with managers,directors, and business-unitteams.The specRequirements, PRDs, and specs,written down.The buildStarts once the spec iswritten.
Where this part fits: the written spec sits between discovery and the first line of code.

Why AI work needs the spec more than ordinary software does

With a normal web form, the code mostly does what you tell it. If the spec is thin, you find the gap when a button does the wrong thing, and you fix it. A model behaves differently. It will produce a fluent answer to almost any input, including inputs you never planned for. It will guess when it lacks information. Its output shifts when you change the prompt, the model version, or the documents it reads. A thin spec on an AI project rarely fails loudly. It fails quietly, in an answer that sounds right.

So I write the limits down first. The spec is where the business decides what the AI may do and what it must never do. It also names who checks the work. Those decisions get much harder to change once users have seen the tool.

There is a second reason, and it is about people. In regulated healthcare, several people outside engineering usually need to agree before anything ships: the business owner, someone responsible for compliance, someone responsible for security. A written spec gives them one thing to read and sign. Without it, each of them approves a different picture of the tool, and you find out at launch.

Without a specSomeone sees a demo, and a developeropens an editor.A few weeks later there is a workingprototype.Each approver pictures a differentversion of the tool.Nobody can say whether it isfinished, and the gaps show up atlaunch.With a specThe limits are written down first: whatthe AI may do and must never do.The business owner, compliance, andsecurity read and sign one document.The person who checks the work is named.The hard decisions are made beforeusers have seen the tool.Without a specSomeone sees a demo, and adeveloper opens an editor.A few weeks later there is aworking prototype.Each approver pictures a differentversion of the tool.Nobody can say whether it isfinished, and the gaps showup at launch.With a specThe limits are written down first:what the AI may do and must neverdo.The business owner, compliance,and security read and sign onedocument.The person who checks the work isnamed.The hard decisions are madebefore users have seen thetool.
Building with a spec versus building without one.

What goes in an AI spec

A spec does not need to be long. It needs to answer eight questions, and I do not start building until each one has an answer. Sometimes the answer is "unknown, and here is how we will find out," and that counts.

The specThe problemThe usersData andsensitivitySuccessmetricsWhat the AImust never doHuman reviewpointsHow it will betestedCost limitsThe specThe problemThe usersData andsensitivitySuccessmetricsWhat the AImust neverdoHumanreviewpointsHow it willbe testedCost limits
Eight questions every AI spec answers before the build starts.
The eight questions, and what the spec says for each
QuestionWhat the spec says
The problemWhat work happens today, who does it, and what it costs them in time or errors.
The usersWho will use the tool, who will read its output, and who is affected without ever opening it.
The data and its sensitivityEvery source the tool reads, who owns it, and whether it can contain patient or personal data.
Success metricsThe number that should move, with a baseline taken before launch.
What the AI must never doThe hard limits, written as rules a test can check.
Human review pointsWhere a person approves, edits, or rejects the AI's work.
How it will be testedThe checks that run before release, and what counts as a pass.
Cost limitsWhat a run, a user, or a month may cost, and what happens at the limit.

Watch out

The first two look obvious, and they are where most specs are weakest. "Help the proposal team answer RFPs faster" is a goal. A problem statement says how long a response takes today, where the hours go, and which part of the work a tool could take over. Users need the same care: the person who reads the AI's output is often not the person who asked for the tool.

Data comes first

In healthcare I treat data sensitivity as the first design decision. Any free-text field might hold patient data, even when the form tells people not to type it there. So the spec names each source, says whether it can carry protected health information, and says what the system does if that data shows up. That answer shapes the architecture: which models you can call, what gets redacted before a prompt leaves, what gets logged.

Each datasourceNamed in the spec,with its owner.Can it carrypatient orpersonal data?Yes: it shapes the designWhich models you can callWhat gets redacted before aprompt leavesWhat gets loggedNo: still plan for itAny free-text field might holdpatient data anywayThe spec says what the systemdoes if it shows upEach datasourceNamed in the spec, with itsowner.Can it carrypatient orpersonal data?Yes: it shapes the designWhich models you can callWhat gets redacted before aprompt leavesWhat gets loggedNo: still plan for itAny free-text field mighthold patient data anywayThe spec says what thesystem does if it shows up
Data sensitivity is the first design decision. The answer shapes the architecture.

Write the limits as rules a test can check

"The AI should be careful with figures" cannot be tested. "Any figure that disagrees with the figures register is flagged to the reviewer, and the tool never changes it" can. I write the never-do list in that second form. Each line names a behavior, the condition that triggers it, and what the system does instead. Written that way, the list is already half of the test plan.

Behavior+What the tool must not do.Change a figureTrigger+The condition that sets therule off.It disagrees with thefigures registerInsteadWhat the system does instead.Flag it to the reviewer andleave the number as it isExample from this partBehaviorWhat the tool must not do.Change a figure+TriggerThe condition that sets the rule off.It disagrees with the figuresregister+InsteadWhat the system does instead.Flag it to the reviewer and leavethe number as it is
A never-do rule names a behavior, the condition that triggers it, and what the system does instead. Written that way, it is half of a test.

The never-do rules for AI tools tend to cover familiar ground. Do not invent a value when an input is missing; ask a person. Do not answer without an approved source. Do not send sensitive data to an outside service. Do not take an action nobody approved. Do not let the model do arithmetic on scores or totals that code can compute.

Decide where people stay in the loop

I want at least one point in every AI tool where a person decides. The spec says where that point is, who the person is, and what they see when they decide. Placing it is a trade-off. Review on every answer adds time. No review on high-stakes output puts the company's name on a model's guess. I put review where a mistake would reach a customer or a patient record, and let low-risk output move on after automated checks.

Automated checksLow-risk output moves on after automatedchecks.A person decidesWhere a mistake would reach a customer or apatient record, a person approves, edits, orrejects the work.Low riskHigh stakesLow riskAutomated checksLow-risk output moves on afterautomated checks.A person decidesWhere a mistake would reach acustomer or a patient record, aperson approves, edits, orrejects the work.High stakes
Where review goes: a person decides wherever a mistake would reach a customer or a patient record.

From RFP Assist

RFP Assist review screen: a question about provider reporting levels, with four facts above the draft: closest match 0.74, source age Current, drawn from 1 entry, fact checks Clear. The answer was already reviewed and accepted.
Each draft shows how close its source match is, how old the source is, and whether a fact check flagged anything. A person decides every answer. Shown with demo data.

Testing and cost belong in the spec

If the spec does not say how the tool will be tested, testing becomes whatever the builder had time for. I write the test plan as acceptance criteria tied to the requirements, and I include the hard cases: a missing input, sensitive data in the wrong field, a model that returns malformed output, a question no approved source covers. I also write down how the report should treat a criterion we cannot test yet. It should say not testable. It should never say passing.

Cost is the item teams forget. A model call is cheap until it runs on every record, every night. The spec should say what the tool may spend and what happens at the limit: stop, warn, or drop to a cheaper model. In my own AI operations platform, a pay-per-use model needs an explicit opt-in and a cost cap before it can run at all. Routing decisions belong here too. Routine questions can go to a cheaper model and hard or high-stakes ones to a stronger one, and the spec should say which is which.

From Command Center OS

Engine lanes panel listing Claude Max, ChatGPT, Gemini, Kiro, OpenRouter, Ollama, OpenCode, and Cursor with subscription or metered badges.
Demo dataEight engine lanes. Subscription lanes come first; metered lanes are opt-in and capped; the app refuses to start if a metered Anthropic key appears.

From discovery to a spec people can sign

The spec comes out of discovery. I come out of those calls with workflow notes, the names of the data owners, a list of what people want the tool to do, and a shorter list of what they need it to do. The spec is where those two lists get separated.

  1. Discovery

    Calls with managers, directors, and business-unit teams map the workflow, pain points, data sources, and success metrics.

  2. Data review

    Each source is named with its owner, and any patient or personal data is flagged before design starts.

  3. Limits

    Never-do rules, human review points, and cost limits are written as statements a test can check.

  4. Success and tests

    Success metrics get a baseline, and acceptance criteria say what counts as a pass.

  5. Scope

    Scope in and scope out are written down. The out-of-scope list is never empty.

  6. Sign-off

    The business owner, compliance, and engineering approve the spec. Then the build starts.

How an AI idea becomes a spec a team can sign.

Tip

Scope out matters as much as scope in. AI projects grow because a model can do so many adjacent things, and each one looks cheap to add. Writing down what the first version will leave out is the most reliable way I know to ship the first version.

Before you write the first line of code

0 of 11 done

Keep the spec current

A spec will change. Discovery misses things, and users teach you more in the first week of testing than in a month of meetings. When the facts change, I update the spec first and the code second, so the document stays the thing everyone can point to.

Next: building beyond the demo

A good spec gets you to a working first version. Production is a longer road. The distance between a demo that impresses a room and a tool a regulated company can run every day is where many AI projects stall, and Part 6 is about closing it.