From AI Idea to ProductionPart 5 of 20
Spec Before Code: Writing Down What the AI May and May Not Do
Why an AI project needs a written spec before any code, and what it should cover: data, limits, human review, testing, and cost. Part 5 of 20.

AI projects that go wrong usually go wrong in the first week, long before anyone writes a prompt. Someone sees a demo, and a developer opens an editor. A few weeks later there is a working prototype, and nobody can say whether it is finished, because nobody wrote down what finished means.
This is Part 5 of From AI Idea to Production: the document that comes before the code. For each approved app on my team, I run discovery calls with managers, directors, and business-unit teams. We map the workflow, the pain points, the data sources, and the success metrics. I turn what we learn into requirements, PRDs, and specs. Then the build starts.
Why AI work needs the spec more than ordinary software does
With a normal web form, the code mostly does what you tell it. If the spec is thin, you find the gap when a button does the wrong thing, and you fix it. A model behaves differently. It will produce a fluent answer to almost any input, including inputs you never planned for. It will guess when it lacks information. Its output shifts when you change the prompt, the model version, or the documents it reads. A thin spec on an AI project rarely fails loudly. It fails quietly, in an answer that sounds right.
So I write the limits down first. The spec is where the business decides what the AI may do and what it must never do. It also names who checks the work. Those decisions get much harder to change once users have seen the tool.
There is a second reason, and it is about people. In regulated healthcare, several people outside engineering usually need to agree before anything ships: the business owner, someone responsible for compliance, someone responsible for security. A written spec gives them one thing to read and sign. Without it, each of them approves a different picture of the tool, and you find out at launch.
What goes in an AI spec
A spec does not need to be long. It needs to answer eight questions, and I do not start building until each one has an answer. Sometimes the answer is "unknown, and here is how we will find out," and that counts.
| Question | What the spec says |
|---|---|
| The problem | What work happens today, who does it, and what it costs them in time or errors. |
| The users | Who will use the tool, who will read its output, and who is affected without ever opening it. |
| The data and its sensitivity | Every source the tool reads, who owns it, and whether it can contain patient or personal data. |
| Success metrics | The number that should move, with a baseline taken before launch. |
| What the AI must never do | The hard limits, written as rules a test can check. |
| Human review points | Where a person approves, edits, or rejects the AI's work. |
| How it will be tested | The checks that run before release, and what counts as a pass. |
| Cost limits | What a run, a user, or a month may cost, and what happens at the limit. |
Watch out
The first two look obvious, and they are where most specs are weakest. "Help the proposal team answer RFPs faster" is a goal. A problem statement says how long a response takes today, where the hours go, and which part of the work a tool could take over. Users need the same care: the person who reads the AI's output is often not the person who asked for the tool.
Data comes first
In healthcare I treat data sensitivity as the first design decision. Any free-text field might hold patient data, even when the form tells people not to type it there. So the spec names each source, says whether it can carry protected health information, and says what the system does if that data shows up. That answer shapes the architecture: which models you can call, what gets redacted before a prompt leaves, what gets logged.
Write the limits as rules a test can check
"The AI should be careful with figures" cannot be tested. "Any figure that disagrees with the figures register is flagged to the reviewer, and the tool never changes it" can. I write the never-do list in that second form. Each line names a behavior, the condition that triggers it, and what the system does instead. Written that way, the list is already half of the test plan.
The never-do rules for AI tools tend to cover familiar ground. Do not invent a value when an input is missing; ask a person. Do not answer without an approved source. Do not send sensitive data to an outside service. Do not take an action nobody approved. Do not let the model do arithmetic on scores or totals that code can compute.
Decide where people stay in the loop
I want at least one point in every AI tool where a person decides. The spec says where that point is, who the person is, and what they see when they decide. Placing it is a trade-off. Review on every answer adds time. No review on high-stakes output puts the company's name on a model's guess. I put review where a mistake would reach a customer or a patient record, and let low-risk output move on after automated checks.
From RFP Assist

Testing and cost belong in the spec
If the spec does not say how the tool will be tested, testing becomes whatever the builder had time for. I write the test plan as acceptance criteria tied to the requirements, and I include the hard cases: a missing input, sensitive data in the wrong field, a model that returns malformed output, a question no approved source covers. I also write down how the report should treat a criterion we cannot test yet. It should say not testable. It should never say passing.
Cost is the item teams forget. A model call is cheap until it runs on every record, every night. The spec should say what the tool may spend and what happens at the limit: stop, warn, or drop to a cheaper model. In my own AI operations platform, a pay-per-use model needs an explicit opt-in and a cost cap before it can run at all. Routing decisions belong here too. Routine questions can go to a cheaper model and hard or high-stakes ones to a stronger one, and the spec should say which is which.
From Command Center OS

From discovery to a spec people can sign
The spec comes out of discovery. I come out of those calls with workflow notes, the names of the data owners, a list of what people want the tool to do, and a shorter list of what they need it to do. The spec is where those two lists get separated.
Discovery
Calls with managers, directors, and business-unit teams map the workflow, pain points, data sources, and success metrics.
Data review
Each source is named with its owner, and any patient or personal data is flagged before design starts.
Limits
Never-do rules, human review points, and cost limits are written as statements a test can check.
Success and tests
Success metrics get a baseline, and acceptance criteria say what counts as a pass.
Scope
Scope in and scope out are written down. The out-of-scope list is never empty.
Sign-off
The business owner, compliance, and engineering approve the spec. Then the build starts.
Tip
Scope out matters as much as scope in. AI projects grow because a model can do so many adjacent things, and each one looks cheap to add. Writing down what the first version will leave out is the most reliable way I know to ship the first version.
Before you write the first line of code
0 of 11 done
Keep the spec current
A spec will change. Discovery misses things, and users teach you more in the first week of testing than in a month of meetings. When the facts change, I update the spec first and the code second, so the document stays the thing everyone can point to.
Next: building beyond the demo
A good spec gets you to a working first version. Production is a longer road. The distance between a demo that impresses a room and a tool a regulated company can run every day is where many AI projects stall, and Part 6 is about closing it.