# AI spec: (project name)

| | |
| --- | --- |
| Business owner | |
| Version | |
| Date | |
| Status | Draft |

> Answer every section. "Unknown, and here is how we will find out" counts as an answer.
> From the AI Spec Kit at markholland.tech/resources. By Mark Holland. CC BY 4.0: copy and adapt it, with credit.

## 1. The problem

<!-- What work happens today, who does it, and what it costs them in time or errors. A goal is not a problem statement: say how long the work takes today, where the hours go, and which part a tool could take over. -->

## 2. The users

<!-- Who will use the tool, who will read its output, and who is affected without ever opening it. -->

## 3. The data and its sensitivity

<!-- Every source the tool reads, who owns it, and whether it can contain personal or regulated data. Any free-text field might hold sensitive data, even when the form tells people not to type it there. -->

| Source | Owner | Can it hold personal or regulated data? | What the system does if that data shows up |
| --- | --- | --- | --- |
| | | | |

## 4. Success metrics

<!-- The number that should move, with a baseline taken before launch. Add a guard metric that must not get worse. -->

| Metric | Baseline today | Target | How it is measured after launch |
| --- | --- | --- | --- |
| | | | |

## 5. What the AI must never do

<!-- Write each rule so a test can check it: the behavior, the condition that triggers it, and what the system does instead. -->

| The AI must never | When | Instead, the system |
| --- | --- | --- |
| Invent a value | A required input is missing | Asks a person for it |
| | | |

## 6. Human review points

<!-- Where a person approves, edits, or rejects the AI's work. Put review where a mistake would reach a customer or a sensitive record. -->

| Where in the workflow | Who reviews | What they see when they decide |
| --- | --- | --- |
| | | |

## 7. How it will be tested

<!-- The checks that run before release, and what counts as a pass. Include the hard cases: a missing input, sensitive data in the wrong field, malformed model output, a question no approved source covers. A criterion you cannot test yet is reported as not testable, never as passing. -->

## 8. Cost limits

| Per run | Per user | Per month | At the limit (stop, warn, or cheaper model) |
| --- | --- | --- | --- |
| | | | |

## Scope

**In the first version**

-

**Left out of the first version** (this list is never empty)

-

## Sign-off

| Role | Name | Date |
| --- | --- | --- |
| Business owner | | |
| Compliance | | |
| Engineering | | |

## Before the first line of code

- [ ] The problem is written in terms of today's work and what it costs in time or errors.
- [ ] Every user, and every reader of the AI's output, is named.
- [ ] Every data source is listed with its owner and whether it can contain personal or regulated data.
- [ ] There is a plan for sensitive data that shows up where it should not.
- [ ] Each success metric has a baseline and a way to measure it after launch.
- [ ] The never-do list is written as rules a test can check.
- [ ] Each human review point names who reviews and what they see.
- [ ] The test plan covers missing inputs, malformed model output, and questions no source answers.
- [ ] Cost limits are set, with a stated behavior at the limit.
- [ ] Scope out is written down, and the list is not empty.
- [ ] The business owner, compliance, and engineering have approved the spec.
