Mark HollandSenior AI Solutions Engineer

How an idea becomes a working AI app

  1. Idea
  2. Spec
  3. Build
  4. Test
  5. Ship
  6. Improve

How I build

Keeping AI Honest

AI models hallucinate. They write an answer that sounds right and isn't: a refund policy that doesn't exist, a figure nobody measured, a court case that was never decided. You can't train that out of a model. So I build systems that stop a hallucination before it reaches anyone, and catch it fast when one gets through.

What I test for: the three AI risks I guard against

Five layers between a question and an answer

Every question passes through five layers before anyone sees an answer. If one layer misses a problem, the next one can still catch it.

Question

“Can a customer who cancels early get a full refund?”

  1. Layer 01

    Ground it

    Answers come from approved sources only

    A model is the AI program that writes the answer. Left alone, it answers from what it picked up in training, and that can be out of date or wrong. I connect it to a set of approved documents and have it answer only from those.

    • The business decides which documents count
    • The model looks up the answer before it writes
    • Change a document and the answers change with it

    Where I built this

    • RFP AssistRFP Assist drafts answers only from the company's approved answer library.
  2. Layer 02

    Cite it

    Every answer shows where it came from

    Each answer points back to the document it used, so a reader can check it quickly. If no approved source covers the question, there is no answer. The system says "I don't know" and hands the question to a person instead of guessing.

    • Each answer links to its source
    • No source means no answer
    • "I don't know" is an allowed reply
    • Unanswered questions go to a person

    “Can a customer who cancels early get a full refund?”

    If no approved document covers early cancellation, the reply is "I don't know. Sending this to a person." instead of a guessed refund policy.

    Where I built this

    • RFP AssistRFP Assist shows the library answers each draft came from and how closely they matched.
  3. Layer 03

    Check it

    Automatic checks and a second opinion

    Before an answer moves on, software checks it. A privacy gate stops personal details, such as names and record numbers, from reaching an outside AI service. On high-stakes questions, other AI models review the answer as well, so one model's mistake doesn't go out on its own.

    • Built-in checks flag answers that break the rules
    • A privacy gate keeps personal data out
    • Several models review high-stakes answers
    • The code does the math, not the model
    • Answers must match a fixed structure or they don't move on

    Where I built this

    • LawLynxLawLynx sends high-stakes questions to a three-model review panel (Claude, GPT, Gemini), and its privacy gate blocks all 18 HIPAA personal identifiers.
    • PresidioFlowPresidioFlow swaps personal health data for tokens, stand-in codes that mean nothing on their own, and blocks data from leaving.
    • AI Use Case Submission PlatformWhen the AI Use Case Submission Platform scores an idea with a model, the model rates each dimension with a reason and a confidence value, and the app computes the weighted score and the outcome. If the model returns unexpected fields or unreadable values, the app falls back to its rule engine and flags the result as a fallback.
  4. Layer 04

    Approve it

    A person signs off on high-stakes work

    Some answers carry real weight, such as a price or a promise to a customer. For those, the AI writes the draft and a person decides whether it goes out. The AI saves the slow first pass, and the person keeps the final say.

    • High-stakes answers wait for a reviewer
    • Nothing high-stakes goes out on the AI's word alone
    • The reviewer sees the draft before anyone else does

    Where I built this

    • RFP AssistA reviewer approves every RFP Assist answer, and response time still fell from 3 to 4 months to 2 to 3 weeks.
  5. Layer 05

    Watch it

    Test every change, monitor what runs

    An app that worked last month can break after next week's change. So every code change goes through automated tests before release. Once an app is live, I monitor how it behaves and keep an audit trail of what happened.

    • Tests run on every code change, before release
    • Automated browser runs click through the app like a person
    • Monitoring shows what each live AI app is doing
    • An audit trail records each step

    Where I built this

    • Application Testing EngineThe Application Testing Engine runs on every code change, from static analysis and an AI code audit to automated browser runs that click every link and button.
    • Observability & Governance EngineThe Observability and Governance Engine, stage one built: every AI call is recorded with the data versions behind it, and in a later stage answers can be held for a person before they are seen.
    • PresidioFlowPresidioFlow keeps a tamper-evident audit ledger and produces signed evidence packs mapped to SOC 2, the HIPAA Security Rule, and ISO 27001.

Where hallucinations come from, and what stops each one

Each row is one way a model hallucinates, the defense that stops it, and the project where I built that defense.

  1. How a model gets it wrong
    It answers from memory instead of your documents
    What stops it
    Ground it: it may only use approved sources it looked up first.
    Where I built it
    • RFP AssistRFP Assist drafts only from the approved answer library.
  2. How a model gets it wrong
    It fills a gap with a confident guess
    What stops it
    No source, no answer. "I don't know" is an allowed reply, and a missing input produces a question for a person, never a made-up value.
    Where I built it
    • AI Use Case Submission PlatformWhen the inputs for a value estimate are missing, the AI Use Case Submission Platform writes "Requires stakeholder conversation to quantify" instead of inventing a number.
  3. How a model gets it wrong
    It slips a wrong number or claim into a fluent answer
    What stops it
    Every figure is checked against a register of confirmed figures and flagged where the reviewer reads it.
    Where I built it
    • RFP AssistRFP Assist's claims register flags a figure nobody has confirmed, and raises it again before export.
  4. How a model gets it wrong
    It gets the math or the rules wrong
    What stops it
    The model judges, the code computes. Scores, totals, and outcomes are calculated in code from the model's structured ratings.
    Where I built it
    • AI Use Case Submission PlatformWhen the AI Use Case Submission Platform scores with a model, the model rates each dimension with a reason and a confidence value. The app, not the model, computes the weighted score and the outcome.
  5. How a model gets it wrong
    It returns something malformed or off-script
    What stops it
    Every answer is checked against a fixed structure. If it fails, a rule-based fallback runs and the result is labeled as a fallback.
    Where I built it
    • AI Use Case Submission PlatformThe AI Use Case Submission Platform falls back to its rule engine and flags the result when the model returns unexpected fields or unreadable values.
  6. How a model gets it wrong
    One model has a blind spot
    What stops it
    Different models give a second opinion on high-stakes questions.
    Where I built it
    • LawLynxLawLynx sends high-stakes questions to a three-model panel (Claude, GPT, Gemini).
  7. How a model gets it wrong
    It's right today and wrong after next week's change
    What stops it
    The same checks rerun on every change, before release.
    Where I built it
  8. How a model gets it wrong
    A report quietly overstates what was checked
    What stops it
    Every report says what it could not verify.
    Where I built it
    • Application Testing EngineThe Application Testing Engine lists what each run could not check.
    • PresidioFlowPresidioFlow's audit check marks the snapshot's signature "not checked" instead of implying it passed.

How I know it's working

These are the numbers I use to tell whether the defenses are holding.

Grounded-answer rate
The share of answers that cite an approved source.
"I don't know" rate
How often the system declines to answer, and whether each decline was right.
Reviewer correction rate
How often a person changes the AI's answer. A rising rate is usually the first sign of trouble, before anyone files a complaint.
Flagged claims
Figures the claims check could not confirm.
Test set accuracy
A fixed set of questions with known answers, rerun on every prompt or model change.

Where this is going

  • Observability & Governance EngineStage one of the governance engine is built; the next stage adds checks in the request path once two weeks of real traffic have been measured.

How the pieces fit

Here is the same idea drawn as one system. Each box is a step a question passes through, and every step leaves a record.

How the pieces fit

Question

Someone asks in plain words

Privacy gate

Removes personal details before any model sees the text

Router

Sends easy questions to cheaper models, hard ones to stronger ones

Approved sources

The only documents the model may use

Model

The AI program that writes the draft

Cheaper model

Fast and low cost, for routine questions

Stronger model

Costs more, for hard or high-stakes questions

Checks

Rule checks and a second opinion on the draft

Human approval

A person signs off on high-stakes answers

Answer

Goes out with its source attached

Audit trail

A record of every step, kept for review

Monitoring

Watches live apps for errors, cost, and slowdowns

Trade-offs I weigh

  • Cost vs. accuracy

    The strongest models cost more per question. I send routine questions to cheaper models and save the strong ones for questions where a mistake would hurt.

  • Speed vs. review

    A human review adds time. I put it where the stakes justify the wait and let low-risk answers go out after the automatic checks.

  • Only the access it needs

    An agent is an AI program that can take actions, like reading files or sending email. I give each one only the access its job requires, so a mistake stays small.

You can't get hallucinations to zero. You can make sure one gets caught before it reaches a customer.

Three AI risks I build against: hidden instructions, AI that acts without asking, and private data that slips out. The guards in my apps, and their limits.

Read what i test for

Talk to me about your AI work

I can walk you through how these layers would fit your team's work. The project pages below show where I built each one, including the governance engine, whose first stage is built.

Email me