Mark HollandSenior AI Solutions Engineer

How I Build

What I test for

OWASP's Top 10 for LLM applications lists the ways AI apps tend to fail, and three of them shape how I build for regulated healthcare. For each one, here is the risk in plain words and the controls that exist in my apps today, each shown with its app's real stage.
  • 3risks covered
  • 17guards in place
  • 7apps protected

Coverage map

Which app has a guard for which risk. Each dot links to that guard.
AppHidden instructionsActing without askingPrivate data leaks
Command Center OS23
RFP AssistNoneNone
Use Case Platform2
Platform MCPNone
LawLynxNoneNone2
PresidioFlowNoneNone
Prompt LibraryNoneNone

Risk 1

Hidden instructions that trick the AI

Prompt injection is text that tells a model to ignore its rules, hidden in an email, an uploaded file, or a web page the model reads. The model cannot reliably tell its owner's instructions from an attacker's, and OWASP lists prompt injection first in its Top 10 for LLM applications. Since no filter catches every hidden instruction, the controls below limit what a tricked model can do.

  1. Hidden instruction

    Text in an email, a file, or a web page tells the model to ignore its rules.

  2. The model reads it

    It cannot reliably tell its owner's instructions from an attacker's.

  3. Where the guards step in: A person or a check

    Depending on the app: drafts and plans wait for approval, answers come only from approved sources, or code checks the model's reply.

  4. The action is held

    The hidden instruction meets a person or a check before an action goes out.

How hidden instructions work, and where the guards in my apps step in.

Guards in my apps

  • An approval inbox card with a drafted email reply, its passed checks, and a panel of what the drafter knew.Demo data

    Command Center OSIn development

    Emails, posts, and memory edits all arrive as drafts, and nothing sends until I approve it.

  • An agent wizard showing its plan and waiting for approval before it runs.Demo data

    Command Center OSIn development

    Agents ask a few questions, show the plan, and wait; nothing runs until I approve the plan.

  • RFP Assist review screen: a question about provider reporting levels, with four facts above the draft: closest match 0.74, source age Current, drawn from 1 entry, fact checks Clear. The answer was already reviewed and accepted.Shown with demo data.

    RFP AssistProduction

    RFP Assist drafts answers only from the company's approved answer library, and a reviewer accepts, edits, or rejects every answer.

  • The dashboard: tiles counting ideas by outcome and a ranked queue of demo use cases with their scores.Shown with demo data.

    AI Use Case Submission PlatformProduction

    The model rates each dimension with a reason, the app computes the score and the outcome, and a reply with unexpected fields or unreadable values falls back to the rule engine and is flagged as a fallback.

Risk 2

AI that acts without asking

Excessive agency means a model has more power than its job needs: it can send email, change records, or run code, and it acts on a bad instruction or its own mistake. OWASP lists it in the same Top 10. The fix is fewer permissions, and a person in front of any action that leaves the system.

  1. Too much power

    The model can send email, change records, or run code.

  2. A bad instruction or a mistake

    It acts on an attacker's words or on its own error.

  3. Where the guards step in: Fewer permissions, a person at the door

    Depending on the app: external actions wait as drafts, tools sit under permission levels, or production changes need a person.

  4. The action waits

    No external send or production change goes out until someone approves it.

How an AI ends up acting without asking, and where the guards in my apps step in.

Guards in my apps

  • An approval inbox card with a drafted email reply, its passed checks, and a panel of what the drafter knew.Demo data

    Command Center OSIn development

    Command Center OS holds every external action as a draft with automated checks until I approve it, and the trust ramp can never remove the gate on external sends.

  • A trust promotion card proposing verified-auto for one skill, with evidence: 18 consecutive approvals, no rejections, no edits.Demo data

    Command Center OSIn development

    After a streak of clean approvals the system can only draft a card proposing more autonomy for a skill, with the evidence; I decide, and any rejection, or three edits in ten runs, puts the skill back on drafts.

  • Mission Control board with sessions grouped by business, each tagged with its engine, and drafts waiting for approval on the right.Demo data

    Command Center OSIn development

    The same core backs an edition that can never run code, and that boundary has its own tests.

  • Enterprise Platform MCPIn development

    Enterprise Platform MCP puts its 13 tools under three permission levels (read, write, and promote), and nothing is promoted to production without a person approving it.

  • A workflow board with columns for new, in review, decided, and on hold, holding demo idea cards.Shown with demo data.

    AI Use Case Submission PlatformProduction

    Build tickets stay drafts, scaffold notes are never applied to a code repository, and the monthly update is copied and posted by a person.

Risk 3

Private data that slips out

Sensitive information disclosure, in OWASP's words, is private data leaving the place it belongs: patient details sent to an outside model, written into a log, or repeated in an answer. In healthcare that data is often PHI, and once it reaches an outside service it cannot be pulled back.

  1. Private data comes in

    A patient note, a name, or a record number lands in a prompt.

  2. Where the guards step in: A gate before the model

    Identifiers are redacted or tokenized, or the item is held for a person to review.

  3. The model gets placeholders

    It works on text with the private details taken out, or never sees the item at all.

  4. No gate, no request

    In LawLynx, if the gate cannot run in production, the request stops.

How private data slips out, and where the guards in my apps step in.

Guards in my apps

  • LawLynx AI Legal Assistant: the chat workspace where a lawyer asks a question about a matter.

    LawLynxIn development

    A PII gate redacts all 18 HIPAA identifiers before content reaches an external model, and if the gate cannot run in production, the request stops.

  • LawLynx AI Legal Assistant: the chat workspace where a lawyer asks a question about a matter.

    LawLynxIn development

    LawLynx also has an egress guard with its own enforcement tests.

  • PresidioFlow audit snapshot check: chain linkage, per-record re-hash, and integrity checksum all pass for 341 entries; the signature is marked not checked.Test data

    PresidioFlowPrototype

    PHI is tokenized before remediation, so the LLM repairs structure it cannot read, and a test_no_phi_egress build gate fails if PHI could leave.

  • The intake form on its first step, with a five-step progress bar and a completeness meter.Shown with demo data.

    AI Use Case Submission PlatformProduction

    Ideas that touch patient or personal data are held for compliance review, never scored, and never sent to a model.

  • The intake form on its first step, with a five-step progress bar and a completeness meter.Shown with demo data.

    AI Use Case Submission PlatformProduction

    A redactor masks names, emails, birth dates, and record-number-like tokens before any model prompt and before logs are written; the code describes it as a heuristic screen, not a guarantee.

  • Prompt Library submission held for compliance review: the gate lists a name with a date of birth, a date of birth, an SSN-like value and a member ID, and the draft is kept below with its title and body.Shown with demo data.

    Prompt, Skill & Agent LibraryProduction

    A submission that looks like PHI or PII is held for review with the draft kept, and a person clears, overrides, or keeps each hold, with every override recorded.

  • Enterprise Platform MCPIn development

    The shared core applies the same policy and PHI guard to every request, whichever of the four adapters it comes through.

  • Engine lanes panel listing Claude Max, ChatGPT, Gemini, Kiro, OpenRouter, Ollama, OpenCode, and Cursor with subscription or metered badges.Demo data

    Command Center OSIn development

    Credentials sit in an encrypted secrets proxy that agents never see.

Illustrative Testing Engine report for a sample claims portal: 3 must fix, 7 to investigate, 21 improvements, 112 checks passed, with the must-fix findings listed first.
Sample reportAn illustrative report. Findings are ranked must fix, needs investigation, and could be better, and every report lists what the run could not verify.

How I check the apps

Application Testing Engine

Production

The Application Testing Engine I built at work runs static analysis, an AI code audit, unit, API, accessibility, performance, and security tests on the AI Solutions team's apps. Its AI code audit reads the code and flags places where user input can redirect the model, the opening for hidden instructions, and places where prompts or patient data can leak into logs. It also flags unchecked cost, unhandled failures, and AI calls that cannot be traced. It does not test what an agent is allowed to do while it runs.

Each guard above comes from an app on this site, shown at its real stage. If you want to talk through how one of them works, write to me.

How I buildContact