Mark HollandSenior AI Solutions Engineer

Application Testing Engine

ProductionBuilt at work for a national healthcare data company

QA is not the builder grading their own work. The Testing Engine runs on its own whenever an app lands in Azure DevOps and on every pull request, uses the app the way a person would, and tells the team what it could not check as well as what failed.

A real Testing Engine run against the Prompt Library on local synthetic data: 3 must fix, 21 to investigate, 91 checks passed, 9 failed, 40 pages crawled, 111 routes and 54 endpoints found, with what has to be fixed first.

What it does

  • Runs automatically when an app is loaded into Azure DevOps and on every pull request: static analysis, AI code audit, unit, API, accessibility, performance, and security tests, plus Playwright runs that click every link and button and use the app like a human, then sends the team a detailed report.
  • Reads acceptance criteria from Azure DevOps work items and reports each as passing, failing, regressed, not testable, or not covered.

From my resume.

Who it is for

The AI Solutions team and every developer who opens a pull request on its apps.

How it works

Application Testing Engine, step by step
  1. Trigger

    An app is loaded into Azure DevOps, or a pull request opens.

  2. Audit

    Static analysis and an AI code audit read the code.

  3. Test

    Unit, API, accessibility, performance, and security tests run.

  4. Use it

    Playwright clicks every link and button, like a person would.

  5. Criteria

    Each acceptance criterion reads as passing, failing, regressed, not testable, or not covered.

  6. Report

    The team gets a detailed report, with the must-fix items first.

The real screens

A real Testing Engine run against the Prompt Library on local synthetic data: 3 must fix, 21 to investigate, 91 checks passed, 9 failed, 40 pages crawled, 111 routes and 54 endpoints found, with what has to be fixed first.
A real run against the Prompt Library: a plain verdict first, then counts, then what has to be fixed. Shown with demo data.
Findings from the same run, ranked: must fix items such as a duplicate test name and an unused import, then items to investigate such as an unhandled error on an empty list and undocumented response fields, each with a category and a confidence percentage.
The same run's findings, ranked, each with its category and how sure the engine is. Shown with demo data.
Illustrative Testing Engine report for a sample claims portal: 3 must fix, 7 to investigate, 21 improvements, 112 checks passed, with the must-fix findings listed first.
  1. Ranked findings. Must-fix items come first, each with the phase that found it and how sure the engine is.

  2. Suggested fix. Each must-fix finding comes with a suggested fix.

Sample reportAn illustrative report. Findings are ranked must fix, needs investigation, and could be better, and every report lists what the run could not verify.

The decision that matters

Report what was not tested

A criterion the engine cannot observe is reported as not testable or not covered, never as passing. A green report that hides its gaps is worse than a red one.

Built with

  • Python
  • Azure DevOps
  • Playwright
  • pytest