Mark HollandSenior AI Solutions Engineer

Command Center OS

In developmentMy own system. Runs locally today; moves to a VPS when finished

One cockpit that runs Claude Code, Codex, and Gemini side by side, routes each job to the cheapest engine that can do it, and never sends anything without my approval.

Command Center OS home in its light theme: a launch composer with project, engine buttons for Auto, Claude, and GPT, a prompt box, Ask the Brain, inbox triage, and a list of live sessions with drafts to review.
commits since June 2026
1,200+
automated tests
2,700+
engine lanes
8
model families and tiers
4 and 10

Counted from the repository, as of September 25, 2026.

On the resume

My own AI operations platform: a harness of harnesses that runs Claude Code, Codex, and Gemini sessions side by side, routes each job to the cheapest capable model under per-role limits, keeps a Git-backed memory with human-approved edits, and puts every external action behind an approval gate. Runs locally today; VPS deployment and team use are next.

The problem

Eight to ten editor windows, and no shared memory

Every project ran in its own editor window with its own agent. Prompts went to the wrong project, each agent started from nothing, and nothing coordinated them against one plan. I wanted one screen that runs every project as a live agent session, remembers what happened, and asks me before anything leaves.

The same cockpit on its Living Sky theme, a coastline photo behind the launch composer, Ask the Brain, inbox triage, and live sessions.
Demo dataOne screen for every project: launch a session, ask the Brain, triage the inbox, and watch live sessions. The optional Living Sky theme follows the real time of day (background photo from Unsplash).

Harness of harnesses

Every coding agent, one control surface

Claude Code, Codex, and Gemini CLI run as real sessions today, each as an interchangeable lane behind one approval gate and one memory. OpenCode and Cursor are supported lanes that switch on once their credentials exist.

  • Launch

    Pick a project, an engine, optional skills or MCP servers, and a first prompt. The session opens in the project's own folder with that project's instructions.

  • Isolate

    Each run is a Room with its own engine, model, and permissions. Rooms never share a prompt window; they coordinate through a shared ledger.

  • Monitor

    An event bus writes every session's output and state to the database. A classifier turns each terminal into a status dot (running, needs input, ready, idle, or error), and a watchdog stops runs that stall.

  • Steer

    Watch or type into any session from the browser, or inject a prompt, pause, kill, or switch engines mid-run. A switch passes the same billing and safety checks as the first route.

  • Collect

    When a Claude Code run finishes, its answer is read from Claude Code's own transcript and filed as a draft in the approval inbox.

You

Cockpit (Next.js)

  • Launch composer
  • Mission Control
  • Inbox
  • Brain

Brain API (FastAPI, 127.0.0.1)

  • Router and billing guard
  • Model policy by role and business
  • Session manager
  • Event bus to the database
  • Approvals
  • Memory

CLI sessions in tmux

  • Claude Code
  • Codex
  • Gemini CLI
  • OpenCodeneeds credentials
  • Cursorneeds credentials

API lanes, no terminal

  • OpenRoutermetered, capped
  • Ollamafree
  • Output and state flow back through the event bus to the status dots
  • A finished Claude Code transcript becomes a draft in the inbox
  • Memory is injected into every session at start
  • No metered Anthropic key
  • Secrets proxy
  • Per-role model policy
DiagramA harness of harnesses. Solid parts are built; dashed parts are off or designed.
Engine lanes panel listing Claude Max, ChatGPT, Gemini, Kiro, OpenRouter, Ollama, OpenCode, and Cursor with subscription or metered badges.
Demo dataEight engine lanes. Subscription lanes come first; metered lanes are opt-in and capped; the app refuses to start if a metered Anthropic key appears.
A live Claude Code terminal session shown inside the cockpit.
Test dataA real Claude Code session running inside the cockpit, attachable from the browser or raw tmux.
Mission Control board with sessions grouped by business, each tagged with its engine, and drafts waiting for approval on the right.
Demo dataSessions across businesses, each on its own engine, with drafts waiting on the right.

Mixture of models

The cheapest engine that can do the job

Every job goes to the cheapest engine that can do it. Admins limit which model families and tiers each role may use, and every run records why it went where it did.

  • Cheapest capable first

    Only engines with the capability a job needs are candidates, sorted by cost: the free Ollama lane, then flat-rate subscriptions, then capped pay-per-use lanes. A metered Anthropic key is blocked for my own work.

  • Four families, ten tiers

    Claude, GPT, Gemini, and an Ollama family each have a tier ladder, cheapest first, with a default reasoning effort per tier. The Claude ladder runs from Haiku to a Fable-class top tier.

  • Limits per role and business

    Per-role allow-lists for families, models, and tiers are checked on every dispatch, with per-business overrides. A denied request fails closed with a 403 and an audit entry.

  • Receipts

    Each run stores a routed-because receipt: family, tier, effort, and the reason.

  • Escalation Roadmap

    Designed so that a complexity check picks the tier and effort for each request, starting at Haiku and escalating through Sonnet and Opus to Fable when a cheaper tier fails its checks, with a fallback family during an outage.

Email replies draft on the Haiku tier, at a third of Sonnet's token price, inside a monthly budget per business that warns at 80% and stops at the cap.

  1. Capability

    Generation, reasoning, or code execution? Engines without it drop out.

  2. Fast path

    Triage, classify, label, summarize, or extract go to cheap routing lanes.

  3. Cost order

    Free Ollama lane, then flat subscription, then flat headless, then metered (opt-in and capped). Metered Anthropic is blocked.

  4. Policy

    Role allow-list for family, model, and tier; business overrides; the stricter cap wins. A denial returns 403 and an audit entry.

  5. Receipt

    routed: family/tier · effort, and the reason.

Worked examples

  • “Triage these 40 emails and label them.”

    Fast path, cheapest lane first

    Ollama lane, $0. Claude Code on the subscription if Ollama is down.

  • “Fix the failing test.”

    Needs code execution

    Claude Code on the flat subscription.

  • “Reply to a customer email.”

    Drafting policy per message type

    Haiku tier, inside a $10 monthly business cap.

  • Roadmap

    Complexity rises or checks fail: escalate Haiku, Sonnet, Opus, Fable.

DiagramHow a job finds its model. Solid parts are built; dashed parts are off or designed.
Model policy settings: tier ladders for the Claude, GPT, Gemini, and local families, cheapest first, with a reasoning effort per tier and primary and fallback families.
Demo dataTier ladders per family, cheapest first, up to the top Fable tier. Per-business overrides and per-role limits are enforced on every dispatch.
A session transcript headed by its routing receipt, routed claude/default with the operator pin, and the injected memory the session booted with.
Test dataEvery run records why it ran where it did, and every new session boots with the goal, memory index, learned rules, and recent journal. The goal text is hidden here.

Mixture of agents

Agents that plan, debate, and wait

Several agent patterns run through the cockpit, and each one stops for a person before anything consequential happens.

  • Guided agents

    An agent asks a few intake questions, shows its plan, and waits. Nothing runs until I approve the plan, and the wizard never sends.

  • Orchestrator skills

    A skill can chain other skills, such as inbox triage feeding the daily brief.

  • Debate

    Two Rooms can argue a plan through questions and answers on a shared task ledger, never a shared prompt, and hand back one reconciled note.

  • Orchestrator Feed

    When a Claude Code session spawns its own subagents, the cockpit shows each one live.

  • Cross-project bar

    Type a plain-English launch and it finds the right project and engine, or answers a question from the run history.

  • Orchestrated swarms Roadmap

    Designed so that a Fable-class orchestrator breaks a prompt into parts and spawns cheaper subagents, such as OpenAI GPT models or Claude Sonnet, reporting into one central feed.

An agent wizard showing its plan and waiting for approval before it runs.
Demo dataAgents ask a few questions, show the plan, and wait. Nothing runs until I approve the plan.

The approval gate

Nothing leaves without a human

Emails, posts, and memory edits all arrive as drafts. Each skill runs three to five automated checks before I see a draft, retries once if a check fails, and then waits for me to approve, edit, reject, or regenerate it.

  • One inbox

    Every approval lands in one keyboard-driven inbox, grouped by business, with a panel that shows what the drafter knew.

  • One winner

    Each approval can be decided only once, so a double click cannot send an email twice.

  • Earned trust

    After a streak of clean approvals, the system drafts a card proposing that a skill move to verified-auto, with the evidence. If I approve it, low-risk internal drafts from that skill auto-approve, each with an audit entry.

  • Automatic demotion

    Any rejection, or three edits in ten runs, puts the skill back on drafts.

  • External sends stay manual

    The trust ramp can never remove the gate on external sends.

External sends always need a human.

  1. Draft

    A skill drafts the email, post, or memory edit.

  2. Checks

    Three to five automated checks per skill.

  3. Self-heal

    One retry if a check fails.

  4. Human

    Approve, edit, reject, or regenerate.

  5. Act

    Send, publish, or commit.

  6. Audit

    Every decision is logged.

Earned trust

  1. A streak of clean approvals
  2. A promotion card with the evidence
  3. I approve the promotion
  4. Low-risk internal drafts auto-approve, with an audit entry

Any rejection, or three edits in ten runs, demotes the skill automatically.

DiagramThe approval gate and earned trust. Solid parts are built; dashed parts are off or designed.
An approval inbox card with a drafted email reply, its passed checks, and a panel of what the drafter knew.
Demo dataEach draft shows what the drafter knew and which checks passed.
A trust promotion card proposing verified-auto for one skill, with evidence: 18 consecutive approvals, no rejections, no edits.
Demo dataAutonomy is earned with evidence and granted by a human. External sends always stay manual.

The memory system

Memory that shows its work

The system of record is a Git wiki plus the operational database. Every new session starts from it, and it changes only through edits I approve.

  • Session start

    Every new session boots with the goal, the memory index, learned rules, and the last ten journal entries, capped at 12,000 characters.

  • Approved edits

    A memory review arrives as a diff I can edit. Approving it commits the change with the approver in the commit trailers; rejecting it changes nothing.

  • Procedures from real runs

    A finished run becomes a standard operating procedure after approval, linked to the run that produced it and marked fresh or stale.

  • Time machine

    Any wiki page can be viewed as it was on a past date, and each fact links to the log it was learned from.

NextNightly consolidation, semantic recall, and prediction tracking are built and tested; switching them on is next.

A run happens; its events and transcript are saved

Procedure

  1. SOP mined from the run
  2. Approved by me
  3. Filed in the wiki, fresh or stale

Memory edit

  1. Review proposes a diff
  2. I edit and approve
  3. Git commit with the approver

Journal

  1. System events logged
  2. Progress, roadblocks, decisions
  3. Last ten ride into the next session

The next session starts with the goal, memory index, learned rules, and recent journal

The Brain shows every page, where each fact came from, and any page on a past date

Built and tested, switching on next

  • Nightly consolidation
  • Semantic recall
  • Learned-rule capture
  • Prediction tracking
DiagramHow memory is written and read. Solid parts are built; dashed parts are off or designed.
The Brain tab showing a wiki page as it was on a past date, with a time machine date picker and a memory activity rail linking each fact to its source.
Demo dataSee any page as it was on a past date; every fact links to where it was learned.
A memory change draft shown as a diff with approve, edit, and reject controls.
Demo dataMemory edits arrive as diffs I can edit before they are committed.
A standard operating procedure generated from a real run, marked fresh, with a re-verify button.
Test dataReal runs become SOPs, marked fresh or stale against the latest run.

Operating the businesses

The businesses run from the same screen

The cockpit also runs the work around the code: what each business does, which parts a skill can take, and which stay with a person.

  • Work Map

    Every task a business does, marked automate (runs a skill) or keep-human, with gaps flagged as skills to build.

  • Connectors

    Gmail is live, with read access and approved sends. GoHighLevel reads and signed webhooks are live. Credentials sit in an encrypted secrets proxy that agents never see.

  • Inbox triage

    Scheduled Gmail triage classifies mail and drafts replies into the approval inbox.

  • Command palette

    One search across skills, MCP servers, plugins, actions, and sessions. The registry finds installed skills and MCP servers and turns them into buttons.

Work Map listing a business's tasks by category, each marked AUTOMATE or KEEP-HUMAN, with gaps flagged to build a skill.
Demo dataEvery task a business does, marked automate or keep-human, with gaps flagged.

Every day

A day in the cockpit

The routine parts run every day: a brief for each business, a sorted inbox with replies already drafted, a plain sentence when a number moves, and an optional background that follows the time of day.

  • Daily brief

    One brief per business, and one across all of them, in three parts: what needs me, what is worth knowing, and what is steady. It is filled in from the day's records by a template, with no model writing it, so it cannot invent a number.

  • Gmail triage

    Every hour, triage reads the inbox read-only, sorts mail into needs a reply, urgent, and newsletters or FYI, and drafts replies that wait in the approval inbox. Nothing sends until I approve it.

  • Watchtower

    Watchtower remembers the normal level of each number it tracks. When one drifts, it says so on Home in a plain sentence, such as a count running well below its 30-reading average.

  • Living Sky

    An optional background whose sun, moon, and horizon follow the real time of day where you are, from dawn through midday and golden hour to night. There are seven scenes to pick from, and choosing Light or Dark turns it off.

  • Break nudge

    After about 90 minutes of continuous work, a gentle nudge suggests stepping away.

The cockpit on the Coast scene at midday: blue sky, surf, and gulls behind the same page.

Demo dataThe same Coast scene through the day. Dawn and golden hour relight the midday photo; night has its own. Photos from Unsplash, under the Unsplash License.
The Portfolio panel on Home with three demo businesses and the brief across all of them: Northwind Studio has 4 items waiting for approval; the other two have no approval history yet.
Demo dataThe brief across all businesses, one line each, rendered from the day's records.
The Inbox triage card on the app's demo mailboxes: 6 messages pulled, 4 that need attention marked needs reply or urgent, newsletters and FYI hidden, and 4 drafts waiting.
Demo dataTriage on the app's demo mailboxes: mail sorted into needs a reply, urgent, and newsletters or FYI, with four drafts waiting for approval.
Appearance settings: Auto (Living Sky), Light, and Dark; a scene grid with Mountains, Coast, City, Rolling Hills, Golden Gate, New York City, and Washington DC; and the break nudge set to 90 minutes.
Demo dataSeven scenes to pick from, or Light or Dark to turn the sky off. The break nudge waits 90 minutes by default. Scene photos from Unsplash, under the Unsplash License.

How it is built

Built like it matters

Python with FastAPI, async SQLAlchemy, and Alembic on PostgreSQL; a Next.js and React cockpit over WebSockets; CLI sessions driven through tmux on WSL.

  • Loopback only

    The API answers only on this machine, blocks any website it does not know, and gives each live terminal its own signed ticket.

  • Secrets

    Credentials are encrypted at rest and read-only by default; write access is an explicit, audited grant. A secret scan covers every Git ref.

  • A hard boundary

    The same core also backs an edition that can never run code. Code execution is a capability the operator build grants itself, and the boundary has its own tests.

  • Tested and audited

    More than 2,700 automated tests. An external audit raised 28 findings, and the first remediation phase merged on September 25, 2026.

  • Sign-in Roadmap

    Designed so that each person signs in, through single sign-on for a team, before it moves to a VPS.

Built with

  • Python and FastAPI
  • Async SQLAlchemy and Alembic
  • PostgreSQL
  • Next.js and React
  • WebSockets
  • tmux sessions
  • Git-backed memory

How I build it

Command Center OS is built the way it is meant to run: models directing models, with a second vendor checking the risky parts.

How I build it
  1. Orchestrate

    A frontier-model orchestrator (Opus or Fable class) plans each pass and directs the work.

  2. Implement

    Sonnet implementation agents write the code, and verification agents check it against the plan.

  3. Review

    Codex reviews anything that touches approvals, credentials, or auth, adversarially. Ninety-five commits reference Codex review rounds.