Command Center OS
In developmentMy own system. Runs locally today; moves to a VPS when finishedOne cockpit that runs Claude Code, Codex, and Gemini side by side, routes each job to the cheapest engine that can do it, and never sends anything without my approval.

- commits since June 2026
- 1,200+
- automated tests
- 2,700+
- engine lanes
- 8
- model families and tiers
- 4 and 10
Counted from the repository, as of September 25, 2026.
On the resume
My own AI operations platform: a harness of harnesses that runs Claude Code, Codex, and Gemini sessions side by side, routes each job to the cheapest capable model under per-role limits, keeps a Git-backed memory with human-approved edits, and puts every external action behind an approval gate. Runs locally today; VPS deployment and team use are next.
The problem
Eight to ten editor windows, and no shared memory
Every project ran in its own editor window with its own agent. Prompts went to the wrong project, each agent started from nothing, and nothing coordinated them against one plan. I wanted one screen that runs every project as a live agent session, remembers what happened, and asks me before anything leaves.

Harness of harnesses
Every coding agent, one control surface
Claude Code, Codex, and Gemini CLI run as real sessions today, each as an interchangeable lane behind one approval gate and one memory. OpenCode and Cursor are supported lanes that switch on once their credentials exist.
Launch
Pick a project, an engine, optional skills or MCP servers, and a first prompt. The session opens in the project's own folder with that project's instructions.
Isolate
Each run is a Room with its own engine, model, and permissions. Rooms never share a prompt window; they coordinate through a shared ledger.
Monitor
An event bus writes every session's output and state to the database. A classifier turns each terminal into a status dot (running, needs input, ready, idle, or error), and a watchdog stops runs that stall.
Steer
Watch or type into any session from the browser, or inject a prompt, pause, kill, or switch engines mid-run. A switch passes the same billing and safety checks as the first route.
Collect
When a Claude Code run finishes, its answer is read from Claude Code's own transcript and filed as a draft in the approval inbox.
You
Cockpit (Next.js)
- Launch composer
- Mission Control
- Inbox
- Brain
Brain API (FastAPI, 127.0.0.1)
- Router and billing guard
- Model policy by role and business
- Session manager
- Event bus to the database
- Approvals
- Memory
CLI sessions in tmux
- Claude Code
- Codex
- Gemini CLI
- OpenCodeneeds credentials
- Cursorneeds credentials
API lanes, no terminal
- OpenRoutermetered, capped
- Ollamafree
- Output and state flow back through the event bus to the status dots
- A finished Claude Code transcript becomes a draft in the inbox
- Memory is injected into every session at start
- No metered Anthropic key
- Secrets proxy
- Per-role model policy



Mixture of models
The cheapest engine that can do the job
Every job goes to the cheapest engine that can do it. Admins limit which model families and tiers each role may use, and every run records why it went where it did.
Cheapest capable first
Only engines with the capability a job needs are candidates, sorted by cost: the free Ollama lane, then flat-rate subscriptions, then capped pay-per-use lanes. A metered Anthropic key is blocked for my own work.
Four families, ten tiers
Claude, GPT, Gemini, and an Ollama family each have a tier ladder, cheapest first, with a default reasoning effort per tier. The Claude ladder runs from Haiku to a Fable-class top tier.
Limits per role and business
Per-role allow-lists for families, models, and tiers are checked on every dispatch, with per-business overrides. A denied request fails closed with a 403 and an audit entry.
Receipts
Each run stores a routed-because receipt: family, tier, effort, and the reason.
Escalation Roadmap
Designed so that a complexity check picks the tier and effort for each request, starting at Haiku and escalating through Sonnet and Opus to Fable when a cheaper tier fails its checks, with a fallback family during an outage.
Email replies draft on the Haiku tier, at a third of Sonnet's token price, inside a monthly budget per business that warns at 80% and stops at the cap.
Capability
Generation, reasoning, or code execution? Engines without it drop out.
Fast path
Triage, classify, label, summarize, or extract go to cheap routing lanes.
Cost order
Free Ollama lane, then flat subscription, then flat headless, then metered (opt-in and capped). Metered Anthropic is blocked.
Policy
Role allow-list for family, model, and tier; business overrides; the stricter cap wins. A denial returns 403 and an audit entry.
Receipt
routed: family/tier · effort, and the reason.
Worked examples
“Triage these 40 emails and label them.”
Fast path, cheapest lane first
Ollama lane, $0. Claude Code on the subscription if Ollama is down.
“Fix the failing test.”
Needs code execution
Claude Code on the flat subscription.
“Reply to a customer email.”
Drafting policy per message type
Haiku tier, inside a $10 monthly business cap.
- Roadmap
Complexity rises or checks fail: escalate Haiku, Sonnet, Opus, Fable.


Mixture of agents
Agents that plan, debate, and wait
Several agent patterns run through the cockpit, and each one stops for a person before anything consequential happens.
Guided agents
An agent asks a few intake questions, shows its plan, and waits. Nothing runs until I approve the plan, and the wizard never sends.
Orchestrator skills
A skill can chain other skills, such as inbox triage feeding the daily brief.
Debate
Two Rooms can argue a plan through questions and answers on a shared task ledger, never a shared prompt, and hand back one reconciled note.
Orchestrator Feed
When a Claude Code session spawns its own subagents, the cockpit shows each one live.
Cross-project bar
Type a plain-English launch and it finds the right project and engine, or answers a question from the run history.
Orchestrated swarms Roadmap
Designed so that a Fable-class orchestrator breaks a prompt into parts and spawns cheaper subagents, such as OpenAI GPT models or Claude Sonnet, reporting into one central feed.

The approval gate
Nothing leaves without a human
Emails, posts, and memory edits all arrive as drafts. Each skill runs three to five automated checks before I see a draft, retries once if a check fails, and then waits for me to approve, edit, reject, or regenerate it.
One inbox
Every approval lands in one keyboard-driven inbox, grouped by business, with a panel that shows what the drafter knew.
One winner
Each approval can be decided only once, so a double click cannot send an email twice.
Earned trust
After a streak of clean approvals, the system drafts a card proposing that a skill move to verified-auto, with the evidence. If I approve it, low-risk internal drafts from that skill auto-approve, each with an audit entry.
Automatic demotion
Any rejection, or three edits in ten runs, puts the skill back on drafts.
External sends stay manual
The trust ramp can never remove the gate on external sends.
External sends always need a human.
Draft
A skill drafts the email, post, or memory edit.
Checks
Three to five automated checks per skill.
Self-heal
One retry if a check fails.
Human
Approve, edit, reject, or regenerate.
Act
Send, publish, or commit.
Audit
Every decision is logged.
Earned trust
- A streak of clean approvals
- A promotion card with the evidence
- I approve the promotion
- Low-risk internal drafts auto-approve, with an audit entry
Any rejection, or three edits in ten runs, demotes the skill automatically.


The memory system
Memory that shows its work
The system of record is a Git wiki plus the operational database. Every new session starts from it, and it changes only through edits I approve.
Session start
Every new session boots with the goal, the memory index, learned rules, and the last ten journal entries, capped at 12,000 characters.
Approved edits
A memory review arrives as a diff I can edit. Approving it commits the change with the approver in the commit trailers; rejecting it changes nothing.
Procedures from real runs
A finished run becomes a standard operating procedure after approval, linked to the run that produced it and marked fresh or stale.
Time machine
Any wiki page can be viewed as it was on a past date, and each fact links to the log it was learned from.
NextNightly consolidation, semantic recall, and prediction tracking are built and tested; switching them on is next.
A run happens; its events and transcript are saved
Procedure
- SOP mined from the run
- Approved by me
- Filed in the wiki, fresh or stale
Memory edit
- Review proposes a diff
- I edit and approve
- Git commit with the approver
Journal
- System events logged
- Progress, roadblocks, decisions
- Last ten ride into the next session
The next session starts with the goal, memory index, learned rules, and recent journal
The Brain shows every page, where each fact came from, and any page on a past date
Built and tested, switching on next
- Nightly consolidation
- Semantic recall
- Learned-rule capture
- Prediction tracking



Operating the businesses
The businesses run from the same screen
The cockpit also runs the work around the code: what each business does, which parts a skill can take, and which stay with a person.
Work Map
Every task a business does, marked automate (runs a skill) or keep-human, with gaps flagged as skills to build.
Connectors
Gmail is live, with read access and approved sends. GoHighLevel reads and signed webhooks are live. Credentials sit in an encrypted secrets proxy that agents never see.
Inbox triage
Scheduled Gmail triage classifies mail and drafts replies into the approval inbox.
Command palette
One search across skills, MCP servers, plugins, actions, and sessions. The registry finds installed skills and MCP servers and turns them into buttons.

Every day
A day in the cockpit
The routine parts run every day: a brief for each business, a sorted inbox with replies already drafted, a plain sentence when a number moves, and an optional background that follows the time of day.
Daily brief
One brief per business, and one across all of them, in three parts: what needs me, what is worth knowing, and what is steady. It is filled in from the day's records by a template, with no model writing it, so it cannot invent a number.
Gmail triage
Every hour, triage reads the inbox read-only, sorts mail into needs a reply, urgent, and newsletters or FYI, and drafts replies that wait in the approval inbox. Nothing sends until I approve it.
Watchtower
Watchtower remembers the normal level of each number it tracks. When one drifts, it says so on Home in a plain sentence, such as a count running well below its 30-reading average.
Living Sky
An optional background whose sun, moon, and horizon follow the real time of day where you are, from dawn through midday and golden hour to night. There are seven scenes to pick from, and choosing Light or Dark turns it off.
Break nudge
After about 90 minutes of continuous work, a gentle nudge suggests stepping away.




How it is built
Built like it matters
Python with FastAPI, async SQLAlchemy, and Alembic on PostgreSQL; a Next.js and React cockpit over WebSockets; CLI sessions driven through tmux on WSL.
Loopback only
The API answers only on this machine, blocks any website it does not know, and gives each live terminal its own signed ticket.
Secrets
Credentials are encrypted at rest and read-only by default; write access is an explicit, audited grant. A secret scan covers every Git ref.
A hard boundary
The same core also backs an edition that can never run code. Code execution is a capability the operator build grants itself, and the boundary has its own tests.
Tested and audited
More than 2,700 automated tests. An external audit raised 28 findings, and the first remediation phase merged on September 25, 2026.
Sign-in Roadmap
Designed so that each person signs in, through single sign-on for a team, before it moves to a VPS.
Built with
- Python and FastAPI
- Async SQLAlchemy and Alembic
- PostgreSQL
- Next.js and React
- WebSockets
- tmux sessions
- Git-backed memory
How I build it
Command Center OS is built the way it is meant to run: models directing models, with a second vendor checking the risky parts.
Orchestrate
A frontier-model orchestrator (Opus or Fable class) plans each pass and directs the work.
Implement
Sonnet implementation agents write the code, and verification agents check it against the plan.
Review
Codex reviews anything that touches approvals, credentials, or auth, adversarially. Ninety-five commits reference Codex review rounds.


