Skip to content

Approach

Spend less. Ship more. Keep the receipts.

AI made building cheap and made measuring cheap. Most teams took the first gift and skipped the second. My work is both: run AI-assisted delivery the way a good engineering manager runs a team, with scope, review, and verification, and measure what it really costs per thing you actually ship.

Human in the loop

AI makes attempts cheap. Acceptance is the unit that costs money.

live simulation

Candidates

0

Accepted

0

$ / accepted

Illustrative: $0.06 per attempt. The cheap part is the attempt; the real unit is what gets accepted.

Principles

Six rules I use on my own products first

01

Measure the unit that matters

Not tokens, not seats, not cost per image. Dollars per accepted output: a merged change, a shipped panel, a resolved request, including the rerolls and the review time.

Receipt: TigerMill defines it; most teams don’t. →

02

Keep spend out of the agent’s toolbelt

Agents can inspect, claim, and close gates. Anything that costs real money happens behind a separate, deliberate step with a named owner.

Receipt: 13 MCP tools, none of which can generate an image. →

03

Gates with attribution, not blanket approval

Let agents approve objective checks, record who acted every time, and keep subjective acceptance with a named human.

Receipt: An agent once approved art I rejected. The log said so. →

04

Verified live beats “tests passed”

Typechecks that block deploys, checking the change on production, and screenshots an agent actually reads. New kinds of users walk paths you never tested.

Receipt: An MCP edit path surfaced a stale edge cache in TendForm. →

05

Make the numbers true before making them better

Require every change to show its measurement, then read it. The first finding is usually that a trusted number was wrong.

Receipt: TrendVesting’s confidence scores ran 55 points hot. →

06

Rebuild smaller

The cheapest system to run is the one with less in it. Delete code, collapse services, and put the edge in front of anything that doesn’t need a server.

Receipt: This site was rebuilt in a day, screenshot-checked on mobile before every deploy. →

The delegation stack

Who does what

The same structure runs TendForm, TrendVesting, TigerMill, and this site. For a client engagement it gets scoped down: agents touch only the repositories and commands we agree on, with no production credentials and no ability to spend.

Read the playbook
  1. 1

    You

    Set the goal, the budget, and what “accepted” means.

  2. 2

    Lead agent

    One persistent lead per repository, in its own worktree, with notes that survive restarts.

  3. 3

    Workers

    Short-lived agents for research, drafting, and implementation, each scoped to specific files.

  4. 4

    Gates

    Typecheck, build, verified-live checks, screenshot review, and human sign-off on anything subjective or paid.

  5. 5

    CI/CD

    Every commit builds; main deploys. Commit bodies carry the reasoning and the measurements.

Why me

Three scales, one discipline

I started as a middle-school teacher who built a grammar game because twelve-year-olds told me the lesson was boring. That instinct, make the complex thing legible to the people who have to live with it, is still the job. I’ve done it for a $10–12M-a-quarter AWS bill at Roku, across a ~700K-server fleet at Apple, and dollar by dollar on products whose invoices I pay myself.

Every AI system has two architectures: the one in the diagram, and the one on the invoice.

Start with one system

The Unit Economics Sprint takes one AI workload and gives you a defensible cost per unit and the top three levers, in a day.

from $5,000

Rates & fit check

Cash-pay advisory
from $250/hr

Check fit