FIELD REPORT 16 · AI SYSTEMS · HIGH TICKET SALES CLIENT

The agent operated
ad account

I wrote the rules for a six figure monthly ad account, then built the agents that run them. Two judgment reviews a day make the cuts and budget steps the rules allow, log every action with its before and after values, and page a person only when a call is outside their authority.

2xdaily reviews,
every action logged
ONE DAY OF THE ACCOUNT · THE SCHEDULE IN ROUND TIMES
00 03 06 09 12 15 18 21 JUDGMENT REVIEW REPORTS ROLLOUT WATCHDOG 06:00 watchdog raise, push days only 07:00 morning report 08:00 creative report, Mondays 08:30 judgment review, morning 10:00 ops flags 16:00 rollout ladder 20:30 judgment review, evening 23:00 watchdog brake, push days only
READ THE REPORT

I wrote the rules first. The agents came second.

The client sells a high ticket offer through booked sales calls. Paid media runs at six figures a month on Meta and Google, across several ad accounts, with new creative going in all the time.

An account that size asks the same questions every day. Which ads crossed the line and should stop? Which earned more budget, and was the last raise long enough ago? Answered from memory, those calls drift. A favorite ad gets a day it didn't earn, and a good morning gets a raise that should have waited.

So I wrote the doctrine down with a number in every rule. Then I built agents in Claude Code and Python that run it twice a day, act inside it, log what they did, and page a person when a call is outside their authority.

ROLEWrote the doctrine,
built the agents
SPENDSix figures a month,
Meta + Google
RHYTHMTwo judgment
reviews a day
WHEN2025 to now

Every rule in the doctrine has a number in it.

The agents read this file at the start of every run. Under it, the same logic runs as a simulator. Set the inputs to see which rule fires and the log line the agent would write.

R1Kill rule

Spend reaches 3x the target cost per qualified application with no qualified application: cut. No waiting for another day of data.

R2Raise rule

Budget goes up in steps of 25%, at least 48 hours apart. Cuts are exempt from the wait.

R3Spend gated test windows

A test is judged on spend, never on days live. Nothing is cut or raised before it has spent 3x its target cost per qualified application.

R4Never touch list

Named campaigns the agent can read but never change. If a rule fires on one, it pages a person with its recommendation.

R5Daily change cap

Each review can move the account's net daily budget only so far. Anything past the cap waits for the next review or a person.

● RUN THE DOCTRINEONE AD · ILLUSTRATIVE INPUTS · $400 A DAY BUDGET

TRY A CASE, OR MOVE THE INPUTS

$300
$950
0
72h
ON THE NEVER TOUCH LIST
HEALTH CHECK

THE CHECKS, IN ORDER

  1. H0
    HEALTH CHECKPacing, disapprovals, tracking heartbeat and billing all green.
    PASS
  2. R3
    TEST WINDOWClosed at $900. Spent $950, 3.17x target.
    CLOSED
  3. R1
    KILL LINE0 qualified at 3.17x target. Cut, no waiting.
    FIRES
  4. T
    TARGETNot needed once the kill line fires.
    NOT REACHED
  5. R2
    RAISE CLOCKCuts skip the 48 hour wait.
    SKIPPED
  6. R4
    NEVER TOUCHNot listed. The agent may act.
    CLEAR

THE CALL

CUT

R1 · KILL RULEThe ad spent 3x its target without one qualified application. The agent pauses it now and does not wait for more data.

SPENT VS WINDOW
$950 of $900 · 3.17x
COST PER QUALIFIED
none yet
DAILY BUDGET
$400 → $0

THE LOG LINE · EXAMPLE FORMAT

08:30 am_review ad_0417 action=CUT rule=R1_kill_3x spend=$950 qa=0 target=$300 (3.17x) status=ACTIVE→PAUSED budget=$400→$0 actor=agent

Illustrative inputs for one ad with a fixed $400 daily budget, so the before and after values are easy to read. The live agent reads trailing windows by lane and by ad from the data lake. R5 applies to a whole review, so a single ad never trips it here.

Same loop, twice a day, inside a fence.

A scheduled headless run in Claude Code and Python. Everything inside the dashed line it may do alone. Anything outside it becomes a page to a person.

INSIDE THE LINE · THE AGENT ACTS ALONE THE NEXT RUN STARTS FROM THIS LOG 01READ the doctrine fileand the last log 02CHECK health acrossevery account 03PULL lanes and adsfrom the lake 04DECIDE run the rulesin order 05ACT cuts, budgetsteps, launches 06LOG before and after,read back RED CHECK OUTSIDE ITS AUTHORITY 07 PAGE A PERSON with a recommendation OUTSIDE THE LINE A PERSON DECIDES INSIDE THE LINE · ACTS ALONE 01READdoctrine and the last log 02CHECKhealth, every account 03PULLlanes and ads, the lake 04DECIDErun the rules in order 05ACTcuts, steps, launches 06LOGbefore, after, read back NEXT RUN STARTS FROM THIS LOG 07 PAGE A PERSON with a recommendation RED CHECK OR OUTSIDE ITS AUTHORITY: A PERSON DECIDES
SPEND PACING

Each account against the spend expected by that hour.

DISAPPROVALS

New rejections and policy flags since the last run.

TRACKING HEARTBEAT

Conversion events still arriving from the site and the CRM.

BILLING

Account status, payment method and spending limits.

HEALTH CHECK · STEP 02 · EVERY ACCOUNT, EVERY RUNA red item stops the run before any rule fires.

THE AGENT DOES ALONE
  • Cut at the kill line, the moment it's crossed
  • Raise 25% when the 48 hour clock allows
  • Hold, and write down why
  • Move a cleared file up one rung
THE AGENT PAGES A PERSON FOR
  • Anything on the never touch list
  • A red health check
  • A move past the daily change cap
  • Any call the doctrine doesn't cover
● THE ACTION LOG · ONE MORNING REVIEWEXAMPLE LOG FORMAT · ILLUSTRATIVE ROWS
Example action log from one morning review, with illustrative rows
TIMEOBJECTACTIONRULEBEFOREAFTERWHY
08:31Test lane ad setCUTR1 kill ruleActive, $250/dPaused, $03.1x target spent, 0 qualified
08:31Scale ad set ARAISER2 raise rule$800/d$1,000/dUnder target, 52h since the last raise, value read back
08:32Scale ad set BHOLDR2 48 hour clock$600/d$600/dUnder target, raised 30h ago, next step in 18h
08:32New fileLAUNCHLadderRung 1Rung 2Cleared rung 1, cooldown over, a review slot open
08:33Protected campaignPAGER4 never touch$1,200/d$1,200/dAn ad in it crossed the kill line. A person decides.
08:33Net, this reviewNETR5 daily cap$2,850/d$2,800/d−$50 a day across the rows above, inside the cap

Budgets are illustrative and add up: the cut removes $250 a day and the raise adds $200 (25% of $800), so the review nets out at −$50 a day. Every row carries its before and after value, and the after value is read back from the ad platform, never assumed.

New creative earns its way up, one rung at a time.

Every file starts at the bottom, where a rejection costs little, and climbs only after surviving the rung it's on. Rung names and settings are an example of the structure.

  1. RUNG 0Compliance grade

    The grader scores the copy before anything uploads.

    TO ENTERSCORE ABOVE 80
  2. RUNG 1Test account

    A low risk account. New concepts go in together, so reviews run in parallel.

    GATESREVIEW CAPDAILY ADD LIMIT
    REJECTED ↓ COUNTS ONE
  3. RUNG 2Second test account

    A second review on a different account, and the first qualified signal.

    GATESCOOLDOWN 48HHOLD AFTER APPROVAL
    REJECTED ↓ COUNTS TWO
  4. RUNG 3Scale account

    Only files that cleared both test rungs. One at a time, on a floor budget.

    GATESONE FILE AT A TIMECOOLDOWN 48H
  5. RUNG 4Main account

    Proven winners only. Nothing is tested here.

    GATESWINNERS ONLY
RETIRED

Two rejections and a file stops climbing for good. A retired file never goes back up the ladder; the idea goes back to the edit.

Example settings: at most two files in review on a rung at once, a 48 hour cooldown between adds on the same rung, a hold of a few hours after an approval before the next add.

● RUNG 0 · THE COMPLIANCE GRADERCLAUDE API · ILLUSTRATIVE COPY AND SCORES
  1. CLEARBook a call and see if it's a fit.
  2. RED · CHANGEGuaranteed results in 30 days.Promises an outcome. This line has to change before upload.
  3. YELLOW · RISKAre you tired of doing it all yourself?Asks the reader about a personal attribute, which Meta reviews closely.
  4. YELLOW · RISKOnly 3 spots left this week.Scarcity has to be true on the day the ad runs.
GO 0 50 80 100 FIRST PASS 58 AFTER THE FIX 84
FIRST PASS58No go. One red line, two yellow.
AFTER THE FIX84Go. The copywriter changed the red line.

The grader reviews and never rewrites. Every flag carries a reason, and a person makes the change. I calibrated its scoring against ads the platforms had actually approved.

Nobody has to ask how yesterday went.

Three reports go to people on a fixed rhythm. A watchdog script handles time boxed pushes, and the review reads its log.

MORNING REPORT07:00 · DAILY

The whole funnel before the team starts.

From ad spend to closed deals in funnel order, by account and by lane. It lands at 7 AM, before the morning review.

CREATIVE REPORTMONDAYS

Patterns, not one lucky ad.

EVERY AD, EVERY WEEKROLLED UP THREE WAYS

BY ANGLE BY HOOK BY AUDIENCE PER AD: HOOK RATE, HOLD RATE

Hook rate and hold rate for every ad, rolled up by angle, hook and audience, so the next shoot starts from what keeps working.

OPS FLAGS10:00 · DAILY

One flag per live ad.

EXAMPLE ROWSONE PER LIVE AD

  1. ad_0412under targetIN
  2. ad_0415inside its test windowWATCH
  3. ad_0417R1: 3.2x target, 0 qualifiedOUT
  4. ad_0420cost drifting upWATCH
  5. ad_0423raised this morningIN

IN keeps running. WATCH is inside its test window or drifting. OUT is cut or retired, with the rule that did it. The ops team gets the verdicts without opening an ad account.

THE WATCHDOGPUSH DAYS

Scheduled budget moves for time boxed pushes.

ONE PUSH DAYSHAPE, NOT DATA

BRAKE RAISE 18:0023:0006:0010:00 PACE

An overnight brake and a morning raise, plus pacing checks through the day. It writes absolute budgets, never deltas, so a rerun can't double a change. The review reads its log and pages only on an error.

Two things broke. Both are rules now.

Agents fail quietly. These two taught me the most, so they stay in the report.

POSTMORTEM 01 · THE MISSING LOG

The review acted, then never wrote it down.

WHAT HAPPENED
A headless review ran on schedule and made its changes. It never wrote its log entry.
HOW I CAUGHT IT
The next review compares the live account to the last logged state. They didn't match, so it rebuilt the missing entry from the live account before it touched anything.
ROOT CAUSE
Logging was the last step of the run, separate from the actions. The actions finished. The last step didn't.
THE FIX
The log write is part of the action now. Each change writes its own line with before and after values as it happens, and a change without its line counts as failed.
POSTMORTEM 02 · THE HALF READ TRACKER

Every ad looked 2.4x better than it was.

WHAT HAPPENED
A creative tracker read two of four ad accounts. Every ad in it looked 2.4x better than it was.
HOW I CAUGHT IT
Before anyone acted on it. The tracker's spend didn't tie out to the ad accounts' own totals.
ROOT CAUSE
Spend came from two accounts while results came from all four, so every cost per result was divided by too little spend.
THE FIX
The tracker counts its accounts before it computes anything, and it joins spend and results on the same key. A short read now fails loudly.

Two reviews a day, and every move on the record.

The result is an operating rhythm. Judgment calls happen on a schedule and inside written limits, and anyone can read back why a budget moved.

2xjudgment reviews a day, morning and evening
3xtarget cost per qualified application with none qualified: cut
+25%per budget step, one step at a time
48hminimum between raises. Cuts don't wait.
THEN

The calls happened when someone logged in, and the reasons lived in their head.

NOW

The calls happen at 08:30 and 20:30 against written rules. Every action has a line with its before and after values, and anything outside the rules goes to a person.

CLAUDE CODEPYTHONCLAUDE APIBIGQUERYAPPS SCRIPTLOOKER STUDIOMETA MARKETING API

The rules, schedules and failures come from my own build notes and review logs. Times are the schedule in round numbers. Budgets, log rows, ad copy and scores are illustrative. The client's name and account figures stay private.

An agent needs three things before it touches a budget.

01

Write the rules before the agent. Give every rule a number, or the agent will pick one for you.

02

Log the before and after value of every action, and read the after value back from the platform. The next run starts from that log.

03

Give the agent a boundary and a pager. Inside the line it acts without asking. Outside it, it stops and pages a person with a recommendation.

NEXT · FIELD REPORT 17 · FOUNDERSeven Figures on Shopify and AmazonMy own ecommerce business: about $1M in revenue in its first full year across Shopify and Amazon FBA, then sold.

Happy to walk through any of these live.