A one-person studio, with real money on the line
Execution can be delegated. Judgment cannot. So I built a studio where agents execute and criticize, and money only moves on a human call.
- Role
- I built and operate the entire studio. The brand is my wife's business; the operation runs on systems I made.
- Context
- a small family e-commerce brand (pet fashion, print on demand), run outside working hours, with our own ad budget
- System
- an operations cockpit plus a plugin with 16 skills and 4 fixed-role agents, published as open source
- Period
- November 2025 to present, in operation
The result, with its frame
August 2026 closed with 36 paid orders and a ROAS of 3.23, against a breakeven of 2.50 recalculated for that month's basket. It landed inside the 20% profit target for the first time, with no budget step, so the gain came from efficiency rather than from spending more. Even so the lock on scaling stayed shut, because the rule asks for the new-customer CPA to hold two weeks first.
The frame matters because the numbers lie easily. In the week of 17 July the blended ROAS was 2.47, above breakeven, and my own weekly briefing called it illusory, because it came entirely from one ad. Without that ad, the rest of the account ran at 0.75. Open audiences never worked here either. Five ad sets, more than R$600, all below breakeven.
What we watch
ROAS is the headline, not the dashboard. Every day and every week the agents read CPA and ROAS per ad, CTR and frequency to catch fatigue early, average order and items per order, and a CAC breakeven recalculated each month from the real basket.
On the site, other agents cross GA4 with Meta. They read sessions, add-to-cart per session and checkout-to-purchase, per ad destination, to tell whether a problem lives in the ad or in the page. The next front is video, and the first real clip ran in August, with a doctrine of its own.
Everything is read campaign by campaign. Below, the account as a whole, June to August. July has no closed ledger, so its items per order are estimated from attributed revenue and the average item price, and it keeps June's breakeven, the rule in force until the August basket closed.
| Jun | Jul | Aug | |
|---|---|---|---|
| CPA | R$108 | R$69.52 | R$60.84 |
| ROAS | 1.30 | 2.82 | 3.23 |
| Average order | R$140.66 | R$195.74 | R$196.32 |
| Items per order | 1.3 | ≈2.0, estimated | 2.03 |
| CAC breakeven | R$58.86 | R$58.86 | R$78.58 |
| ROAS breakeven | 2.39 | 2.39 | 2.50 |
The choice
I choose by looking at the roughs, never by reading a description. Here are two routes for the same brief, three Shih Tzu shirts on a surface of the house. The linen sofa, where the dog actually climbs, or the light porcelain floor.
I picked the sofa. It keeps the warm, textile feel of the winning parent ad. And its mid-tone sand fixes a finding from the previous round, a surface so light it lost the scroll against the white feed, without breaking the brief's veto on dark backgrounds.
ROUTE 1 · LINEN SOFA · CHOSEN
ROUTE 2 · PORCELAIN FLOORThe studio
The whole cycle runs on agents with fixed roles, like a real studio. Diagnosis reads the ad account and the funnel. Strategy prioritizes the test backlog. In production, an art director proposes visual routes anchored in reference, a producer executes the candidates, and an adversarial validator tries to kill every piece before I see it.
The approved campaign is born paused, because turning it on is always a human gesture. Then measurement runs against fixed rules, and what the operation teaches gets promoted into versioned methodology, like software.
Judgment, built in
The validator is fail-closed. Any blocking criterion kills the piece, and the verdict is written down. This one failed on the first pass because the AI drew a dog with no eyes and the body of another breed, and the text scrim had a hard edge. One regeneration later it passed with zero warnings. Then the approval package set the campaign up at R$15 a day, paused, waiting for me.
FAILEDITERATION 1 · FAILED: ANATOMY AND SCRIM
PASSEDITERATION 2 · PASSED, ZERO WARNINGSWhat the operation taught
The rule I value most in version 1.3.0 of the method is that a measurement only becomes a rule after three questions. How many cases, with what variance? Does the sample cover the variants that matter? What cheap counter-example would refute it, and was it checked?
My hardest problem here is not the design. It is handing my judgment to the AI, describing what I approve and what I reject in a way it can apply on its own, learn from what passed, and keep that memory between cycles. That work is what became the rules, the contract tests and the changelog.
Results
- August 2026: ROAS 3.23 against a breakeven of 2.50, inside the profit target for the first time. Scaling still locked by rule.
- A losing channel named and closed: link-click campaigns brought 793 clicks and zero sales, and are now banned as a sales channel by rule.
- One night of operation in July produced 18 learnings; only one already lived in the method. The ones that are methodology went into version 1.3.0 of the open-source plugin; the business numbers stayed in the workspace, by contract.
Limits and learning
The scale is small, and declared. A niche brand with our own budget, two to three thousand reais a month in ads. The case is not about size. It is about a system that keeps me honest with small numbers, which is exactly when it is easiest to fool yourself.
The images are generated. The validator catches anatomy and legibility, but not taste. Taste is still the human gate, and it is the part I would least like to automate.