Skip to content

A one-person studio, with real money on the line

Execution can be delegated. Judgment cannot. So I built a studio where agents execute and criticize, and money only moves on a human call.

Role
I built and operate the entire studio. The brand is my wife's business; the operation runs on systems I made.
Context
a small family e-commerce brand (pet fashion, print on demand), run outside working hours, with our own ad budget
System
an operations cockpit plus a plugin with 16 skills and 4 fixed-role agents, published as open source
Period
November 2025 to present, in operation

The result, with its frame

August 2026 closed with 36 paid orders and a ROAS of 3.23, against a breakeven of 2.50 recalculated for that month's basket. It landed inside the 20% profit target for the first time, with no budget step, so the gain came from efficiency rather than from spending more. Even so the lock on scaling stayed shut, because the rule asks for the new-customer CPA to hold two weeks first.

The frame matters because the numbers lie easily. In the week of 17 July the blended ROAS was 2.47, above breakeven, and my own weekly briefing called it illusory, because it came entirely from one ad. Without that ad, the rest of the account ran at 0.75. Open audiences never worked here either. Five ad sets, more than R$600, all below breakeven.

Line chart with three months of ROAS in 2026 against the breakeven line: June 1.30, below the 2.39 breakeven; July 2.82 from the Meta pixel; August 3.23 against a recalculated breakeven of 2.50.
CHART · ROAS BY MONTH, JUNE TO AUGUST 2026. JUNE AND AUGUST FROM THE MONTHLY P&L, JULY FROM THE META PIXEL

What we watch

ROAS is the headline, not the dashboard. Every day and every week the agents read CPA and ROAS per ad, CTR and frequency to catch fatigue early, average order and items per order, and a CAC breakeven recalculated each month from the real basket.

On the site, other agents cross GA4 with Meta. They read sessions, add-to-cart per session and checkout-to-purchase, per ad destination, to tell whether a problem lives in the ad or in the page. The next front is video, and the first real clip ran in August, with a doctrine of its own.

Everything is read campaign by campaign. Below, the account as a whole, June to August. July has no closed ledger, so its items per order are estimated from attributed revenue and the average item price, and it keeps June's breakeven, the rule in force until the August basket closed.

JunJulAug
CPAR$108R$69.52R$60.84
ROAS1.302.823.23
Average orderR$140.66R$195.74R$196.32
Items per order1.3≈2.0, estimated2.03
CAC breakevenR$58.86R$58.86R$78.58
ROAS breakeven2.392.392.50

The choice

I choose by looking at the roughs, never by reading a description. Here are two routes for the same brief, three Shih Tzu shirts on a surface of the house. The linen sofa, where the dog actually climbs, or the light porcelain floor.

I picked the sofa. It keeps the warm, textile feel of the winning parent ad. And its mid-tone sand fixes a finding from the previous round, a surface so light it lost the scroll against the white feed, without breaking the brief's veto on dark backgrounds.

Route 1 rough: three folded black Shih Tzu t-shirts on the seat of a sand-coloured linen sofa, soft daylight, blurred backrest above.ROUTE 1 · LINEN SOFA · CHOSENRoute 2 rough: the same three t-shirts laid on a light grey porcelain floor, flatter light, harder surface.ROUTE 2 · PORCELAIN FLOOR
ROUGHS · ONE BRIEF, TWO ROUTES, SEPTEMBER 2026

The studio

The whole cycle runs on agents with fixed roles, like a real studio. Diagnosis reads the ad account and the funnel. Strategy prioritizes the test backlog. In production, an art director proposes visual routes anchored in reference, a producer executes the candidates, and an adversarial validator tries to kill every piece before I see it.

The approved campaign is born paused, because turning it on is always a human gesture. Then measurement runs against fixed rules, and what the operation teaches gets promoted into versioned methodology, like software.

A closed multi-agent loop: brief, art director, producer and an adversarial validator that can kill a piece, followed by a highlighted human decision gate, a campaign always created paused, measurement against unit economics, and learning promoted to a versioned method that feeds the next brief.
DIAGRAM · THE STUDIO LOOP AND THE HUMAN GATE

Judgment, built in

The validator is fail-closed. Any blocking criterion kills the piece, and the verdict is written down. This one failed on the first pass because the AI drew a dog with no eyes and the body of another breed, and the text scrim had a hard edge. One regeneration later it passed with zero warnings. Then the approval package set the campaign up at R$15 a day, paused, waiting for me.

Rejected render: a French bulldog lying belly-up on a bed under the headline, drawn with no visible eyes and an elongated body of another breed; the darkened band behind the text ends in a hard horizontal line.FAILEDITERATION 1 · FAILED: ANATOMY AND SCRIMApproved render: the same headline over a brindle French bulldog sleeping on its side on a bed, eyes closed, compact body, four paws, soft gradient behind the text.PASSEDITERATION 2 · PASSED, ZERO WARNINGS
RENDERS · THE VALIDATOR'S VERDICT, JULY 2026

What the operation taught

The rule I value most in version 1.3.0 of the method is that a measurement only becomes a rule after three questions. How many cases, with what variance? Does the sample cover the variants that matter? What cheap counter-example would refute it, and was it checked?

My hardest problem here is not the design. It is handing my judgment to the AI, describing what I approve and what I reject in a way it can apply on its own, learn from what passed, and keep that memory between cycles. That work is what became the rules, the contract tests and the changelog.

Results

  • August 2026: ROAS 3.23 against a breakeven of 2.50, inside the profit target for the first time. Scaling still locked by rule.
  • A losing channel named and closed: link-click campaigns brought 793 clicks and zero sales, and are now banned as a sales channel by rule.
  • One night of operation in July produced 18 learnings; only one already lived in the method. The ones that are methodology went into version 1.3.0 of the open-source plugin; the business numbers stayed in the workspace, by contract.

Limits and learning

The scale is small, and declared. A niche brand with our own budget, two to three thousand reais a month in ads. The case is not about size. It is about a system that keeps me honest with small numbers, which is exactly when it is easiest to fool yourself.

The images are generated. The validator catches anatomy and legibility, but not taste. Taste is still the human gate, and it is the part I would least like to automate.