Most accounts refresh creative the way people water houseplants: sporadically, on guilt, after visible wilting. The result is a familiar cycle - a control fatigues for weeks while nobody notices, panic produces a batch of new creatives testing five things at once, one "wins" for unknowable reasons, and the cycle resets. The alternative is a cadence: a standing rotation with defined slots, disciplined challengers, kill criteria written before data arrives, and a fatigue watch that feeds the queue automatically. This article assembles it, as the testing chapter of the AI advertising pillar.
The vibes-refresh problem
The failure has two halves. Detection: fatigue is gradual - frequency creeps, response sags - so week-to-week eyes miss it and the account pays a decaying control for a month before anyone reacts. Response: because refreshes are panic events, they arrive as batches testing everything at once (new angle, new format, new hook, simultaneously), which produces winners nobody can learn from. Both halves are cadence problems, and both dissolve when detection is computed and the response is a queue that was ready before the fatigue arrived.
The model: slots and challengers
Structure the account's creative as slots: an audience-placement combination that spends meaningfully (prospecting/feed, retargeting/story, whatever your structure serves) - each holding a control (the current best) and a bounded stream of challengers (one or two at a time, never five). Challengers earn the control seat by beating it under the kill criteria; deposed controls retire with their run logged. The model's virtues are boring and decisive: every test has a defined opponent, spend concentrates enough per challenger to reach significance, and the account always knows which creative is load-bearing where - which is exactly what the PPC audit's creative-health surface reads.
One variable per test
The discipline the panic-batch violates: a challenger differs from its control on one axis - the angle (per the angle map's open ground), the hook, the format, or the visual system. Not two. The reason is learning, not purity: a challenger that changes angle and format and wins tells you nothing about why, so the next test starts from zero again. Single- variable tests accumulate a causal ledger - this audience responds to reliability angles, stories beat feed for this offer - and the ledger is what makes month twelve's testing smarter than month one's. Generative production makes the discipline cheap: rendering one-variable variants to spec is exactly what the Design agent does from a brief.
Kill criteria, written in advance
The fatigue watch
Fatigue detection is a computation, not a feeling: per creative, per slot, track frequency and response weekly and fit the trend - response falling while frequency climbs is the signature, visible in the data a fortnight before it is visible in the topline. The standing mission flags creatives entering decline, which does two things: the slot's next challenger launches before the control collapses (overlap, not gap), and the retirement queue stays stocked with evidence. This is the same computed-not-eyeballed principle as the audit - and it is what turns the cadence from a calendar into a feedback loop.
The weekly cadence, assembled
| Beat | What happens | Source |
|---|---|---|
| Field read | Ad libraries diffed; angle map updated | Creative intelligence loop |
| Fatigue watch | Curves computed; declining controls flagged | Standing mission |
| Test decisions | Bounds-hit challengers promoted or killed per criteria | Pre-registered criteria |
| New challengers | Next open-angle variants rendered and drafted, paused | Angle map + Design agent |
| Review | Five-point read; un-pause | Paused-draft workflow |
Every beat is a mission except the review, which is the point: the human hour per week goes to judgment - promote or kill, which angle next, approve the drafts - while the reading, computing and rendering run on schedule from the Advertising shelf. Run this for two quarters and the compounding is visible: a causal ledger of what works per slot, controls that never silently decay, and a testing program whose next move is always already queued.
Frequently asked questions
How often should ad creative be refreshed?
Continuously, via a standing rotation: each slot holds a control and one or two challengers, with the fatigue watch launching the next challenger before the control declines - overlap, never gap.
What is the slots-and-challengers model?
Structure creative by audience-placement slots, each with a current-best control and a bounded challenger stream competing under pre-registered criteria. It concentrates spend, defines every test’s opponent, and builds a causal ledger.
Why one variable per test?
Multi-variable winners teach nothing - the next test restarts from zero. Single-variable challengers accumulate causal knowledge about angles, hooks and formats per audience, which is what compounds.
What are kill criteria?
Three numbers written before launch: spend bound, time bound, and the deciding metric with threshold. When bounds hit, the decision executes - which is what prevents sunk-cost extensions and impatient kills.
How is creative fatigue detected early?
By computing the curve: response falling while frequency climbs, tracked weekly per creative per slot. The signature shows in data roughly two weeks before it shows in topline results.
Sources
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.