The problem this solves
Outbound programs generate exactly the data needed to improve them, and almost none of it gets used. Replies, positive rates, meetings booked - it all sits in the sequencer's reporting tab, glanced at when a campaign ends, summarized as "that one did okay." Next month's targeting is then chosen the way last month's was: by instinct, by whoever argued hardest, by inertia.
The teams that compound are the ones that run the loop: rank what converted, diagnose what did not, reallocate, repeat. The loop stalls in practice because the retro is genuinely tedious - joining results across segments and message angles, separating deliverability problems from targeting problems from copy problems, and turning the diagnosis into next month's concrete plan. It is an afternoon of analysis with no natural owner, so it happens once and then never again.
This mission is that afternoon, on demand. Hand it your results - pasted or attached - and it ranks segments and angles by what actually converted, diagnoses the losers honestly (was it the list, the message, or the inbox?), and produces next month's plan: what to double, what to rewrite, and the three drafts to lead with.
How the mission runs
- Ingest the results. Reply, positive-reply, and meeting data by segment and message angle - pasted into the mission or attached as a CSV - becomes the ground truth. The mission computes over the actual rows, so the ranking is arithmetic rather than recollection.
- Rank what converted. Segments and angles are ranked by the metric that matters to you - positive replies or meetings per hundred sends - with volume-aware honesty: a 2-for-20 segment is flagged as promising-but-thin, never crowned over a 30-for-600 workhorse.
- Diagnose the losers. Underperformers get a differential diagnosis: deliverability checks on the sending domains separate inbox problems from message problems, and the losing copy is read against the winners to name what differs - claim, proof, ask, or fit.
- Write the reallocation plan. The findings become next month's concrete plan: which segments get doubled volume, which get paused, which angles get rewritten and how, and what single experiment the month should carry. Each call cites the numbers behind it.
- Draft the leads. The three drafts to lead with next month are written out in full - built from the winning angles, aimed at the doubled segments - so the plan ends one edit away from execution rather than as advice.
The prompt
This is the exact objective the agent receives. Swap the obvious placeholders for your own domain, segment or channel and run it as-is from the console, Slack, or the API.
What comes back
An outbound retro that ends in a plan: segments and angles ranked by real conversion with volume honesty, a diagnosis separating list, message, and inbox problems, a cited reallocation plan for next month, and the three lead drafts written in full. Run it monthly and the program compounds instead of drifting.
Make it yours
- Run it after every distinct campaign rather than monthly when your volumes are high enough for the numbers to settle faster.
- Chain it with the URL-to-pipeline playbook: the retro picks the doubled segment, and the pipeline playbook builds its next list.
- Add the deliverability sweep as a standing precheck, so inbox problems are caught before they masquerade as message problems in the next retro.
Frequently asked questions
What data does it need, exactly?
Sends, replies, positive replies, and meetings, broken down by segment and by message angle - the export every sequencer produces. Attach the CSV or paste the table; the mission computes over the rows exactly rather than skimming them.
How does it tell a bad list from a bad message?
By triangulating: deliverability checks on the sending domain rule the inbox in or out, reply rate versus positive-reply rate separates attention problems from offer problems, and cross-reading losing copy against winning copy on the same segment isolates the message variable.
Can it work with small numbers?
Yes, with honesty about what small numbers can say. The ranking flags thin samples explicitly and the plan leans on direction rather than false precision - doubling a promising-but-thin segment is framed as the experiment it is.