The problem this solves
Sending-tool debates run on vibes. Someone read a thread praising one platform's deliverability; someone else has a habit investment in the incumbent; the pricing pages disagree about what matters. The honest answer - that the right tool depends on your domain, your list, and your sequence - is only reachable by experiment, and the experiment almost never happens because setting it up properly is fussy: identical conditions, a fair split, and a scoring plan agreed before the results start rolling in.
Sloppy split tests are worse than no test, because they produce confident wrong answers. If the halves differ in list quality, if the sequences drift apart during setup, or if unverified addresses seed one half with extra bounces, the comparison measures the setup errors rather than the tools. Fairness is a discipline: same list randomly divided, same copy character for character, same schedule, and a list cleaned before the split so neither half carries hidden handicaps.
The other failure arrives at the end: two weeks later, the results exist and nobody agreed on what winning means. Open rates favor one tool, replies the other, and the debate resumes exactly where it started. The scoring criteria have to be written down before the first send, or the test decides nothing. A checklist agreed in advance is the cheapest referee available.
How the mission runs
- Verify the list before anything splits. Email Verification runs on all 100 contacts from the thread first, so undeliverable addresses exit before the division. A shared cleaning pass guarantees neither half inherits a hidden bounce handicap, which is the quiet way most homemade split tests invalidate themselves.
- Split evenly and fairly. The verified list is divided into two randomized halves, balanced rather than sliced in original order - because source lists are usually sorted by something, and 'first fifty versus last fifty' would quietly test list position instead of the tools.
- Load identical campaigns into Lemlist and Instantly. The same three-touch sequence is created in both tools, character for character, with matching schedules and settings as far as each platform allows. Where a platform forces a difference, it is documented, so any result divergence can be traced to the tools rather than to silent setup drift.
- Hand over the comparison checklist. The mission closes with a written scoring plan for the two-week mark: which metrics to compare - delivery, placement-sensitive opens, replies, bounces - what each divergence would mean, and what sample-size caveats apply at 50 sends per arm. The definition of winning is fixed before the first send.
- Export the split record as CSV. A CSV records which contact went to which tool, preserving the experiment's ground truth. When the results are read, there is no ambiguity about who was in which arm, and the same record makes a follow-up test on the losing half trivially easy to run.
The prompt
This is the exact objective the agent receives. Swap the obvious placeholders for your own domain, segment or channel and run it as-is from the console, Slack, or the API.
What comes back
Two live campaigns - half the verified list in Lemlist, half in Instantly, identical sequence and schedule - plus a CSV recording the exact split and a written checklist for the two-week readout: metrics to compare, what each gap would mean, and the caveats that keep 50-per-arm conclusions honest. The tool debate ends with your own data instead of someone else's thread.
Make it yours
- Re-run the test with your next campaign's list to check whether the winner holds across segments before you consolidate spend.
- Test one tool against itself with two different sending domains instead, isolating domain reputation as the variable while holding the platform and the copy constant.
- Extend the readout to include reply quality, since a tool that earns fewer but warmer replies may be the actual winner.
Frequently asked questions
Is 50 contacts per arm enough to decide?
Enough to detect large differences - a deliverability gap or a broken configuration shows up clearly - and honest about small ones, which the checklist's caveats spell out. Treat the first run as a screening test; if the tools land close, the follow-up on a second list settles it.
Do both campaigns launch automatically?
Both are assembled and staged for your review first, like every send-adjacent mission. You confirm the two campaigns match and release them together, since a staggered start would add a timing variable the test design just worked to eliminate from the comparison.
What if the platforms cannot be configured identically?
Genuine platform differences are part of what the test measures, and forced setup differences are documented in the handover so you can weigh them at readout. The discipline is transparency: every known asymmetry is written down before results arrive, keeping the conclusion honest.