"Personalization" in outbound has meant, for a decade, a first-name token and a company-name token in a template everyone recognizes on sight. That version is dead - filtered by machines enforcing Google's sender guidelines, ignored by humans, and actively harmful to email deliverability through the engagement signals it fails to earn. What replaced it in programs that work is narrower and more honest: reference something real. One observable, relevant fact about the account, connected to a problem you can help with. This article is the craft of that sentence - the evidence hierarchy, the ethical line, and the rules that keep it honest at scale.
What died: token-swap personalization
The economics that killed the template are worth stating because they also explain the replacement. When every team can send thousands of token-swapped emails, the tokens carry zero information - the recipient learns nothing about whether you understand their situation, so the reply rate converges on the spam baseline, so senders compensate with volume, so filters tighten, so the baseline falls further. The only exit from the spiral is mail that demonstrates understanding, and understanding cannot be tokenized - it has to be evidenced. Which conveniently is exactly what signal-based targeting produces as a by-product: every account arrives with the reason it was chosen.
The evidence hierarchy
| Tier | Examples | How to reference |
|---|---|---|
| What they published | Job postings, funding announcements, launches, blog posts | Directly and by name - they chose to make it public |
| What they observably do | Tech stack, live ads, content strategy, hiring velocity | Directly, framed as observation of the company |
| What they consume | Intent signals, topic research | Obliquely only: inform timing and offer, never “I saw you researching” |
The top tier is the workhorse - the postings playbook exists because published evidence supports the most direct, least awkward reference. The middle tier ("you are running Meta ads with a single creative angle" - readable from the Meta Ad Library) works when framed as professional observation. The bottom tier is the one teams misuse: intent data should move an account up the queue and shape the offer - it should never appear in the copy, because "we noticed your company researching X" reads as surveillance even when technically aggregate.
The line: observed versus creepy
Inside the line, confidence is warranted: companies publish postings, run ads and announce funding precisely because they are public acts. The squeamishness some teams feel about referencing them is miscalibrated - the recipient knows their posting is public, and relevance built on it reads as diligence, not intrusion.
Offer-matching: the signal sets the ask
The most underused personalization lever is not the opener - it is the offer. A signal reveals a situation, and situations want different things: an account deep in comparison research wants the honest comparison, not a demo; a company that just posted for its first marketing operations hire wants "what that hire will need in their first quarter"; a fresh raise wants acceleration material, not education. Matching the ask to the situation is what makes the email useful - and useful mail earns the replies that reputation feeds on. A touch whose offer ignores its own evidence ("saw you're hiring - want a demo?") wastes the signal it opened with.
The drafting rules
Five rules produce the shape that works. Evidence first: the observable fact is sentence one - not your company, not "hope this finds you well". One claim: connect the evidence to exactly one problem; two problems is a newsletter. One ask: sized to the evidence - a strong signal earns a meeting ask; a weak one earns "worth a look?". No flattery: "impressive growth!" is filler that signals template; the evidence itself is the compliment. The screen test: every sentence should survive the recipient pulling up the evidence beside it - which is also the QA method: draft with the evidence file open, per the same citation discipline that governs every agent deliverable.
Doing it at scale without faking it
The scaling problem is real - evidence-grounded drafting takes minutes per account that template blasting takes seconds for - and it is exactly the shape agents solve honestly. A mission carries each account's evidence file (the signals that selected it, per the intent data and postings pipelines), drafts against the rules above with the evidence in context, and parks the batch for approval - so a human reviews finished, grounded drafts instead of writing them. The failure mode to refuse is the counterfeit version: templates with an evidence-shaped blank ("[PERSONALIZATION]") filled by a model guessing from the company name. The test is provenance: if the draft cannot show the source of its first sentence, it is a token-swap with better prose. The Outbound shelf holds the runnable versions; every one of them carries the evidence through to the draft.
Frequently asked questions
What does evidence-based personalization mean?
Opening with one observable, relevant fact about the account - a posting, an announcement, a visible practice - connected to a problem you can help with, and matching the offer to the situation the evidence reveals.
What evidence is acceptable to reference in cold email?
Public-professional information the company chose to publish or visibly does: postings, funding news, launches, tech stack, live ads. Never an individual’s browsing, and intent data only shapes timing and offer - it never appears in copy.
Why not mention intent data directly?
"We noticed your company researching X" reads as surveillance even when aggregate, and it burns trust for every future touch. Let intent move the account up the queue and pick the offer; let published evidence carry the opener.
How does this scale beyond a few accounts a day?
With missions that carry each account’s evidence file into drafting and park batches for approval - review replaces writing. What does not scale honestly is templates with model-guessed personalization blanks.
What reply rates should evidence-first outbound expect?
Multiples of token-template baselines, with the leverage in signal freshness and offer-match rather than copy cleverness. Measure per signal family, and let the funnel math - not anecdotes - set expectations.
Sources
- Google - Email sender guidelines (engagement and complaint thresholds)
- M3AAWG - anti-abuse best practices for senders
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.