Good data hygiene is a scheduled routine with named owners, not a one-time cleanup project. The routine below runs in six passes: audit, normalize, duplicate review, suppression, missing-field fill, and freshness check. Each pass has a defined output, and the key discipline is measuring quality against the records that survive cleanup, not the raw count you started with. A database that looks 90 percent complete before deduplication can look very different once duplicates and suppressed records are removed from the denominator.
The reason this matters is simple: metrics computed on dirty denominators hide problems. If 60 of your contacts have blank emails but 100 of your records are duplicates or suppressed anyway, the blank-email rate you should care about is the one among usable records. The worked example below makes this concrete.
Why hygiene fails as a project and works as a routine
A familiar failure mode is the quarterly heroic cleanup. Someone exports the CRM, spends a week in spreadsheets, imports the fixes, and declares victory. Three months later, decay has undone the work: people changed jobs, companies renamed themselves, new imports arrived with inconsistent formatting, and nobody owned the interval between cleanups.
A routine fixes this by shrinking each pass until it fits inside a sustainable cadence. Instead of cleaning 50,000 records once a quarter, you clean the records touched or created since the last pass, plus a rotating slice of the backlog. The backlog can shrink steadily, and new debt gets retired on a rhythm instead of piling up between heroic efforts.
Understanding where each field value came from also matters more than teams expect. When two records disagree about a job title, you cannot resolve the conflict without knowing which source supplied each value and when. That is a data provenance question, and it is worth capturing source and timestamp metadata during enrichment rather than reconstructing it later.
The six-pass routine
Pass 1: Audit
Pull counts before touching anything. Total records, records by lifecycle stage, records with blank values in each critical field, records not updated in a defined window, and candidate duplicates by your matching keys. This snapshot is your baseline; every later pass reports against it.
Pass 2: Normalize
Standardize formats so later matching works: country and state values to a single convention, phone numbers to one format, company names trimmed of legal suffixes in a normalized shadow field (keep the original), domains lowercased and stripped of protocols. Normalization is low-risk and reversible if you write normalized values to dedicated fields rather than overwriting originals.
Pass 3: Matched duplicate review
Generate duplicate candidates using stable keys: email for contacts, domain for companies, and fuzzy name-plus-location as a secondary signal. Then review before merging. Merges blend activity histories, deal associations and ownership, and unwinding a bad merge is painful enough that you should treat it as irreversible in practice. Queue candidates by confidence, auto-approve nothing that touches an open opportunity, and log every merge decision.
If you sync with an external system, preserve stable identifiers so future imports update existing records instead of spawning new duplicates. HubSpot's import documentation, for example, emphasizes matching on stable identifiers and checking current file requirements before a bulk import (HubSpot). The general lesson applies to any CRM: dedup work is wasted if the next import cannot recognize the records it should update.
Pass 4: Suppression
Remove or flag records you should not use: opt-outs, bounced addresses, competitors, records failing compliance requirements, and roles outside your market. Suppression is distinct from deletion. A suppressed record stays in the system so future imports and enrichment jobs recognize it and do not reintroduce it as a fresh contact. Treat this pass as a cleanup audit, not the enforcement layer: suppression flags should also be checked live at every send or export, not only on the pass cadence.
Note what suppression does not tell you. A mailbox that verifies as deliverable is not thereby a permitted or well-placed send. Google's sender guidelines cover authentication, reputation and unsubscribe handling as separate obligations (Google); deliverability and consent are their own workstreams, not byproducts of clean data.
Pass 5: Missing-field fill
With duplicates and suppressed records out of the way, measure completeness on what remains and fill gaps in priority order: fields that gate routing and segmentation first, nice-to-have fields last. A waterfall enrichment approach, where multiple sources are tried in sequence until a verified value is found, raises fill rates without paying for every source on every record. Record which source filled each field so the next conflict is resolvable.
Pass 6: Freshness and backlog
Tag every record with a last-verified date per critical field, not just a last-modified date, since automation can touch a record without verifying anything. Define a staleness window per field type. As an illustrative operating choice, treat titles and current employers as faster-decaying than industry or company size, and tune the windows to what you actually observe in your own data. Records past their window enter the backlog, and each cycle retires a fixed slice of it.
Worked example: measuring on the right denominator
Illustrative example. A team audits a 500-record segment before a campaign:
- 500 records in the raw segment
- 40 records identified as duplicates and merged after review: 500 - 40 = 460
- 20 records suppressed (opt-outs and out-of-market roles): 460 - 20 = 440 eligible records
- 60 of the 440 eligible records have blank email fields
The blank-email share that matters is 60 / 440 = 13.64 percent, not 60 / 500 = 12 percent. The eligible-denominator figure is higher and more honest: it tells you that roughly one in seven records you could actually use is missing its most important field. Enrichment effort now has a precise target of 60 records, and the post-fill audit can report an exact new completeness rate instead of a vague improvement claim.
Owners and cadence
The specific assignments below are illustrative; adapt them to your team's structure and volume.
| Pass | Suggested owner | Illustrative cadence | Output |
|---|---|---|---|
| Audit | RevOps | Monthly | Baseline metrics snapshot |
| Normalize | RevOps or data engineer | Weekly, on new/changed records | Normalized shadow fields |
| Duplicate review | RevOps + record owners | Weekly queue review | Merge log with decisions |
| Suppression | Marketing ops | Weekly | Updated suppression flags |
| Missing-field fill | RevOps | Monthly, priority fields first | Fill-rate report by field |
| Freshness/backlog | RevOps | Monthly rotating slice | Backlog burn-down count |
One person should own the routine end to end even when passes are delegated. Shared ownership of a maintenance loop often turns into no ownership.
Tradeoffs and failure handling
Review depth versus throughput. Reviewing every duplicate candidate is safe but slow. A reasonable compromise is confidence tiers: exact-key matches on records with no open deals can move faster, while fuzzy matches and deal-attached records always get human eyes. That policy is something your team enforces operationally; no tool enforces judgment for you.
Over-suppression. Aggressive suppression rules can quietly remove valid records. Audit a sample of suppressed records each quarter and record the suppression reason in a field your team configures and validates so mistakes are findable.
Enrichment conflicts. When a new value contradicts an existing one, do not silently overwrite. Prefer the source with stronger provenance and a more recent verification date, and log the decision. This is where the source-and-timestamp metadata from Pass 5 pays off; see the data enrichment glossary entry for how enrichment layers interact with existing values.
Backlog stall. If the freshness backlog stops shrinking, the slice size is too ambitious or the cadence too optimistic. Cut the slice and keep the rhythm. A small pass that happens beats a large pass that slips.
Where autonomous agents fit
The repetitive parts of this routine, candidate detection, normalization, multi-source field fill, and freshness monitoring, are exactly what autonomous agents handle well when the pipeline feeds your existing systems rather than a separate dashboard. AstroFabric's agents can enrich and verify records against multiple data types and stream structured results into your CRM with approval-gated writes, so the human review steps above stay human. If you are building the surrounding pipeline, the data infrastructure for prospecting overview shows how discovery, verification and delivery connect.
Run one audit pass this week on a single active segment, compute your eligible-record completeness rate, and you will know exactly where the next hour of hygiene work should go. When you are ready to automate the fill and verification steps, sign up for AstroFabric and point an enrichment playbook at that segment.
FAQ
How often should a CRM data hygiene routine run? Match cadence to decay speed and field importance. A light weekly pass on active segments plus a deeper monthly or quarterly full cycle works for many teams. Sustainability beats ambition: pick the shortest cadence your owners can actually keep.
Should duplicate merges ever be automated without review? Automated candidate detection is fine; automated merging of ambiguous matches is where damage happens. Treat merges as effectively irreversible, require review for anything touching open deals or active sequences, and keep a merge log.
Does verifying an email address mean the contact is safe to email? No. Verification estimates whether a mailbox is likely to accept messages, and results can be inconclusive. It is not inbox placement, not proof the address belongs to the person you expect, and not permission to send. Consent and sender reputation are separate obligations.
Sources
Every playbook on this blog ships as a runnable mission.
Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.