B2B Data Providers: A Buyer's Evaluation Playbook

How to evaluate B2B data providers on usable, ICP-fit records instead of database size, with a 100-record pilot method and a usable-record cost worksheet.

GuideBY THE ASTROFABRIC TEAM · SEP 14, 2026 · 8 MIN READ

The right way to compare B2B data providers is to measure the cost of a usable, accepted record for your specific ICP, not the size of the vendor's database or the headline price per record. A large database with a low sticker price can still be expensive when few returned records pass your verification and fit checks. This playbook shows how to run a controlled 100-record pilot, how to read the nested counts it produces, and how to compare providers on source breadth, identity resolution, freshness and delivery.

Why database size is the wrong yardstick

A large database can help with breadth, but its total size does not show coverage of your particular market. A specialist database can also cover a narrow ICP well. Check three things behind the headline count.

First, overlap. Providers may draw on overlapping sources. Measure how many additional accepted records a second provider adds within the segments you need. Second, staleness. A database can be enormous and still carry job titles that are two role changes old. Third, relevance. Coverage of the global economy is irrelevant if your ICP is 400 mid-market logistics companies in North America. What you need is depth exactly there.

A pilot on your target accounts helps evaluate these issues together. Pair the results with a review of contracts, delivery requirements and ongoing operating costs.

The four dimensions that actually separate providers

Before running numbers, decide what you are grading. These four dimensions give the pilot a practical starting point.

Source breadth. Does the provider combine multiple data types, or is it strong in one category and thin elsewhere? For account-level work you may need firmographic, technographic, hiring and funding data together. A provider that nails contact data but has no company-level signal layer forces you to stitch vendors, which is fine if you plan for it and painful if you do not. Understanding data provenance matters here: a vendor that can explain where a field came from is easier to audit than one that cannot.

Identity resolution. When you submit a company domain and a target role, does the provider return the right person at the right legal entity, or a name-alike at a subsidiary in another country? Entity confusion can send a correct-looking value to the wrong record. Check entity keys automatically and manually review an ambiguous sample.

Freshness. Ask how the provider detects change, not how often it claims to refresh. Role changes, layoffs and acquisitions invalidate records continuously. A vendor that re-verifies on access behaves differently from one that batch-refreshes quarterly.

Delivery. Data that cannot reach your systems cleanly is a liability. Look for API access, webhook delivery, idempotent writes and the ability to land records in your CRM and operational sheets rather than in another isolated interface. If your team runs waterfall enrichment across multiple sources, delivery mechanics decide whether the waterfall is automatable at all.

The 100-record pilot, step by step

Here is the method, with an illustrative example throughout. All numbers below are hypothetical and exist to show the arithmetic, not to describe any real vendor.

Pull 100 accounts or contacts at random from your genuine ICP definition. Do not let the vendor supply the sample, and do not cherry-pick well-known companies, because an easy segment can hide gaps in the customers you actually target. Submit the same 100 records to each provider under evaluation, then score the results through three nested gates:

  1. Matched. The provider returned a record for the input at all.
  2. Verified. The returned record passed your verification checks, such as mailbox verification for emails and a manual spot check on titles.
  3. ICP fit. The verified record actually matches your targeting criteria: right role seniority, right company segment, right geography.

Suppose the pilot returns these results:

GateCountRate vs previous gateRate vs total
Submitted100100%
Matched8282.0%82%
Verified7085.4% (70 of 82)70%
ICP fit6288.6% (62 of 70)62%

The counts are nested: every ICP-fit record is also verified, and every verified record was matched. That nesting is the point. It tells you where quality leaks. In this example, 18 records never matched, 12 matched records failed verification, and 8 verified records turned out not to fit the ICP, often because the person had changed roles or the company had been misclassified. Each leak has a different fix. A low match rate warrants checking both input identifiers and coverage. Verification failures may reflect stale addresses or inconclusive checks. Fit failures call for a review of targeting rules, classification and entity matching; the counts alone do not identify the cause.

Usable-record cost: the number that decides the deal

Now attach money. In this hypothetical pricing model, the provider charges $0.31 per input record submitted, including inputs that do not produce an accepted result. The pilot cost 100 × $0.31 = $31.00. The usable output was 62 records, so the usable-record cost is $31.00 ÷ 62 = $0.50 per record your team can actually work.

Compare that with a second illustrative vendor at $0.20 per record. Cheaper on paper. But if its pilot yields only 32 ICP-fit records, the math flips: 100 × $0.20 = $20.00, and $20.00 ÷ 32 = $0.625 per usable record. The vendor with the 55 percent higher input price is 20 percent cheaper per accepted record in this example, and the downstream costs of the low-yield vendor, wasted outreach, rep time and CRM pollution, are not even in that figure yet.

A worksheet for your pilot:

LineVendor AVendor B
Records submitted100100
Price per record
Total pilot cost
Matched / verified / ICP fit
Usable-record cost (total ÷ ICP fit)
Fields delivered per usable record
Delivery path into your systems

Conditional recommendations by workflow

There is no universal winner, only fits. If your motion is rep-driven prospecting where sellers search and pull contacts themselves, an integrated search tool such as Apollo, which offers company and person search for B2B prospecting (Apollo), fits the daily workflow. If your team builds bespoke enrichment logic and wants hands-on control of each step, a workflow workspace such as Clay, which describes enrichment, agents, signals and API interfaces (Clay), fits builders who enjoy assembling the machine.

If you want the objective-to-dataset motion, where you describe a target and autonomous agents discover, verify, enrich and score records, then stream them into your existing CRM, sheets and channels, an agentic layer like AstroFabric fits teams treating data as infrastructure for prospecting rather than a manual task. Run the same pilot regardless of category; the scoring method does not change.

Tradeoffs and failure handling

Pilots have limits worth respecting. A 100-record sample carries statistical noise, so do not call a 62 versus 65 result a reliable difference without examining the paired records and testing a larger sample. Vendors may perform differently across your segments; if you sell into three distinct ICPs, stratify the sample rather than blending them. And a pilot measures a moment in time, so revisit the numbers quarterly, because provider quality drifts.

When a pilot fails, diagnose before you disqualify. A poor match rate on an unusual ICP may mean no vendor covers it well and you need a custom discovery approach. A poor verification rate might trace to your own stale input list rather than the vendor. And never merge pilot data into production CRM records without an approval step, because merges are hard to unwind cleanly. Suppress duplicates, gate the writes, and keep an audit trail of what came from where. Contact enrichment that skips these controls creates cleanup work that erases the savings.

FAQs

How many records should a pilot include? Enough to expose variance without heavy cost. A suggested exploratory starting point is 100 to 300 records drawn from your real ICP; it is not a statistically validated sample-size rule. Small samples make rates unstable; one lucky batch of 25 can hide serious gaps.

Should I run the same sample against every provider? Yes. The sample, field requirements and acceptance rules must be identical across vendors, or the comparison is meaningless. Score everyone against the same definition of a usable record.

Does an email verification pass mean the contact is safe to reach out to? No. Mailbox verification estimates an address's current mail-acceptance status; a pass is not a delivery guarantee. It does not prove the person still holds the role, that you have permission to contact them, or that they have buying interest. Treat it as one gate among several.

Run the pilot on your own ICP

Pick your next renewal or evaluation, pull 100 real target records, and score every candidate provider through the matched, verified and ICP-fit gates before you sign anything. If you want autonomous agents to handle the discovery, verification and enrichment behind that funnel and deliver the results into your existing systems, start with AstroFabric and run the same 100-record test.

Sources

⟨ RUN IT INSTEAD OF READING IT ⟩

Every playbook on this blog ships as a runnable mission.

Open a workspace and the playbook library is waiting - describe the outcome and the agents carry it end to end, on your plan's monthly credits.

⟨ KEEP READING ⟩
ComparisonEnrichment & data

Clay Alternatives: Choose the Right Data Workflow

How to evaluate Clay alternatives by workflow ownership, composition model, data scope and tool fit, with a scored rubric and a worked example.

Aug 13, 2026 · 7 min read
GuideEnrichment & data

Waterfall Enrichment: Provider Order and Stop Rules

Order enrichment providers by marginal cost per accepted record, set stop rules, and track per-stage provenance so waterfall enrichment stays cost-disciplined.

Sep 1, 2026 · 8 min read