Data quality / FIELD GUIDE

What is Data quality?

Data quality is the degree to which data is fit for its intended use, considering properties such as accuracy, completeness, consistency, timeliness, validity and uniqueness.

Key takeaways

  • Quality is fitness for a specific use, not a universal label on a dataset.
  • Measure the dimensions that can change the outcome and retain their separate results.
  • Fix defects at their source and check the final destination, where errors become operational.

Overview

Quality is contextual. A quarterly company-size estimate may support market segmentation but be unsuitable for a precise headcount report. Evaluate the fields and thresholds that matter to the decision. A dataset can be complete and still be inaccurate, or syntactically valid and still describe the wrong entity. Avoid collapsing every dimension into an unexplained overall percentage.

How it works

  1. Define the intended use and the fields that can change the outcome.

  2. Create measurable checks for the relevant quality dimensions.

  3. Sample results, fix root causes and monitor quality as sources or rules change.

Define a quality contract for the workflow

A list used for broad market research can tolerate uncertainty that would be unacceptable in a personalized message. Start by naming the intended decision and the fields that can change it. For account routing, company identity and territory may be critical. For outreach, current employment, channel status and suppression checks may also be required.

Turn those requirements into observable checks. Define what passes, what needs review and what should be excluded. Do not let a strong result on an easy field compensate for failure on a critical one. A record with a perfectly formatted phone number still fails if it belongs to the wrong person.

Illustrative quality checks for a prospect handoff
DimensionQuestionExample failure
AccuracyDoes the value describe the intended entity?A role belongs to a namesake
CompletenessAre the required facts present?No account identifier for delivery
TimelinessIs the evidence current enough?Employment was last observed years ago
UniquenessDoes one entity appear once where intended?The same person is assigned twice
ValidityDoes the value follow the required format?An invalid country code blocks the import

Use a sample that can reveal the real defects

Sample across sources, regions, company sizes and match difficulty. A random sample is useful for an overall estimate, while targeted samples help investigate known failure modes. Keep those purposes separate when reporting results. A deliberately difficult sample should not be presented as a representative market accuracy rate.

Record the reference used to judge correctness and allow unresolved cases. If the reviewer cannot establish the current employer, that is not automatically a correct or incorrect result. Repeat a subset with another reviewer when judgment is involved so disagreements in the reference process do not masquerade as provider defects.

Repair the cause and watch the handoff

An illustrative audit finds 30 duplicate contacts among 500 records. Merging them repairs the current dataset, but the problem will return if the import still creates a new record on every retry. Link each defect category to a root cause: collection, matching, transformation, synchronization or stale evidence. The durable fix often belongs upstream of the cleanup job.

Track the rate of rejected destination records, human corrections and recurring defects after the change. Keep quality results by field and segment so an improving overall average cannot conceal a new problem in one territory. A data-quality program succeeds when decisions and handoffs become more reliable, not merely when a dashboard turns green.

ILLUSTRATIVE EXAMPLE

What this looks like in practice

An account list has every email field populated, but a sample reveals obsolete employers and duplicate people. Its completeness is high while accuracy and uniqueness fail the outreach requirements.

Examples explain the concept; they are not reported customer results.

What to check

Report quality by field, source and segment. Include sample size, test method and uncertainty when presenting an accuracy claim.

Common mistake

Calling a dataset “99% accurate” without explaining what was tested, which fields were included or how the sample was selected.

Data quality vs. Data completeness

Completeness measures whether required information is present. Quality is broader: present values must also be suitable, correct and timely for the task.

Read the Data completeness definition →

Evidence and context

6 dimensions

IBM describes six commonly used dimensions: accuracy, completeness, consistency, timeliness, validity and uniqueness. A complete dataset can still fail the other checks.

Source: IBM

Questions answered

What is Data quality?

Data quality is the degree to which data is fit for its intended use, considering properties such as accuracy, completeness, consistency, timeliness, validity and uniqueness.

Can data quality be one score?

A summary score can help monitoring, but retain the underlying dimensions and thresholds. Otherwise a strong result in one dimension may hide a critical failure in another.

How often should quality be checked?

Check at ingestion and important handoffs, then monitor on a cadence that reflects how quickly the fields change and how consequential their use is.

What does a 99% data-quality claim mean?

It is incomplete without the tested fields, population, sample size, reference standard and treatment of unknowns. It might describe syntax validity, correct identities or something else entirely. Ask for the measurement definition before comparing claims from different sources or applying the figure to your own workflow.

Should low-quality records be deleted immediately?

Choose the action by defect and purpose. A malformed value may be repairable, a duplicate may need merging and an uncertain match may need review. Preserve necessary evidence and relationships during cleanup. Deleting every failed record can hide the cause and remove information needed to prevent the same defect returning.

References and further reading

Primary documentation and source material for this topic. Sources checked September 14, 2026; provider requirements can change.

  1. What is data quality?IBM
  2. Data quality dimensionsIBM

Continue reading on the blog

Explore all articles and guides →

Put the concept to work.

Explore the relevant AstroFabric workflow and see how the pieces connect.

Help keep this guide useful. Suggest a correction or browse the full glossary.