Data quality / FIELD GUIDE

What is a Confidence score?

A confidence score expresses a system’s assessed support for a prediction, match or extracted value under a particular method. It is not automatically a calibrated probability that the result is correct.

Key takeaways

  • A confidence score summarizes a model or rule’s support for a particular claim.
  • Its scale and meaning must be defined before it can guide action.
  • Confidence in a match is different from lead value, data freshness or purchase probability.

Overview

Scores can come from rules, model outputs or statistical estimates. Their meaning depends on how they were produced and validated. A score of 0.9 from one system is not necessarily comparable to 90 from another. Calibration asks whether outcomes assigned a given probability are correct at roughly that frequency over suitable test data.

How it works

  1. Define what the score refers to and which evidence produces it.

  2. Validate score bands against labeled outcomes representative of real use.

  3. Set action and review thresholds based on the cost of errors.

Name the claim the score refers to

A score might describe whether two company records match, whether an extracted field is supported or whether an email classification is reliable. Those are different targets. A single confidence label on an entire contact can conceal that identity is well supported while current employment is uncertain.

Record the score’s producer, version and interpretation. Some scores are weighted heuristics; others are model outputs or calibrated probabilities. A number between zero and one is not automatically a probability. Review the documentation and validation evidence before presenting the score to users as a percentage chance that a claim is correct.

Scores that should not be conflated
ScoreQuestion it may answerDifferent question
Match confidenceDo these records identify the same entity?Is the account commercially attractive?
Extraction confidenceDoes the source support this field?Is the source current enough?
Lead scoreHow should this record be prioritized?Is every underlying fact correct?
Outcome probabilityHow likely is a defined event?Has that event already been confirmed?

Check whether the numbers mean what they suggest

If a system claims calibrated probabilities, review groups of predictions against observed correctness under a reliable reference process. In an illustrative group of 100 matches assigned about 0.8 probability, roughly 80 correct matches would be consistent with that claim, subject to sampling uncertainty. One small group is not enough to establish calibration across every segment.

For a heuristic score, evaluate ranking and threshold behavior instead of demanding probability semantics it never claimed. A useful score can order likely matches without expressing an exact probability. Make that clear in the interface so users do not convert an internal rank into a false sense of certainty.

Set thresholds according to the consequence

A candidate suggestion can use a lower threshold than an automatic merge of customer records. Define the cost of a false acceptance and a missed match, then review representative cases around the proposed cutoff. Keep a middle band for human review when that helps manage uncertainty without discarding useful candidates.

Do not let a high score bypass hard permissions or required evidence. A model’s confidence cannot authorize an export or repair a missing source. Monitor the accepted, rejected and reviewed groups after deployment, and reassess thresholds when the source population or model changes. Confidence is a decision aid, not a substitute for the workflow’s rules.

ILLUSTRATIVE EXAMPLE

What this looks like in practice

A company matcher labels a shared domain and registration number as strong evidence, while a name-only match receives a lower score and goes to review before any CRM merge.

Examples explain the concept; they are not reported customer results.

What to check

Check calibration where probability claims are made, as well as precision at the action threshold. Inspect performance on ambiguous and out-of-distribution records.

Common mistake

Displaying a model’s self-reported confidence as a precise probability without measuring its relationship to correct outcomes.

Confidence score vs. Lead scoring

A confidence score concerns support for a result. A lead score concerns business priority or fit. A high-priority lead can still have low-confidence contact information.

Read the Lead scoring definition →

Questions answered

What is a Confidence score?

A confidence score expresses a system’s assessed support for a prediction, match or extracted value under a particular method. It is not automatically a calibrated probability that the result is correct.

Can scores from different providers be compared?

Only after understanding their definitions and validating them on a common dataset. Identical numeric ranges do not imply identical meaning.

Should a high confidence score bypass approval?

No. Confidence and authorization are different. A correct prediction can still propose an action that requires permission or review.

Can a language model’s self-reported confidence be trusted?

Treat it as an unvalidated signal unless you have evidence that it predicts correctness for the task and population. A fluent explanation or high self-rating can accompany an error. Prefer observable evidence, calibrated evaluation where appropriate and explicit handling of unsupported claims.

Should confidence scores be shown to users?

Show them when they help a meaningful decision and the scale can be explained. Pair the score with reasons, evidence and an action such as review. A precise number without a defined target or validation can mislead more than a clear status such as confirmed, uncertain or unresolved.

References and further reading

Primary documentation and source material for this topic. Sources checked September 14, 2026; provider requirements can change.

  1. AI Risk Management FrameworkNIST
  2. What is AWS Entity Resolution?AWS

Continue reading on the blog

Explore all articles and guides →

Put the concept to work.

Explore the relevant AstroFabric workflow and see how the pieces connect.

Help keep this guide useful. Suggest a correction or browse the full glossary.