Skip to content

An independent study reference written by Dr Phuc V. Nguyen. It is not official subject material — for assessment requirements always follow your subject outline and vUWS.

Data Foundationsintermediate

Data quality dimensions

Quality becomes manageable once it is broken into dimensions that can each be measured. The common set is completeness (are required values present), accuracy (do values match an authoritative reference), consistency (do related values agree with each other and across systems), timeliness (is the data current enough for the decision), uniqueness (is each real-world entity represented once) and validity (does the value obey its format or code list). Each becomes a percentage computed by a stated rule on a named field, so quality moves from an opinion to a monitored number, and a failing threshold points at a specific defect rather than a general complaint.

Try it yourself

Quality dimensions, field by field

Each dimension is a percentage from a stated rule on a named field, with its grain and its denominator shown. Switch rules on and off and watch which measures move. The blend at the end is the number an executive gets, and it cannot tell anyone what to do.

Order records included 12
Rows in the customer table 12
Quarantined rows 2
Rows with an unresolved defect 7
Blended score91.7%
Validity, postcode100.0%Validity, order date100.0%Validity, order amount100.0%Completeness, mobile number91.7%Consistency, postcode across systemsnot assessed—Uniqueness, customer id91.7%Timeliness, captured at91.7%Accuracy, industry code66.7%Accuracy, mobile, postcode, amountnot assessed—blend 91.7%
DimensionFieldGrainScoreDenominator and rule
Validitypostcodeorder records100.0%12 of 12 order records, in the reference list of postcodes
Validityorder dateorder records100.0%12 of 12 order records, between 1 January 1990 and the as-of date
Validityorder amountorder records100.0%12 of 12 order records, the row records a currency and the amount is written as A$0,000.00
Completenessmobile numbercustomers91.7%11 of 12 customers, a non-blank value on the customer row
Consistencypostcode across systemsorder recordsnot assessedthe rule would compare each order postcode with the postcode held for the same customer in a second operational system, over the 12 order records in the working table. This fixture holds one system, so there is nothing to compare, and two systems agreeing would show only that they agree
Uniquenesscustomer idcustomers91.7%11 of 12 customers, 1 minus extra copies of a customer already present, over rows in the customer table
Timelinesscaptured atorder records91.7%11 of 12 order records, age at or under 90 days on 31 March 2026
Accuracyindustry codecustomers66.7%6 of 9 customers, matches the business register extract, over the customers the extract covers
Accuracymobile, postcode, amountnot applicablenot assessedno authoritative reference for these values exists here, and each would need its own reference and its own denominator, so this is reported rather than scored
Rules applied to the raw feed
Missing mobile numbers are always left unresolved. No rule here fills a blank, because a filled blank is an invented fact that everything downstream then treats as evidence. The format rule works the same way. It may rewrite how an amount is written only where the row records its own currency, which every row in this feed does, and a bare numeral on a feed that records no currency would stay flagged, because AUD 1,450 and USD 1,450 are different facts.
Timeliness tolerance (days)90 days
Audit log, 10 entries
ORD-1008 STANDARDISE Standardise amount. arrived as "1450" with no currency marker in the amount string, and the row records its currency as AUD in a field of its own, so the rule rewrites the presentation to A$1,450.00 from a recorded fact and the amount itself is unchanged
ORD-1009 QUARANTINE Postcode reference check. postcode 9999 is not in the reference list, row held out of the working table for review, not deleted and not guessed
ORD-1007 QUARANTINE Order date range check. order date 1 January 1900 is a well-formed date and outside the plausible range, so a format rule passes it and this one does not
ORD-1001 LEAVE UNRESOLVED Customer link. pair scored 0.405, under the review band floor of 0.58, so nothing merges and nobody looks at it. The second record of this customer stays in the working table and uniqueness carries the extra copy. At this threshold no pair of different people is merged anywhere in the candidate file.
ORD-1002 LEAVE UNRESOLVED Customer link. pair scored 0.405, under the review band floor of 0.58, so nothing merges and nobody looks at it. The second record of this customer stays in the working table and uniqueness carries the extra copy. At this threshold no pair of different people is merged anywhere in the candidate file.
ORD-1004 FLAG First-option concentration. Agriculture is first alphabetically and holds 25.0% of the working table, above the 20.0% limit, and the value is left exactly as it is
ORD-1005 FLAG First-option concentration. Agriculture is first alphabetically and holds 25.0% of the working table, above the 20.0% limit, and the value is left exactly as it is
ORD-1006 FLAG First-option concentration. Agriculture is first alphabetically and holds 25.0% of the working table, above the 20.0% limit, and the value is left exactly as it is
ORD-1010 FLAG Freshness check. captured 229 days before 31 March 2026, past the 90 day tolerance, a refresh policy is the fix and no rule can invent a fresh value
ORD-1003 LEAVE UNRESOLVED Missing mandatory field. mobile number is blank, no rule here may invent one, so the gap stays visible in the count
The blend is 91.7%, an unweighted mean of 7 ratios that count different things, 4 of them over order records and 3 over customers. Underneath it, accuracy on the industry code is 66.7% while validity on the postcode sits at 100.0%. Those two numbers call for different work, and the blend asks for none in particular. Consistency is not assessed, because it compares two systems and this fixture holds one. Accuracy on mobile, postcode and amount is not assessed either, because no authoritative reference for those values exists here, and reporting that is more honest than scoring it 100%.
The 14 orders are a stylised fixture written for this widget, not a real extract, and the as-of date is fixed at 31 March 2026. Every row records its own currency, which is what lets a format rule rewrite an amount without assigning one. The business register extract covers 9 of the 14 order records, which are 8 of the 13 customers behind them. The linkage weights, email 0.55, name 0.30 and postcode 0.15, are a stated choice rather than an estimated model, and the fixture also declares which candidate pairs are truly one customer, which is the only reason the merge decisions above can be scored at all.

Why it matters

Saying the data is bad is like saying a car is bad. Useless. Is the fuel gauge wrong, are two tyres flat, or has the registration expired? Each has a different fix and a different urgency. The dimensions are the separate gauges. Once you can say that completeness on mobile numbers is 92 per cent and falling, you have something a person can act on straight away.

Before you read on — recall

A postcode field passes every validation rule, and every address involved was confirmed with the customer in the past seven days, yet parcels are being misdelivered. Which dimension is most likely failing?

Formulas

Completeness
C=npopulatednrequired×100%C = \dfrac{n_{\text{populated}}}{n_{\text{required}}} \times 100\%
Measured on a named field, across the rows where that field is genuinely required. Of 50,000 customer records, 4,100 have no mobile number, so completeness on that field is 45,900 divided by 50,000, about 91.8 per cent. Completeness across a whole table means nothing, because the fields are not equally required.
Uniqueness
U=(1−nduplicatentotal)×100%U = \left(1 - \dfrac{n_{\text{duplicate}}}{n_{\text{total}}}\right) \times 100\%
If 620 of 50,000 customer rows are extra copies of an entity already present, uniqueness is about 98.8 per cent. The figure depends entirely on the matching rule used to decide two rows are the same person, so the rule belongs in the published definition of the measure.
Timeliness as currency against a tolerance
A=tnow−tupdated  ≤  τA = t_{\text{now}} - t_{\text{updated}} \;\le\; \tau
Age AA is the gap between now and the last update, and τ\tau is the tolerance the decision allows. A price list refreshed nightly has an age of up to 24 hours, which is fine for a periodic margin review and unacceptable for a live pricing engine. Timeliness only means something against a stated tolerance.

Worked examples

Scenario

A monthly data quality report shows an overall quality score of 94 per cent. The executive sponsor asks what to do about the missing 6 per cent.

Solution

A blended score cannot answer that, because it hides which dimension failed and on which field. Splitting it shows validity at 99.6 per cent, completeness on mobile number at 91.8 per cent, and uniqueness on customer at 98.8 per cent. Those imply three different responses: a change at the point of capture, a matching and merge exercise, and no action at all for validity. Scores are useful for tracking a trend, and the action always comes from the dimension and field underneath.

Scenario

A customer record shows a date of birth of 1 January 1900. The validation rule passes it.

Solution

The value is valid, because it is a well-formed date inside the accepted range, and it is not accurate, because nobody in the file was born then. It is almost certainly a default keyed to get past a mandatory field. Validity checks format and code lists and is cheap to automate. Accuracy needs comparison against something authoritative, which is expensive, so a practical compromise is to monitor suspicious concentrations. One date holding four per cent of all records is a signal no format rule will ever raise.

Common mistakes

  • ✗Accuracy and validity are the same thing. Validity asks whether a value obeys its format or code list, accuracy asks whether it matches reality, and a well-formed wrong value passes validity easily.
  • ✗A single overall quality score is enough to manage by. Blended scores hide which dimension and which field failed, and the action always comes from that detail.
  • ✗Completeness means no nulls anywhere. A null is only a defect where the field is required for the use in question, and forcing values into optional fields manufactures data that is worse than a null.
  • ✗Timeliness is about how fast the pipeline runs. It is the age of the data relative to the decision, so a slow pipeline that delivers a monthly figure on the second of the month is perfectly timely.

Revision bullets

  • •Completeness, accuracy, consistency, timeliness, uniqueness, validity
  • •Each dimension is a percentage computed by a stated rule on a named field
  • •Validity is cheap to automate, accuracy needs an authoritative reference
  • •Uniqueness depends on the matching rule, so publish the rule with the number
  • •Timeliness is age against a decision tolerance, not pipeline speed
  • •Blended scores hide the field and dimension that actually need action

Quick check

A postcode field passes every validation rule, and every address involved was confirmed with the customer in the past seven days, yet parcels are being misdelivered. Which dimension is most likely failing?

A team reaches 100 per cent completeness on the industry field after making it mandatory at data entry. Sales analysis by industry then gets worse. Why is that plausible?

Connected topics

More in Data Foundations

Sources

  1. Wang & Strong (1996)
    Wang, R. Y., & Strong, D. M. "Beyond Accuracy: What Data Quality Means to Data Consumers." Journal of Management Information Systems, 12(4), 5-33, 1996.
    Derives and groups data quality dimensions empirically rather than asserting them.
  2. ISO/IEC 25012:2008
    International Organization for Standardization. ISO/IEC 25012:2008, Software engineering: Software product Quality Requirements and Evaluation (SQuaRE), Data quality model.
    A standardised data quality model separating inherent characteristics from those that depend on the system holding the data.
  3. DAMA-DMBOK (2017)
    DAMA International. DAMA-DMBOK: Data Management Body of Knowledge, 2nd ed. Technics Publications, 2017.
    Treats dimensions as things to be defined, measured and monitored against thresholds rather than described.
How to cite this page
Dr. Phil's Quant Lab. (2026). Data quality dimensions. Business Analytics Atlas. https://phucnguyenvan.com/analytics_atlas/concept/ba-data-quality-dimensions
Next concept
Data quality as fitness for use
Built by Dr. Phuc V. Nguyen ·Follow on LinkedInWork with PhilEmail