SF-Data-Arch Master Data Management Practice Question
A global retailer is consolidating customer master data from Salesforce, a legacy loyalty system, and an e-commerce platform. The data architect must design a matching strategy that minimizes false positives while still identifying the same customer across sources where names are spelled differently and addresses vary. Which matching approach should the architect recommend?
⚠ Common exam trap
The trap here is assuming that exact matching on one or a few fields is safer, when in cross-source consolidation it creates false negatives, and that fuzzy name matching alone is sufficient, when it actually increases false positives.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Probabilistic matching with a configured match threshold and score bands, supplemented by deterministic rules for high-confidence identifiers like loyalty number.
Cross-source customer matching with inconsistent names and addresses requires probabilistic matching over multiple attributes with a tuned threshold, because exact keys alone miss too many true matches. Adding deterministic rules for unique identifiers like loyalty number anchors high-confidence matches. This hybrid strategy controls false positives through scoring while preserving recall, which is exactly what the scenario demands.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Exact matching on the combination of first name, last name, and postal code.
Why it's wrong here
Requiring exact agreement on three fields is brittle: nicknames, middle names, and address changes cause misses, and common names in dense areas can still collide. It does not handle the stated variation in names and addresses, so it produces both false negatives and potential false positives. This deterministic composite key is too rigid for cross-source customer matching.
- ✗
Fuzzy matching on the full name field alone using a Levenshtein distance threshold.
Why it's wrong here
Name-only fuzzy matching ignores other attributes and can merge different people with similar names, especially in large populations. It also fails when names are recorded in different orders or scripts. Without additional attributes or a threshold strategy across fields, false positives rise. This does not meet the requirement to minimize false positives while handling address variation.
- ✗
Deterministic matching on exact email address only, treating any non-match as a distinct customer.
Why it's wrong here
Exact email matching is deterministic and low risk for false positives, but it fails when customers use different emails across channels or when email is missing, producing many false negatives. The requirement is to match across sources where names and addresses vary, so a single exact key cannot achieve the needed recall. This approach would leave duplicates unmerged and is not appropriate.
- ✓
Probabilistic matching with a configured match threshold and score bands, supplemented by deterministic rules for high-confidence identifiers like loyalty number.
Why this is correct
Probabilistic matching compares multiple attributes with weights and tolerates variation in names and addresses, while a threshold controls false positives. Adding deterministic rules for unique identifiers such as loyalty number captures exact matches with certainty. This hybrid approach balances precision and recall, directly addressing the need to minimize false positives while still matching records across sources with inconsistent spelling and addresses.
About these practice questions
Courseiva writes every SF-Data-Arch question from scratch — 222 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Salesforce exam blueprint
This SF-Data-Arch practice question is part of Courseiva's free Salesforce certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SF-Data-Arch exam.