SF-Data-Arch Data Migration Practice Question
A healthcare provider is migrating 8 million legacy Patient_Visit__c records into Salesforce. Each visit references a legacy numeric provider code, and the legacy system does not expose a stable unique identifier for each provider. The Salesforce Provider__c object already has a custom external ID field named Legacy_Provider_Code__c that is populated. The architect wants to load visits without first exporting provider Salesforce IDs. Which approach should be used?
⚠ Common exam trap
The trap here is assuming that a lookup must always be populated with the 15- or 18-character Salesforce record ID, when an external ID on the parent object can be used instead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Load Patient_Visit__c records with a lookup field populated by the legacy provider code, relying on the external ID field to resolve the relationship during the load.
Relationship fields can be populated with an external ID value when the target object has that field marked as an external ID, letting Salesforce match and set the lookup during the load. Because Provider__c already has Legacy_Provider_Code__c configured as an external ID and populated, the visit load can reference the provider by that code directly, eliminating a separate ID export and join step and keeping the migration efficient at scale.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a junction object between Provider__c and Patient_Visit__c, load the junction records with both legacy codes, and let Salesforce infer the lookups.
Why it's wrong here
A junction object is used to model many-to-many relationships, but a visit belongs to one provider, so this changes the data model unnecessarily. Salesforce does not infer lookups on junction records; each lookup still needs an ID or external ID value. This adds objects, storage, and reporting complexity without solving the original requirement.
- ✓
Load Patient_Visit__c records with a lookup field populated by the legacy provider code, relying on the external ID field to resolve the relationship during the load.
Why this is correct
Salesforce resolves relationship fields by matching the value supplied in the relationship field against the target object's external ID field, so loading visits with the legacy provider code in the Provider__c relationship maps each visit to the correct Provider__c record without a prior ID export. This is exactly the pattern Bulk API 2.0 and Data Loader support for external ID lookups, and it scales for multi-million record loads.
- ✗
Load Patient_Visit__c records into a staging custom object, then use a scheduled Apex job to copy each legacy provider code into a text field and later reconcile manually.
Why it's wrong here
Staging plus scheduled Apex adds latency, consumes Apex CPU and SOQL limits, and still leaves the relationship unresolved until a second pass runs. It does not actually use the existing external ID field to establish the lookup during load, and a manual reconciliation step is not viable for 8 million visits. This introduces unnecessary custom code and operational risk compared with a direct external ID lookup load.
- ✗
Export all Provider__c records with their Salesforce IDs, VLOOKUP the IDs into the visit file, then load visits with the 18-character Salesforce ID in the lookup field.
Why it's wrong here
This approach works but is unnecessary given that Legacy_Provider_Code__c is already an external ID on Provider__c. Exporting and VLOOKUP-joining 8 million rows is error-prone, increases file size, and adds a manual step. External ID relationship loading is the supported, more reliable mechanism when a populated external ID already exists on the target object.
About these practice questions
Courseiva writes every SF-Data-Arch question from scratch — 222 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Salesforce exam blueprint
This SF-Data-Arch practice question is part of Courseiva's free Salesforce certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SF-Data-Arch exam.