Courseiva

CCNA Data Migration Questions

38 questions · Data Migration · All types, answers revealed

1
MCQeasy

A Salesforce data architect is migrating 100,000 Opportunity records from a legacy system. The legacy data includes a custom field 'Legacy_Region__c' that must be mapped to a new picklist field 'Region__c' in Salesforce. The picklist values in Salesforce are: 'North America', 'Europe', 'Asia', 'Latin America'. The legacy data uses values like 'NA', 'EU', 'APAC', 'LATAM'. What is the most efficient way to transform the data during migration?

A.Load the legacy values as-is and then use Data Loader to update the picklist values.
B.Use a formula field in Salesforce to convert the legacy values after loading.
C.Perform the transformation in an ETL tool before loading the data.
D.Create a workflow rule to update the picklist field based on the legacy field.
AnswerC

Using an ETL tool allows the architect to map legacy values to the correct picklist values before loading into Salesforce. This ensures data integrity and avoids load errors due to invalid picklist values. The transformation can be done via lookup tables or scripts within the ETL tool. This is the most efficient and reliable method for large-scale migrations.

Why this answer

The most efficient way to transform legacy values to match Salesforce picklist values is to perform the transformation in an ETL tool before loading. This avoids load errors and ensures data quality. Formula fields, post-load updates, and workflow rules are either not possible or inefficient for this purpose.

Exam trap

The trap here is assuming that Salesforce can automatically map legacy values to picklist values or that post-load updates are efficient.

2
MCQhard

Refer to the exhibit. During a high-volume migration, the team encounters the provided error frequently. What is the most likely cause?

A.The batch size is too small, causing excessive commits to the database.
B.Multiple threads are updating child records that share the same parent account.
C.The API user does not have sufficient permissions to modify the account records.
D.The Salesforce database is experiencing a scheduled maintenance window.
AnswerB

When child records are updated, Salesforce locks the parent record to ensure data integrity. If multiple threads attempt to update children of the same parent simultaneously, they compete for the parent lock, resulting in the UNABLE_TO_LOCK_ROW error. This is a classic concurrency bottleneck in multi-threaded bulk migration processes.

Why this answer

The 'UNABLE_TO_LOCK_ROW' error indicates that multiple threads or processes are attempting to update the same record or its parent records simultaneously. In data migration, this usually happens when records sharing the same parent are processed in parallel. By reordering the data or reducing the concurrency in the bulk load, architects can prevent these deadlocks and ensure successful completion without constant retries or system instability.

Exam trap

Candidates frequently mistake locking errors for permission or data format issues, missing the core architectural problem that concurrent multi-threaded updates on shared parent records cause deadlocks.

3
MCQhard

A company needs to migrate data into a custom object while maintaining the original CreatedDate from a legacy system. Which of the following is required?

A.A custom Apex trigger to overwrite the CreatedDate field.
B.Enable 'Set Audit Fields upon Record Creation' permission.
C.Use the Salesforce Import Wizard with mapped headers.
D.Create a formula field to store the legacy date.
AnswerB

This is the only supported mechanism for populating audit system fields during record insertion. By enabling this permission on the integration user profile, the API is granted the capability to accept values for fields like CreatedDate, allowing for accurate migration of legacy audit data into the platform.

Why this answer

Maintaining the original CreatedDate requires the 'Set Audit Fields upon Record Creation' user permission. Without this, Salesforce ignores the provided CreatedDate and assigns the current server time upon insertion. This permission is a highly sensitive setting and must be enabled specifically for the integration user through the user record or the appropriate permission set, ensuring audit traceability for historical records.

Exam trap

Candidates incorrectly assume that administrative privileges or data loader settings alone can override system fields. They fail to realize that the 'Set Audit Fields' permission is a specific, separate security requirement.

4
MCQmedium

A data architect is migrating 8 million Contact records from a legacy CRM into Salesforce Sales Cloud. The legacy system stores phone numbers as free-text strings in multiple formats (e.g., '(415) 555-1212', '415.555.1212', '4155551212'). After migration, users must be able to search for Contacts by phone number and have the numbers display consistently. Which approach should the architect use to ensure phone numbers are searchable and standardized?

A.Store the original free-text phone numbers in a Text field and create a formula field that formats the number for display.
B.Store phone numbers in the standard Phone field and rely on Salesforce's automatic normalization and search indexing.
C.Create a custom Text field with a validation rule that enforces a specific phone number format, and load the raw values.
D.Pre-process the legacy data in the ETL layer to normalize all phone numbers to a single format, then load them into the standard Phone field.
AnswerD

Normalizing phone numbers in the ETL layer before loading ensures a consistent canonical format at rest, which makes the standard Phone field searchable and displayable uniformly. This approach leverages the built-in indexing of the Phone field and avoids the need for custom parsing logic in Salesforce, meeting both search and consistency requirements.

Why this answer

Normalizing phone numbers during the ETL process before loading them into the standard Phone field ensures that all values conform to one format. This makes the standard field's search indexing effective and provides consistent display without custom code. The other options either leave inconsistent data at rest or rely on features that do not automatically normalize free-text input.

Exam trap

The trap here is assuming the standard Phone field automatically normalizes or reformats inconsistent free-text values during import.

5
Multi-Selecthard

Which THREE factors are essential to consider when performing a multi-org data migration consolidation?

Select 3 answers
A.Handling duplicate record IDs across source orgs.
B.Merging disparate security models and profiles.
C.Mapping custom fields with identical names but different data types.
D.Enabling the 'Set Audit Fields' for every user.
E.Using only the standard Import Wizard for migrations.
AnswersA, B, C

Since Salesforce IDs are only unique within a single org, consolidation will lead to ID collisions. Architects must map original records to new unique external keys and cross-reference them to ensure that parent-child relationships remain intact within the new, consolidated environment after the data is migrated.

Why this answer

Consolidating multiple Salesforce orgs requires meticulous planning around unique ID collisions, metadata differences (such as custom field mapping), and security model alignment. Since each org has its own unique record IDs and potentially overlapping data, an intermediary mapping strategy is vital to ensure that relationships are preserved correctly in the new, unified environment without creating duplicate records or data conflicts.

Exam trap

Candidates often focus only on record migration while neglecting the metadata and security model. They fail to realize that differing profiles and field data types will break the target org structure.

6
MCQhard

During a migration, you encounter an error stating that the record exceeds the maximum character limit for a field. What is the most robust way to handle this?

A.Automatically truncate the text in the ETL process to fit the limit.
B.Increase the Salesforce field length to the maximum possible.
C.Evaluate the data and request business approval for truncation or re-mapping.
D.Skip the records that exceed the character limit.
AnswerC

Data integrity is paramount in any migration. The architect must investigate the source of the excess data and work with stakeholders to decide the best course of action. This ensures that the business is informed of potential data loss and approves the method used, mitigating risks and ensuring compliance with business requirements.

Why this answer

Truncating data without business approval is a data integrity risk. The architect should analyze the source data to determine why it exceeds the limit and work with stakeholders to either shorten the source text, map it to a larger field, or confirm if truncation is acceptable. Simply truncating in the ETL layer may result in the loss of critical information, which can lead to compliance or business issues down the line.

Exam trap

Candidates often choose technical shortcuts like automatic ETL truncation or script-based rejection, missing the requirement to involve business stakeholders for formal impact approval before modifying data.

7
MCQmedium

What is the main advantage of using an ETL tool rather than the standard Salesforce Data Loader for a large-scale enterprise migration?

A.It is always free to use.
B.It provides visual workflow orchestration and error handling.
C.It bypasses Salesforce security controls.
D.It allows direct database access to the Salesforce backend.
AnswerB

ETL tools offer visual design interfaces that make orchestrating complex migrations much easier than the basic, manual steps required by Data Loader. Features like automated retries and detailed logging streamline the migration process, allowing architects to handle errors gracefully and maintain a clear audit trail of all data movements.

Why this answer

Enterprise ETL tools provide advanced capabilities such as graphical orchestration, error logging, automated retry logic, and seamless connectivity to heterogeneous data sources. While Data Loader is sufficient for simple, smaller loads, ETL tools allow architects to manage complex migration workflows, parallel jobs, and sophisticated data transformations across multiple systems, which is required for large-scale enterprise integration projects.

Exam trap

Exam takers frequently choose standard Data Loader for massive enterprise migrations, forgetting that ETL tools offer critical visual workflow orchestration, error handling, and robust transformation capabilities.

8
MCQmedium

A data architect is migrating 50 million records into a custom object. To ensure optimal performance and avoid hitting governor limits, which strategy should be implemented during the initial data load?

A.Enable all existing triggers and validation rules to ensure data quality immediately upon insertion.
B.Use the SOAP API to perform synchronous inserts for real-time validation feedback.
C.Leverage Bulk API 2.0 and disable non-essential automation, triggers, and validation rules.
D.Increase the batch size to the maximum allowed limit for standard REST API calls.
AnswerC

Bulk API 2.0 is designed specifically for large data sets, providing efficient background processing. Disabling triggers and complex validation rules prevents unnecessary processing overhead, reduces CPU consumption, and avoids record locking. This is the industry-standard approach for large-scale data migrations to ensure stability and maximum performance throughout the load process.

Why this answer

Utilizing the Bulk API 2.0 with serial mode for initial loads minimizes contention on shared resources and prevents record locking errors during high-volume inserts. Designing around index selectivity and disabling unnecessary automation like workflow rules or triggers before loading is crucial for performance. This approach ensures data integrity while drastically reducing the time required to complete the migration compared to standard synchronous API methods.

Exam trap

Candidates often forget to disable automation. They attempt to load 50 million records while triggers and validation rules are active, causing massive performance bottlenecks and hitting governor limits almost immediately.

9
Multi-Selectmedium

When migrating data to Salesforce, which TWO strategies help ensure successful relationship mapping between parent and child objects?

Select 2 answers
A.Load child records first, then use Apex to link parents.
B.Use External IDs to map the relationship between parent and child.
C.Perform a two-pass load, starting with parent objects followed by child objects.
D.Query the Salesforce internal IDs after the load to update the child records.
E.Only migrate parent objects to ensure no orphan records exist.
AnswersB, C

External IDs are the standard and most reliable way to maintain relationships during migration. By mapping the legacy parent ID to an External ID field on the Salesforce parent object, the data loader can resolve the lookup field on the child record automatically without needing the internal Salesforce ID.

Why this answer

Mapping relationships is the hardest part of data migration. Using external IDs for lookups ensures that child records are connected to the correct parent regardless of the internal Salesforce ID. Furthermore, performing the load in a specific sequence (parents first, then children) eliminates 'missing parent' errors.

These two strategies ensure that relational integrity is preserved from the legacy system without relying on fragile multi-step manual work or custom post-load reconciliation scripts.

Exam trap

Candidates often attempt to load child records before parent records or rely on Salesforce internal IDs. This leads to broken relationships and failed records because the parent IDs do not yet exist.

10
MCQmedium

A company is migrating legacy account data into Salesforce. They need to map legacy system primary keys to a new custom field to support future integrations. What is the recommended approach to ensure this field supports efficient data retrieval?

A.Create a standard text field and create a custom index request through Salesforce Support.
B.Use a custom field without the External ID attribute and perform lookups via SOQL.
C.Configure the custom field as an External ID with unique constraints enabled.
D.Store the legacy primary key in the standard Name field for each account.
AnswerC

Configuring the field as an External ID automatically creates an index on the database table. This allows for rapid record matching during upsert operations, which is essential for data migrations. Additionally, the unique constraint ensures data integrity by preventing duplicate source records from being inserted into the destination object.

Why this answer

Marking the custom external ID field as 'Unique' and 'External ID' ensures Salesforce automatically indexes the column. This indexing is critical for upsert operations, as it allows the platform to quickly match existing records without performing a full table scan. Proper schema design during migration prevents long-term performance degradation and simplifies the maintenance of relational integrity across interconnected systems.

Exam trap

Candidates often select standard text fields or forget to enable unique constraints, assuming that simply checking the External ID box is enough to guarantee performance and prevent duplicates during data integration upsert operations.

11
Multi-Selectmedium

Which TWO steps are critical when preparing to migrate sensitive PII (Personally Identifiable Information) into Salesforce?

Select 2 answers
A.Perform data masking on source datasets for sandboxes.
B.Use the Bulk API for all PII data transfers.
C.Enable Field-Level Security to restrict access to sensitive fields.
D.Delete all audit logs after the migration is complete.
E.Export all PII to unencrypted CSV files for mapping.
AnswersA, C

Data masking is a mandatory security practice for non-production environments to prevent the exposure of PII. By sanitizing the data before it reaches the sandbox, the risk of data leakage is minimized while still allowing developers to test against realistic, but safe, dataset structures.

Why this answer

When migrating PII, Data Architects must prioritize compliance and data protection. Data masking ensures that non-production environments do not contain real sensitive data, while Field-Level Security (FLS) ensures that only authorized users can access the data within production. These steps are essential to maintaining regulatory compliance, such as GDPR or HIPAA, throughout the lifecycle of the data migration project.

Exam trap

Test-takers often focus solely on migration speed and technical mapping, forgetting compliance requirements like PII masking in non-production environments and proper field-level security.

12
MCQmedium

A Salesforce architect is migrating 1 million Case records with related Case Comments from a legacy system. The legacy system stores comments in a separate table linked by a legacy Case ID. The architect plans to use the Bulk API to load Cases first, then load Case Comments in a second pass. During the test, the architect realizes that the legacy Case ID is not stored in Salesforce after the first load, making it impossible to link comments to the correct Cases. What should the architect have done to enable this relationship?

A.Create a custom external ID field on the Case object to store the legacy Case ID, populate it during the Case load, and then use that field to relate Case Comments during the second load.
B.Export the Salesforce Case IDs after the first load, map them back to the legacy Case IDs in the source system, and then load Comments with the new Salesforce IDs.
C.Load Cases and Case Comments simultaneously using a single Bulk API job with nested JSON to preserve relationships.
D.Use the Salesforce Data Loader's 'Insert' operation for Cases and then use 'Update' for Comments, relying on Salesforce's automatic relationship matching.
AnswerA

Storing the legacy Case ID in an external ID field on Case allows the architect to reference it when loading Case Comments. The Bulk API can use the external ID to look up the parent Case record and establish the relationship. This is a standard practice for migrating related records in multiple passes, ensuring referential integrity without manual intervention.

Why this answer

To link Case Comments to Cases in a two-pass migration, the architect must store the legacy Case ID on the Case record as an external ID. This allows the second load to use that external ID to look up the parent Case and correctly associate comments. Without it, the relationship cannot be established, leading to orphaned comments.

Exam trap

The trap here is thinking that Salesforce or Data Loader can automatically match related records without a stored external ID, when the key must be explicitly persisted for lookups.

13
MCQmedium

A Salesforce architect is migrating 2 million legacy Contact records into a new Salesforce org. The legacy system does not have email addresses for 40% of the contacts, but Salesforce's standard Email field is not required. During a test load of 10,000 records using the Bulk API in serial mode, the load succeeds but the architect notices that duplicate contacts are being created because the legacy system reused a 'legacy_id__c' external ID field for different contacts across regions. What should the architect do to prevent duplicate creation during the full migration?

A.Create a new unique external ID field by concatenating region and legacy_id, populate it in the source data, and use it as the external ID for upsert operations.
B.Set the legacy_id__c field as unique in Salesforce and retry the load; Salesforce will automatically reject duplicate external IDs.
C.Enable the 'Prevent Duplicates' setting on the Contact object and use Data Loader's deduplication feature before loading.
D.Use the Bulk API in parallel mode with a batch size of 1 to ensure each record is processed individually and duplicates are avoided.
AnswerA

The duplicate creation stems from non-unique external IDs. By creating a composite external ID that combines region and legacy_id, the architect ensures each record has a unique identifier for upsert. This allows the Bulk API to correctly match existing records and prevent duplicates, while also preserving the original legacy ID for reference in a separate field.

Why this answer

The duplicate creation is due to the external ID field not being unique across regions. The correct solution is to create a new composite external ID that combines region and legacy ID, ensuring uniqueness. This allows upsert operations to correctly identify records and prevent duplicates, while the original legacy ID can be retained for audit purposes.

Exam trap

The trap here is assuming that enabling duplicate rules or setting a unique constraint on the existing non-unique field will solve the problem, when the source data itself contains duplicates that must be transformed.

14
MCQmedium

You are migrating records that contain multi-select picklist values. What is the key consideration for mapping this data from a legacy system?

A.The values must be separated by commas in the source file.
B.Each value must be imported as a separate child record.
C.The values must be formatted as a semi-colon separated string.
D.The picklist must be converted to a custom object before loading.
AnswerC

Salesforce requires multi-select picklist values to be provided as a string with each value separated by a semi-colon. This is the only format that the Salesforce API accepts for this field type. Transforming the source data to this specific format is a mandatory step in the data migration process.

Why this answer

Multi-select picklists are stored as semi-colon separated strings in Salesforce. When migrating, the source data must be transformed to match this specific delimiter format. Failure to use the exact semi-colon separator, or including values that are not currently defined in the Salesforce picklist configuration, will result in import errors.

It is essential to sanitize and format these values in the ETL layer before the load.

Exam trap

Test-takers frequently assume multi-select picklists are mapped using comma-separated lists or arrays, ignoring Salesforce's strict requirement for a semi-colon delimiter format.

15
Multi-Selectmedium

When planning a large data migration, which TWO tasks are essential to perform before the data load? (Choose two)

Select 2 answers
A.Deactivate all triggers, workflow rules, and flows that could interfere with the load.
B.Increase the Salesforce API limit for the specific migration user.
C.Establish a field mapping document that reconciles legacy data to Salesforce objects.
D.Delete all existing records in the environment to ensure a fresh start.
E.Enable the 'Parallelize' setting on all user profiles.
AnswersA, C

Disabling automation is essential to prevent recursive triggers or unintended updates that degrade performance. When loading large volumes, automated processes can consume CPU and DML limits, causing the load to fail. Clearing the path allows the raw data to be inserted efficiently before re-enabling business logic later.

Why this answer

Preparing the environment by deactivating automation and establishing a clean mapping strategy are foundational steps. Deactivating automation prevents unintended side effects and performance bottlenecks during the load. A robust mapping strategy, including data cleansing and validation, ensures the migration follows business requirements.

These steps minimize downtime and reduce the risk of data corruption, ensuring the system remains performant and the data remains high quality post-migration.

Exam trap

Candidates often forget to deactivate asynchronous automation like triggers and flows, leading to performance bottlenecks, governor limit exceptions, and corrupted legacy data relationships during large data loads.

16
MCQmedium

A Salesforce data architect is migrating 500,000 Lead records from a legacy system. The legacy data contains a field 'Lead_Status' with values such as 'New', 'Working', 'Qualified', 'Unqualified'. Salesforce's Lead Status picklist has values: 'Open - Not Contacted', 'Working - Contacted', 'Closed - Converted', 'Closed - Not Converted'. The architect needs to map the legacy values to the Salesforce picklist values. Which approach should be used to ensure a successful migration?

A.Load the legacy values as-is and then run a batch Apex job to update the Lead Status field.
B.Create a custom field to store the legacy Lead Status and leave the standard Lead Status blank.
C.Modify the Salesforce Lead Status picklist to include the legacy values.
D.Use an ETL tool to map legacy values to the corresponding Salesforce picklist values before loading.
AnswerD

An ETL tool can transform the legacy Lead Status values to match the Salesforce picklist values exactly. This ensures that the standard Lead Status field is populated correctly and avoids load errors. The mapping can be defined in the ETL tool's transformation logic. This is the most reliable method for ensuring data quality.

Why this answer

Mapping legacy values to the standard Salesforce picklist values must occur before loading to avoid errors and ensure data consistency. An ETL tool provides the necessary transformation capabilities. Other options either avoid the mapping, cause load failures, or alter the Salesforce schema unnecessarily.

Exam trap

The trap here is thinking that Salesforce will automatically map legacy picklist values or that the picklist can be easily modified without impact.

17
MCQmedium

What is the primary benefit of using a staging database for data migration?

A.It provides a backup of the source data.
B.It allows for complex data transformation and cleansing.
C.It eliminates the need for field mapping documentation.
D.It speeds up the actual data insertion into Salesforce.
AnswerB

Staging databases are essential for complex transformations that are difficult to perform within Salesforce or CSV tools. They enable SQL-based joins, cleansing scripts, and deduplication logic, ensuring that the data is perfectly structured and validated before it is uploaded to the final production instance for migration.

Why this answer

A staging database provides a clean, neutral environment to perform ETL (Extract, Transform, Load) tasks such as data deduplication, value normalization, and relationship mapping. By transforming the data outside of Salesforce, architects ensure that only clean, verified data enters the target environment. This minimizes the risk of system-level errors and validation failures that occur when attempting to manipulate raw, messy data directly inside the Salesforce platform.

Exam trap

Candidates often confuse the staging database with a backup or a data warehouse. They fail to recognize its primary function as an ETL workspace for cleaning and transforming data before Salesforce ingestion.

18
MCQeasy

A Salesforce data architect is preparing to migrate 5 million Account records from a legacy CRM. The legacy system has a field 'Legacy_Owner__c' that contains the email address of the account owner. During migration, the architect must ensure that the ownership is assigned to the correct active Salesforce User. Which approach should be used to map the legacy owner email to the Salesforce User ID?

A.Load the email addresses into the OwnerId field directly; Salesforce will automatically resolve them to User IDs.
B.Use the Data Loader's 'Bulk API' with the 'Assignment Rule' option enabled to automatically assign owners based on email.
C.Create a custom field on Account to store the legacy owner email, then use a post-load process or Data Loader to update OwnerId based on a User lookup.
D.Load the legacy owner email into the Account Owner field using the 'External ID' feature of the User object.
AnswerC

This approach correctly separates the migration of Account data from the ownership assignment. Storing the legacy email in a custom field preserves the original data for audit, and a subsequent update using a User lookup (via ETL, Data Loader, or Apex) can accurately set OwnerId. This avoids load failures and ensures ownership is assigned to the correct active User. It also allows validation of email-to-user mappings before final assignment.

Why this answer

To map legacy owner emails to Salesforce User IDs, the architect should first load Accounts with the legacy email stored in a custom field. Then, using a lookup or ETL transformation, the correct User IDs can be determined and applied to the OwnerId field. This two-step process avoids errors and ensures accurate ownership assignment.

Exam trap

The trap here is assuming that Salesforce can automatically resolve email addresses to User IDs during a data load without an explicit mapping step.

19
MCQmedium

A multinational corporation is migrating 2 million Account records and 5 million Contact records from a legacy CRM into Salesforce. The legacy system stores the Account's legacy ID on the Contact record as a foreign key. The target Salesforce org uses an external ID field on Account called Legacy_ID__c. What is the most efficient way to associate Contacts with their Accounts during the migration?

A.Migrate Contacts first with a placeholder Account, then migrate Accounts and update Contacts via a batch Apex job.
B.Export the Salesforce Account record IDs after migration, then use those IDs to update the Contact records in the legacy system before migrating Contacts.
C.Use the Data Loader's 'Insert' operation for Contacts and rely on Salesforce's auto-association feature to link Contacts to Accounts based on matching names.
D.Migrate Accounts first, then migrate Contacts with a lookup to Account using the Legacy_ID__c field as the external ID in the Account relationship field.
AnswerD

This is correct because Salesforce allows you to populate a lookup relationship using an external ID field. After Accounts are migrated, Contacts can be loaded with the Account's legacy ID in the Account lookup field, and Salesforce will resolve it to the correct Account record. This avoids the need for a second pass or manual mapping and is efficient for large volumes.

Why this answer

Migrating Accounts first and then using the external ID field Legacy_ID__c in the Contact's Account lookup field allows Salesforce to automatically resolve the relationship. This is the most efficient method because it leverages built-in external ID resolution during data load, eliminating the need for post-migration updates or complex Apex jobs. It ensures referential integrity and scales well for large data volumes.

Exam trap

The trap here is thinking that you need to migrate Contacts first or use custom code to establish relationships, when Salesforce natively supports populating lookups via external IDs.

20
Multi-Selectmedium

A data architect is preparing to migrate 20 million Opportunity records from a legacy CRM into Salesforce. The source data includes Opportunities with related OpportunityLineItem records and historical StageName values. The architect must ensure that the migration preserves data integrity and avoids common Bulk API pitfalls. Which TWO actions should the architect take before initiating the load? (Choose two.)

Select 2 answers
A.Disable all validation rules, workflow rules, and triggers on Opportunity during the load to improve performance.
B.Load Opportunity records first, then load OpportunityLineItem records using the parent Opportunity's External ID to establish the relationship.
C.Set the Batch Size to 1 record to avoid governor limit issues and ensure maximum error isolation.
D.Enable AllOrNone on all Bulk API batches to guarantee that no partial data is committed if a single record fails.
E.Ensure that the External ID field on Opportunity is marked as Unique to support upsert and prevent duplicate Opportunities.
AnswersB, E

OpportunityLineItem records require a parent OpportunityId. By loading Opportunities first with an External ID, the architect can then load line items referencing that External ID, ensuring referential integrity. This two-pass approach is standard for parent-child migrations and avoids orphaned child records.

Why this answer

Loading parents before children using an External ID preserves referential integrity, and marking the External ID as Unique supports upsert and prevents duplicates. Together, these actions address the core data integrity concerns for a large Opportunity migration with related line items.

Exam trap

The trap here is thinking that disabling automation or using AllOrNone is a best practice for data integrity; both can introduce risk or reduce throughput without guaranteeing correctness.

21
MCQhard

A data architect is migrating 10 million Case records and their related Case Comments from a legacy system to Salesforce Service Cloud. The legacy system stores Case Comments in a separate table with a foreign key to the Case. The architect plans to use Bulk API 2.0 for the migration. What is the most critical consideration for maintaining referential integrity between Cases and Case Comments during the load?

A.Load Cases and Case Comments simultaneously using parallel processing to save time.
B.Load Cases first, then load Case Comments with the correct ParentId referencing the Salesforce Case ID.
C.Disable validation rules on Case Comments to avoid errors during the load.
D.Use an External ID field on Case Comments to link to the legacy Case ID without loading Cases first.
AnswerB

To maintain referential integrity, Cases must be loaded first to generate Salesforce IDs. Then, Case Comments can be loaded with the ParentId set to the corresponding Case ID. This requires mapping legacy Case IDs to Salesforce IDs, often via an External ID field on Case. This sequential approach ensures that each comment is linked to an existing Case.

Why this answer

Referential integrity requires that parent records exist before children. Therefore, Cases must be loaded first, and their Salesforce IDs captured to populate the ParentId on Case Comments. Using an External ID on Case can facilitate mapping.

Parallel loading or relying solely on External IDs without loading Cases first would fail.

Exam trap

The trap here is thinking that External IDs can bypass the need for the parent record to exist, or that parallel loading can maintain relationships.

22
MCQhard

A financial services company is migrating 5 million Account records from a legacy CRM to Salesforce. The legacy system has a field 'AnnualRevenue' stored as a string with currency symbols and commas (e.g., '$1,234,567.89'). The target Salesforce field is a Currency field with 2 decimal places. During a test load of 100,000 records using the Bulk API, the architect receives errors indicating 'Invalid currency format'. What is the most efficient way to resolve this before the full migration?

A.Modify the Salesforce field type to Text to accept the legacy format, then create a formula field to display the numeric value.
B.Pre-process the source data using an ETL tool or script to remove non-numeric characters and convert the string to a decimal number before loading.
C.Use Data Loader's 'Transform' feature to strip currency symbols and commas before loading.
D.Load the data as-is and then use a scheduled Apex job to clean up the values after import.
AnswerB

The errors occur because the Bulk API expects numeric values for Currency fields, but the source data contains symbols and commas. Pre-processing the data to remove non-numeric characters and convert to decimal ensures the data matches the target format. This is the most efficient and reliable method, as it addresses the issue at the source and avoids repeated load failures.

Why this answer

The Bulk API rejects records with non-numeric characters in Currency fields. Pre-processing the source data to remove symbols and commas and convert to decimal is the most efficient solution. It ensures data conforms to Salesforce's expected format, preventing load errors and maintaining data integrity without compromising field types or requiring post-load fixes.

Exam trap

The trap here is assuming Data Loader can transform data during load or that changing the field type is a quick fix, when the correct approach is to clean the data before migration.

23
MCQhard

Refer to the exhibit. Why might 'Parallel' concurrency mode cause errors during a large migration?

A.It causes the API to exceed the daily limit for API calls.
B.It leads to record locking contention on parent objects.
C.It prevents the data from being loaded in a specific order.
D.It forces the data to be processed in a single thread.
AnswerB

Parallel processing attempts to update records as fast as possible. If multiple threads attempt to update or insert child records that share the same parent, they will contend for the parent record lock, resulting in 'UNABLE_TO_LOCK_ROW' errors that can stop the migration process.

Why this answer

Refer to the exhibit. The JSON configuration shows a Bulk API 2.0 job using 'Parallel' concurrency mode. The Data Architect must be aware that parallel processing increases the risk of record locking when parent records have many children being processed simultaneously.

This configuration requires careful monitoring of record contention to prevent failures, even though it provides the fastest throughput for independent records.

Exam trap

Candidates often confuse parallel mode performance benefits with safety, assuming it prevents errors. They overlook how simultaneous processing of child records tied to identical parents triggers record locking contentions during high-volume data migrations.

24
MCQmedium

During a migration, you discover that the source system has a different date format than the required ISO 8601 format for Salesforce. What is the most efficient way to handle this?

A.Use an Apex trigger to format the dates after the records are inserted.
B.Transform the data using an ETL tool or script before loading it into Salesforce.
C.Update the Salesforce field type to 'Text' to accommodate any date format.
D.Manually update the records in the Salesforce UI after the migration completes.
AnswerB

Transforming data in the ETL layer is the industry standard for migration. It keeps the Salesforce org clean and prevents errors during ingestion. Using tools such as Informatica, Mulesoft, or custom scripts allows for robust validation and formatting, ensuring the incoming data meets all schema requirements before submission.

Why this answer

Data transformation should ideally occur in the ETL (Extract, Transform, Load) layer before the data reaches Salesforce. By formatting dates during the extraction or transformation process, you ensure that the data conforms to Salesforce requirements before it hits the API. This is more scalable than attempting to force Salesforce to interpret varying formats, which can lead to data loss or import errors during the ingestion process.

Exam trap

Candidates frequently choose to handle formatting inside Salesforce via complex formula fields or triggers rather than cleaning the data upstream.

25
MCQmedium

Which object must be loaded first in a standard Salesforce migration involving Accounts, Contacts, and Opportunities?

A.Opportunities, because they drive revenue reporting.
B.Contacts, to ensure the primary contact is defined.
C.Accounts, to establish the parent records for children.
D.It does not matter if you use the Bulk API.
AnswerC

Accounts serve as the foundation of the relational data model. By loading them first, you generate the Salesforce IDs necessary to map children (Contacts and Opportunities) correctly during subsequent loads. This top-down hierarchical approach is the industry-standard sequence for successful data migration projects in Salesforce.

Why this answer

Accounts must be loaded first because they are the parent objects for both Contacts and Opportunities in the standard data model. Without the Account records present, the foreign key lookups for Contacts and Opportunities will fail, as they have no valid parent ID to link to. Establishing the hierarchy is the fundamental requirement for data referential integrity.

Exam trap

Test-takers sometimes select Opportunities first because they represent revenue, ignoring the fundamental foreign key dependency requiring parent Accounts.

26
MCQhard

During a large data migration, a developer notices that existing Apex triggers are causing severe performance degradation. What is the most effective way to prevent trigger execution without deleting the code?

A.Comment out the trigger logic and redeploy the code.
B.Disable the triggers in the Setup menu under 'Manage Triggers'.
C.Use a Custom Setting to control trigger execution.
D.Remove user permissions for the integration user.
AnswerC

A Custom Setting (or Custom Metadata) provides a clean, highly performant way to toggle Apex execution logic. This allows developers to disable resource-intensive triggers temporarily for the duration of the data load and reactivate them immediately afterward without needing a full code deployment cycle.

Why this answer

The most effective method to mitigate trigger-related performance issues during high-volume loads is to implement a 'Bypass' or 'Switch' pattern using Custom Settings or Custom Metadata. By checking this flag at the start of the trigger, logic can be conditionally skipped. This preserves code integrity while allowing the migration to proceed at maximum speed without the overhead of complex validation or automated process logic.

Exam trap

Candidates often select architectural solutions like permanently deleting or rewriting trigger logic, failing to recognize that configuration-based bypass mechanisms using custom settings are much safer and more efficient.

27
MCQeasy

A data architect is migrating 500,000 records into a custom object that has a validation rule requiring the 'Status__c' field to be either 'Active' or 'Inactive'. The legacy data contains some records with a blank Status__c. The architect needs to ensure the migration succeeds without modifying the validation rule. What should be done?

A.Temporarily deactivate the validation rule during the data load, then reactivate it afterward.
B.Load the records with blank Status__c using the Data Loader's 'Bulk' mode, which bypasses validation rules.
C.Transform the legacy data by setting a default value of 'Inactive' for any records with a blank Status__c before loading.
D.Create a before insert trigger to set a default value for Status__c when it is blank.
AnswerC

This is correct because the validation rule requires a specific value. Transforming the data to populate a valid default ensures that all records pass validation. This approach is simple, does not require changing the validation rule, and maintains data quality. It is a standard data cleansing step in migration projects.

Why this answer

The validation rule requires Status__c to be 'Active' or 'Inactive'. The most straightforward solution is to transform the legacy data by setting a default value of 'Inactive' for blank entries. This ensures all records pass validation without altering the validation rule or adding custom code.

Data transformation during migration is a best practice to meet target system requirements and maintain data integrity.

Exam trap

The trap here is believing that Bulk API bypasses validation rules or that deactivating rules is an acceptable workaround, when the requirement explicitly states not to modify the validation rule.

28
MCQhard

A healthcare provider is migrating 8 million legacy Patient_Visit__c records into Salesforce. Each visit references a legacy numeric provider code, and the legacy system does not expose a stable unique identifier for each provider. The Salesforce Provider__c object already has a custom external ID field named Legacy_Provider_Code__c that is populated. The architect wants to load visits without first exporting provider Salesforce IDs. Which approach should be used?

A.Create a junction object between Provider__c and Patient_Visit__c, load the junction records with both legacy codes, and let Salesforce infer the lookups.
B.Load Patient_Visit__c records with a lookup field populated by the legacy provider code, relying on the external ID field to resolve the relationship during the load.
C.Load Patient_Visit__c records into a staging custom object, then use a scheduled Apex job to copy each legacy provider code into a text field and later reconcile manually.
D.Export all Provider__c records with their Salesforce IDs, VLOOKUP the IDs into the visit file, then load visits with the 18-character Salesforce ID in the lookup field.
AnswerB

Salesforce resolves relationship fields by matching the value supplied in the relationship field against the target object's external ID field, so loading visits with the legacy provider code in the Provider__c relationship maps each visit to the correct Provider__c record without a prior ID export. This is exactly the pattern Bulk API 2.0 and Data Loader support for external ID lookups, and it scales for multi-million record loads.

Why this answer

Relationship fields can be populated with an external ID value when the target object has that field marked as an external ID, letting Salesforce match and set the lookup during the load. Because Provider__c already has Legacy_Provider_Code__c configured as an external ID and populated, the visit load can reference the provider by that code directly, eliminating a separate ID export and join step and keeping the migration efficient at scale.

Exam trap

The trap here is assuming that a lookup must always be populated with the 15- or 18-character Salesforce record ID, when an external ID on the parent object can be used instead.

29
Multi-Selectmedium

A Salesforce architect is planning a migration of 2 million custom object records from a legacy Oracle database. The legacy system has a field 'Status' with values like 'Active', 'Inactive', 'Pending', and 'Archived'. In Salesforce, the Status field is a picklist with values 'Active', 'Inactive', and 'Pending' only. The architect needs to migrate the data while preserving as much information as possible and ensuring data quality. Which two actions should the architect take? (Choose two.)

Select 2 answers
A.Add 'Archived' as a new picklist value in Salesforce to accommodate the legacy data without transformation.
B.Create a custom text field to store the original legacy status value for audit purposes, in addition to mapping to the picklist.
C.Use a validation rule to prevent 'Archived' from being loaded, and manually update those records after migration.
D.Map the legacy 'Archived' value to 'Inactive' in Salesforce and document the transformation for future reference.
E.Exclude records with 'Archived' status from the migration to avoid picklist conflicts.
AnswersB, D

Storing the original legacy status in a separate text field preserves the exact source value for auditing and reporting, while the picklist field is used for operational purposes. This dual approach ensures data quality and traceability without compromising the picklist's integrity. It is a best practice when transforming data during migration to retain historical context.

Why this answer

The architect should map the unsupported 'Archived' value to a valid picklist value like 'Inactive' to maintain data integrity and avoid load failures. Additionally, storing the original legacy status in a separate text field preserves the exact source information for auditing. This combination ensures data quality, traceability, and compliance with Salesforce picklist constraints.

Exam trap

The trap here is assuming that adding a new picklist value or excluding records is acceptable, when the better approach is to transform the data and preserve the original value for audit.

30
MCQeasy

A data architect needs to migrate 500,000 Account records from a legacy system into Salesforce. The legacy system uses a 15-digit alphanumeric string as the unique identifier. The architect must ensure that future updates can match existing Accounts without creating duplicates. Which field configuration should be used on the Account object to support this requirement?

A.Create a custom Text field of length 15 and use it as the Salesforce Record ID.
B.Create a custom Text field of length 15 and make it an External ID.
C.Use the standard Account Number field and set it as an External ID.
D.Create a custom Text field of length 15 and mark it as Unique.
AnswerB

An External ID field is specifically designed to store unique identifiers from external systems and is used by the upsert operation to match records. Making the custom Text field an External ID allows the architect to use upsert with this field as the match key, preventing duplicates during future updates. The length of 15 accommodates the legacy identifier.

Why this answer

To support upsert and prevent duplicates, the legacy identifier must be stored in a field marked as an External ID. A custom Text field of sufficient length, designated as External ID, allows the Data Loader or Bulk API to match existing records based on that field. This is the standard pattern for integrating external systems and maintaining data integrity.

Exam trap

The trap here is confusing the Unique attribute with the External ID attribute; only External ID enables upsert matching.

31
MCQmedium

Which of the following is the most important factor when choosing between the Bulk API 2.0 and the SOAP API for a large data migration?

A.The use of OAuth authentication protocols.
B.The requirement to process records synchronously.
C.The total number of records to be migrated.
D.The availability of a command-line interface.
AnswerC

The volume of records is the deciding factor. Bulk API 2.0 is purpose-built for high-volume jobs, handling large batches asynchronously to optimize server resources. In contrast, the SOAP API is designed for smaller, real-time integrations, and attempting to use it for large migrations causes timeouts and system stress.

Why this answer

The primary factor is the volume of data being moved. Bulk API 2.0 is designed specifically for large datasets (hundreds of thousands to millions of records) and operates asynchronously, reducing the likelihood of request timeouts. The SOAP API, while powerful for real-time transactional operations, is synchronous and poorly suited for large-scale data migrations due to its overhead and batch limitations.

Exam trap

Test-takers often confuse real-time transactional APIs with bulk capabilities, selecting synchronous SOAP APIs for massive datasets instead of recognizing volume as the primary decision driver.

32
MCQhard

A data architect is migrating 20 million Case records from a legacy system into Salesforce. The legacy system has a 'Case_Status__c' field that must map to the standard Case Status picklist. The legacy values include 'Open', 'Closed', 'Escalated', 'Pending', and 'Resolved'. The Salesforce Case Status picklist contains 'New', 'Working', 'Escalated', 'Closed', and 'Pending'. The architect needs to ensure that the migration does not fail due to picklist value mismatches. Which approach should be taken?

A.Load the legacy values as-is into a custom text field and create a formula field to display the mapped Salesforce status.
B.Disable picklist validation on the Case Status field during the migration and re-enable it afterward.
C.Add the missing legacy values ('Open' and 'Resolved') to the Salesforce Case Status picklist before the migration.
D.Use an ETL transformation to map 'Open' to 'New' and 'Resolved' to 'Closed' before loading the data.
AnswerD

This approach correctly handles the mismatch by transforming legacy values to valid Salesforce picklist values. Mapping 'Open' to 'New' and 'Resolved' to 'Closed' aligns with common business semantics and avoids altering the standard picklist. The ETL tool can perform this transformation at scale for 20 million records, ensuring the load succeeds without errors. This is the recommended practice for picklist value discrepancies.

Why this answer

The architect should use an ETL transformation to map legacy picklist values to valid Salesforce values before loading. Mapping 'Open' to 'New' and 'Resolved' to 'Closed' ensures the data conforms to the standard Case Status picklist, preventing load failures and maintaining data integrity. This approach is scalable for 20 million records.

Exam trap

The trap here is assuming that legacy picklist values can be loaded directly or that the standard picklist can be easily modified without side effects.

33
MCQmedium

A healthcare company is migrating 4 million Patient__c records from a legacy system into Salesforce. The legacy extract includes a field named 'last_modified' that is populated for 99.8% of records, but the remaining 0.2% have a null value. The target external ID field requires a unique, non-null value. What should the data architect do to ensure a successful migration?

A.Set the external ID field to allow nulls in the field definition, then migrate the records with null values as-is.
B.Use the standard Salesforce Data Loader's 'Insert' operation instead of 'Upsert' to bypass the external ID requirement.
C.Generate a synthetic unique value for the records with null last_modified by using the legacy primary key concatenated with a timestamp, and map that to the external ID field.
D.Exclude the 0.2% of records with null last_modified from the migration and document them as exceptions for manual entry later.
AnswerC

This is correct because the external ID field must be populated with a unique, non-null value for every record. The legacy primary key is already unique, and adding a timestamp ensures no collision with other synthetic values. This approach preserves referential integrity and allows upserts. It avoids data loss and meets the field constraint without altering the source system.

Why this answer

The external ID field must contain a unique, non-null value for every record to support upsert and maintain data integrity. Generating a synthetic value from the legacy primary key and a timestamp ensures uniqueness and populates the field for the 0.2% of records with missing last_modified. This avoids data loss and meets the field constraint without modifying the source system or relaxing the Salesforce field definition.

Exam trap

The trap here is assuming that external ID fields can contain nulls or that nulls can be ignored, when in fact a unique non-null value is required for upsert operations.

34
MCQhard

A data architect is migrating 30 million Opportunity records into Salesforce. The legacy system has a 'Close_Date__c' field that is a date, but some records have invalid dates such as '0000-00-00' or future dates beyond 2099. The Salesforce Opportunity Close Date field is a standard Date field. The architect must ensure the migration does not fail due to these invalid dates. Which approach should be used?

A.Convert the invalid dates to a default date such as '1900-01-01' to maintain a non-null value.
B.Load the invalid dates as NULL and log the affected records for manual review after the migration.
C.Create a custom text field to store the original date string and load it alongside the standard Close Date field, leaving Close Date blank for invalid records.
D.Use the Data Loader's 'Allow Field Truncation' option to bypass date validation and load the invalid dates as-is.
AnswerB

Setting invalid dates to NULL prevents load failures because Salesforce Date fields accept NULL values. Logging the affected records allows for post-migration cleanup and manual correction. This approach balances data integrity with migration success, ensuring that valid records are loaded while invalid ones are flagged for review. It is a pragmatic solution for large volumes where manual correction before load is impractical.

Why this answer

The architect should set invalid dates to NULL and log the affected records for post-migration review. Salesforce Date fields accept NULL values, so this prevents load failures. Logging allows for manual correction later, ensuring data quality without blocking the migration.

This approach is scalable for 30 million records and maintains the integrity of the standard Close Date field.

Exam trap

The trap here is assuming that invalid dates can be forced into Salesforce using truncation or default values, when the correct approach is to nullify and log them.

35
MCQhard

A data architect is migrating 12 million Opportunity records into a new Salesforce org. The legacy system uses a custom 'Opportunity_Key__c' that must be populated for integration purposes. During a test load using the Bulk API in parallel mode, the team observes that some records fail with 'UNABLE_TO_LOCK_ROW' errors. What is the most likely cause of these errors?

A.Multiple batches in parallel are attempting to update the same parent Account records, causing row-level lock contention.
B.The custom 'Opportunity_Key__c' field is not marked as an external ID, preventing proper indexing and causing locks.
C.The parallel mode uses a single batch that is too large, exceeding the record lock timeout threshold.
D.The Bulk API parallel mode exceeds the daily API request limit, causing lock contention.
AnswerA

When Opportunities are loaded in parallel, Salesforce processes multiple batches concurrently. If these Opportunities share parent Account records, Salesforce must lock those parent records to maintain referential integrity, leading to UNABLE_TO_LOCK_ROW errors when concurrent batches contend for the same parent. This is a classic symptom of parallelism combined with shared parent references.

Why this answer

UNABLE_TO_LOCK_ROW errors during parallel Bulk API loads typically occur when concurrent batches attempt to update child records that reference the same parent records. Salesforce locks parent records to enforce referential integrity, and simultaneous access causes contention. Reducing parallelism or serializing the load for affected objects resolves the issue.

Exam trap

The trap here is attributing row lock errors to API limits or field configuration rather than to concurrent DML on shared parent records.

36
MCQeasy

A data architect is migrating 4 million legacy Customer__c records into Salesforce using the Bulk API. The source system exports a CSV with a column 'Legacy_Id__c' that must remain unique and searchable for downstream integrations. After the initial load, the architect notices duplicate Customer__c records were created because the External ID field was not marked as unique. Which Salesforce field configuration should have been applied to Legacy_Id__c to prevent duplicates during upsert?

A.Set the field type to Text (Encrypted) and enable 'Do Not Track History'.
B.Mark the field as an External ID and select 'Unique' in the field definition.
C.Create a validation rule that compares Legacy_Id__c to existing records.
D.Enable 'Required' on the field and use Data Loader's 'Insert' operation instead of 'Upsert'.
AnswerB

For upsert operations, Salesforce requires an External ID field. Selecting 'Unique' enforces a uniqueness constraint at the database level, so any duplicate Legacy_Id__c values in the source CSV or already in Salesforce cause the record to be treated as an update rather than an insert, preventing duplicate Customer__c records.

Why this answer

An External ID field with the Unique attribute is the correct mechanism for upsert matching and duplicate prevention. When the Bulk API processes an upsert, it uses the External ID to determine whether to insert or update. Without the Unique constraint, Salesforce cannot reliably match records, and duplicate Legacy_Id__c values can result in multiple Customer__c records.

Exam trap

The trap here is assuming that marking a field as Required or creating a validation rule can enforce uniqueness during a data load.

37
MCQeasy

What is the primary benefit of performing a data cleansing exercise before initiating a migration to Salesforce?

A.It eliminates the need for field mapping in the ETL process.
B.It ensures that the target Salesforce Org meets storage limits.
C.It prevents the migration of invalid or duplicate data into the new system.
D.It automatically adjusts the Salesforce schema to fit the legacy data.
AnswerC

Cleansing removes inconsistencies, duplicates, and inaccurate values. By doing this early, you ensure the new Salesforce system is populated with high-quality, trusted data. This is essential for successful adoption and reporting, as users are more likely to trust the system when the records are clean and reliable.

Why this answer

Data cleansing is the process of detecting and correcting corrupt or inaccurate records. Doing this before migration is vital because it prevents the 'garbage in, garbage out' syndrome. High-quality data ensures that the new system is reliable from day one, improves user adoption, and prevents the need for costly post-migration data remediation projects which are significantly more complex once the data is integrated into existing business logic.

Exam trap

Candidates often underestimate the cost of post-migration cleanup. They wrongly assume that fixing data inside Salesforce is easier than cleaning it beforehand, ignoring the complexity of existing business logic and dependencies.

38
MCQhard

A financial services firm is migrating 10 million transaction records into a custom object Transaction__c. The legacy system uses a composite key of AccountNumber and TransactionDate to uniquely identify transactions. The target Salesforce org has a unique external ID field Transaction_Key__c. The architect needs to ensure that re-running the migration does not create duplicate records and that updates to existing transactions are applied. Which approach should be used?

A.Create a custom Apex trigger to check for existing transactions before insert and update them if found.
B.Use the Data Loader's Update operation with a SOQL query to fetch existing records and match them by AccountNumber and TransactionDate.
C.Use the Data Loader's Insert operation and rely on Salesforce's duplicate rules to prevent duplicates.
D.Populate Transaction_Key__c with the concatenation of AccountNumber and TransactionDate, then use the Data Loader's Upsert operation with Transaction_Key__c as the external ID field.
AnswerD

This is correct because the composite key from the legacy system can be concatenated into a single string that populates the unique external ID field. Upsert uses this field to match existing records, preventing duplicates and allowing updates. This approach is standard for migrations where a natural composite key exists and can be represented as a single string within the 255-character limit.

Why this answer

Concatenating the composite key into the external ID field and using Upsert allows the migration to match existing records by that key. Upsert updates existing records and inserts new ones, preventing duplicates. This is the most efficient and standard method for handling composite keys in large migrations, as it leverages Salesforce's built-in matching and bulk processing capabilities without custom code.

Exam trap

The trap here is assuming that duplicate rules or custom triggers can replace the need for an external ID and Upsert operation, when they cannot efficiently handle large-scale updates and inserts.

Ready to test yourself?

Try a timed practice session using only Data Migration questions.