Courseiva

CCNA Large Data Volume Considerations Questions

36 questions · Large Data Volume Considerations · All types, answers revealed

1
MCQeasy

What is the primary architectural goal of using the Salesforce Bulk API 2.0?

A.To provide real-time updates for end-user applications.
B.To bypass all security and validation rules.
C.To efficiently manage large volume data operations.
D.To allow for direct SQL query access to the database.
AnswerC

Bulk API 2.0 is purpose-built to handle millions of records by automatically partitioning the data into optimized batches. This allows for reliable, high-throughput processing while maintaining stability in the Salesforce environment, making it the primary tool for large-scale data imports and exports.

Why this answer

The Bulk API 2.0 is designed specifically for handling large data volumes by automating the batching process. It simplifies the developer experience while optimizing for high throughput. By offloading the batching logic to the Salesforce platform, it ensures that data loads occur in a manner that respects system resources and governor limits, which is the foundational requirement for managing enterprise-scale data in the cloud.

Exam trap

Candidates often confuse the Bulk API with standard REST/SOAP APIs, failing to realize the Bulk API is specifically engineered for throughput, not for real-time or low-latency requests.

2
Multi-Selectmedium

A data architect is planning to implement a Skinny Table for a custom object with 10 million records to improve query performance. Which two considerations are critical when designing the Skinny Table? (Choose two.)

Select 2 answers
A.The Skinny Table requires a synchronization mechanism to keep data current with the main object.
B.The Skinny Table can be used as the primary data source for all reporting to replace the main object.
C.The Skinny Table should be indexed on all fields to maximize query flexibility.
D.The Skinny Table must include all fields from the original object to maintain data integrity.
E.The Skinny Table should include only the fields required for the most common queries.
AnswersA, E

Because a Skinny Table is a separate custom object, it does not automatically reflect changes to the main object. A synchronization process, such as a trigger or batch job, must be implemented to update the Skinny Table when the main object changes. Without synchronization, the Skinny Table would become stale and produce incorrect query results.

Why this answer

A Skinny Table is a denormalized copy of a subset of fields from a large object, used to accelerate specific queries. The two critical considerations are selecting only the fields needed for common queries to minimize row size, and implementing a synchronization mechanism to keep the copy up to date. Without these, the Skinny Table would either be too large to be effective or become inconsistent with the source data.

Exam trap

The trap here is assuming a Skinny Table must mirror all fields or can replace the main object entirely, when it is a targeted performance optimization.

3
MCQeasy

A data architect is reviewing a custom object that has grown to 8 million records. Users report that list views and reports are slow because they sort on a text field, Priority__c, which has only three possible values. What should the architect recommend to improve performance?

A.Enable Divisions on the object and assign records to divisions based on Priority__c.
B.Request a custom index on Priority__c to speed up sorting.
C.Convert Priority__c to a picklist field with a restricted set of values to improve selectivity.
D.Remove sorting on Priority__c from list views and reports, and instead sort on an indexed field or filter on a more selective field.
AnswerD

Sorting on a low-cardinality field like Priority__c forces the query optimizer to scan many records because the field is not selective. Removing the sort or replacing it with a sort on an indexed, high-cardinality field can dramatically improve performance. Filtering on a more selective field also reduces the number of records that need to be sorted.

Why this answer

For large data volumes, query performance is heavily influenced by field selectivity. A field with only three distinct values is not selective, so sorting or filtering on it forces the platform to scan many records. The best approach is to avoid sorting on such fields and instead use an indexed, high-cardinality field for sorting or filtering.

This reduces the data set the query must process.

Exam trap

The trap here is assuming that any field used in sorting or filtering should be indexed, even when it has very low cardinality.

4
MCQhard

When dealing with high-frequency updates on a parent object, what is the best practice to prevent lock contention?

A.Always update the parent in the child's trigger.
B.Use an asynchronous approach to aggregate parent updates.
C.Set the parent record to read-only.
D.Convert the lookup relationship to a master-detail.
AnswerB

Asynchronous processing allows updates to be queued and batched, preventing multiple concurrent transactions from fighting over the same parent row. This pattern is essential for high-frequency updates, as it serializes the parent-level changes and allows the system to process them efficiently without risking transactional row locks.

Why this answer

To minimize lock contention, avoid updating the parent record every time a child record is updated. Instead, use an asynchronous pattern like a batch job or a queueable apex to aggregate updates and perform a single update to the parent. This reduces the number of locks requested on the parent object, significantly decreasing the likelihood of row-locking errors during high-volume operations.

Exam trap

Candidates often suggest using triggers for synchronous roll-ups, which causes massive row-locking issues when multiple child records are updated simultaneously in high-volume environments.

5
MCQhard

Refer to the exhibit. A batch job is failing consistently with the provided error message. The job updates child records that have a Master-Detail relationship with a parent Account. What is the primary cause of this error?

A.The batch job is exceeding the heap size limit.
B.Multiple concurrent threads are updating child records pointing to the same parent.
C.The sharing rules on the Account object are too complex.
D.The batch job is processing records in an incorrect order.
AnswerB

Master-Detail relationships enforce a lock on the parent record whenever a child is modified. When multiple batches process child records belonging to the same parent record at the same time, they compete for this parent-level lock, resulting in the UNABLE_TO_LOCK_ROW exception due to contention.

Why this answer

This error occurs due to row-level locking when multiple threads attempt to update records associated with the same parent simultaneously. In Salesforce, updating a child record in a Master-Detail relationship implicitly locks the parent record. When high-volume parallel batches attempt to update children of the same parent, they contend for this single lock, causing transaction failures.

Architecturally, you must serialize these updates or reduce parallel threads to avoid collision.

Exam trap

Candidates often assume that Apex batch limits or query timeouts caused the failure, missing how parent-child Master-Detail record locks trigger row contention under high-concurrency threading.

6
MCQhard

What is the consequence of having a 'non-selective' query running on an object with 50 million records?

A.The query will fail immediately with a compilation error.
B.The query will run faster because it ignores indexes.
C.The query will be blocked or timed out by the platform.
D.The query will automatically create an index for the fields used.
AnswerC

To protect system resources, Salesforce restricts queries that are non-selective on large datasets. If a query cannot be optimized via an index, it is likely to be blocked or time out, ensuring that the system maintains consistent performance for all users in the multi-tenant environment.

Why this answer

Non-selective queries force the system to perform a full table scan. In an environment with millions of records, this consumes massive resources, often leading to query timeouts, increased CPU usage, and potential platform instability. Preventing this is the primary goal of data architecture, as it protects the multi-tenant environment and ensures that individual system processes do not consume resources allocated for other tenants or internal tasks.

Exam trap

Candidates often assume non-selective queries will simply run slowly. They fail to realize that the Salesforce platform proactively blocks or times out these queries to protect the multi-tenant environment.

7
MCQmedium

When dealing with LDV, what is the primary benefit of using External Objects via Salesforce Connect instead of standard Salesforce tables?

A.External objects are automatically indexed by Salesforce.
B.It keeps the data within Salesforce for better reporting performance.
C.It avoids storing massive amounts of data in the Salesforce database.
D.External objects allow for full Apex trigger functionality.
AnswerC

By keeping high-volume, low-access data in an external system, you preserve Salesforce storage and avoid the performance overhead of managing massive tables. This architecture ensures that core business objects remain performant and responsive, while secondary data remains available through a seamless, integrated user experience.

Why this answer

External Objects allow organizations to view and interact with data stored in external systems without importing it into Salesforce. This strategy keeps the Salesforce database lean, avoiding the storage limits and performance degradation associated with managing hundreds of millions of records locally. It is the architectural standard for offloading non-critical, high-volume historical data while keeping it accessible within the platform's user interface.

Exam trap

Candidates often assume External Objects are used for performance speed, rather than storage management. They fail to realize that external calls actually add latency compared to querying local indexed data.

8
MCQmedium

What is a 'Selective Query' in the context of Salesforce LDV?

A.A query that uses exactly five fields in the WHERE clause.
B.A query that returns fewer than 100 records total.
C.A query that utilizes indexed fields to filter within defined thresholds.
D.A query that uses only SOQL and no Apex logic.
AnswerC

A selective query is one that effectively uses indexes to limit the search space to a small percentage of total records. By staying within the platform's selectivity thresholds, the query optimizer can retrieve results quickly without the performance penalty associated with full table scans on massive data sets.

Why this answer

A selective query is one that hits an index and returns a small, efficient subset of records. Salesforce imposes selectivity thresholds; if a query exceeds these (e.g., usually 30% of the first million records), it is considered non-selective. Ensuring queries are selective is the single most important factor in maintaining performance for large datasets, as it allows the database to avoid exhaustive scanning of rows.

Exam trap

Candidates often confuse 'selective' with 'simple' queries. They fail to understand that selectivity is a mathematical threshold based on the percentage of records returned, not just query complexity.

9
MCQeasy

Which object type is most likely to cause performance issues in an LDV environment if not managed correctly?

A.A flat custom object with no relationships.
B.An object with many lookup relationships to other parent objects.
C.An object containing only standard picklist fields.
D.A setup object maintained by the system administrator.
AnswerB

Objects with many lookup relationships are problematic because every record update might involve locking multiple parent records. This leads to severe contention during large data imports and complex queries. These objects require careful planning, such as implementing indexing and optimizing batch processes to minimize the locking overhead.

Why this answer

Highly relational objects, especially those with many lookup relationships, are most prone to performance issues. Each lookup relationship can act as a locking point, and queries involving these fields often require complex joins. By understanding which objects are 'hot' in terms of frequency of access and relationship density, architects can prioritize them for indexing, archiving, and skinny table strategies to maintain system stability.

Exam trap

Candidates often assume that only 'large' objects cause issues. They fail to realize that objects with many lookup relationships create locking contention and join complexity that drastically slows down performance.

10
MCQhard

Refer to the exhibit. In a 50 million record Account table, why might this query perform poorly?

A.The CreatedDate field is not indexed by default in Salesforce.
B.The custom field is likely missing an index, causing a full table scan.
C.The Id field should be the only field in the WHERE clause.
D.Apex triggers are preventing the query from completing.
AnswerB

Without an index on the custom field, the query optimizer cannot efficiently narrow down the search space. When dealing with millions of records, the lack of an index forces the system to scan every row, which is the primary cause of slow query execution and potential timeout errors.

Why this answer

This query uses a filter on CreatedDate (a system-indexed field) and a custom field (Custom_External_ID__c). If Custom_External_ID__c is not indexed, the query optimizer must perform a full table scan. Even with a system index, if the combination of filters is not selective enough, the query will exceed the platform's selectivity threshold, leading to a non-selective query error or severe performance degradation.

Exam trap

Candidates often assume that any field used in a filter is automatically indexed by the platform. They overlook the need for custom indexes on non-standard fields in large tables.

11
MCQhard

A company is migrating 100 million records into Salesforce. Which TWO actions should the architect take to optimize performance and prevent row locking?

A.Enable all custom indexes before the data load.
B.Disable unnecessary triggers during the data load.
C.Use the Bulk API 2.0 with serial mode.
D.Reduce the number of indexes on the target object.
E.Increase the batch size to 10,000.
AnswerB, D

Triggers run for every record processed, consuming CPU and database resources while creating row locks on parent or related objects. Disabling them allows the system to focus exclusively on record insertion, significantly increasing throughput and avoiding potential deadlocks caused by concurrent logic execution during the migration.

Why this answer

High-volume data loads require careful management of record locking and indexing. Disabling triggers during the load prevents unnecessary processing and lock contention. Furthermore, reducing the number of indexes on the target object minimizes the overhead required for every insert operation, as Salesforce must update every index on the object for each new record processed during the migration process.

Exam trap

Candidates often overlook the impact of indexes on DML performance, assuming more indexes are always better, when in reality, every index adds overhead during mass record inserts.

12
MCQmedium

Which platform feature is best used to move data off-platform to avoid Large Data Volume issues?

A.Salesforce BigObjects.
B.Salesforce Connect.
C.Data Loader command line.
D.Platform Events.
AnswerB

Salesforce Connect integrates external data into the org via external objects. This allows companies to keep data outside of the Salesforce database, preventing native performance degradation while ensuring users still have access to the information, which is a perfect pattern for large-scale data archival and management needs.

Why this answer

Salesforce Connect allows organizations to display data from external systems as if it were stored natively in Salesforce, without actually consuming storage or impacting governor limits. By using external objects, companies can maintain massive amounts of historical data off-platform while still providing users with a seamless interface for viewing and reporting on that data within the native environment.

Exam trap

Candidates often suggest archiving data to a Big Object or a standard custom object, which still consumes storage, rather than using Salesforce Connect to keep data off-platform.

13
MCQmedium

Which strategy should be employed when designing a data archiving solution for a high-volume object to ensure continued system performance?

A.Store all historical data in a hidden custom object within Salesforce.
B.Regularly delete records that are older than three years.
C.Offload historical records to an external system or Big Object.
D.Use Salesforce Sharing Rules to hide older records from users.
AnswerC

This is the correct approach to maintain performance. Offloading data to an external repository or Big Object reduces the record count in the transactional object, keeping indexes lean and queries fast, while still allowing access to historical data when needed for compliance or analytical purposes.

Why this answer

Archiving is critical for maintaining performance in Salesforce. By moving stale data to an external data store or a Big Object, you keep the active 'working set' of records small. This ensures that queries, reports, and DML operations remain fast.

An effective strategy involves identifying criteria for record age or status and offloading these records periodically to prevent the primary object from reaching a state that degrades system performance.

Exam trap

Many candidates incorrectly recommend creating more custom indexes or skinny tables instead of removing historical data from the active database entirely.

14
MCQhard

Which design pattern effectively handles high-volume record updates while avoiding 'Too many SOQL queries' errors?

A.Perform a SOQL query inside the trigger loop.
B.Use a Map to cache related records.
C.Use the Future method for every record update.
D.Call the update DML statement inside the loop.
AnswerB

Caching related parent records in a Map ensures that SOQL queries are performed only once for the entire batch. This minimizes resource consumption and prevents governor limit violations, allowing developers to process thousands of records efficiently, which is vital for high-volume data operations in the Salesforce platform.

Why this answer

The use of Maps for caching parent records is essential to avoid repeated querying. By querying all necessary parent records once, storing them in a Map, and accessing them by ID in a loop, you reduce the SOQL query count to one. This pattern is the industry standard for handling large collections of records without hitting governor limits.

Exam trap

Candidates commonly write queries inside iterative loops to fetch related parent data, quickly exhausting SOQL query limits instead of utilizing collection mapping techniques.

15
MCQmedium

When designing a system that requires frequent querying of very large objects, which approach provides the best performance while maintaining data integrity?

A.Always use SOQL with complex joins across many tables.
B.Implement custom indexing or Skinny tables.
C.Move all data to a custom object to simplify the schema.
D.Use the Salesforce REST API for all data retrieval.
AnswerB

Custom indexing and Skinny tables are the most effective native ways to optimize read performance. By providing the database with direct access paths or denormalized data sets, these features allow the query optimizer to return results rapidly, avoiding the performance pitfalls of full table scans.

Why this answer

Denormalization via techniques like Skinny tables or indexing is the standard way to optimize read performance for large volumes. These strategies shift the cost from query time to write time, which is usually preferable in high-volume systems where users need quick access to data. By aligning the database schema with the specific query patterns of the application, you minimize the work the database must do to return results.

Exam trap

Candidates often select standard sharing rules or caching mechanisms, missing that schema-level adjustments are required for massive data volume queries.

16
MCQmedium

A Salesforce org has a custom object Log__c with 5 million records. The object has a lookup to Case. Users report that when they view a Case record, the related list of Log__c records takes a long time to load. The data architect decides to create a custom index on the Case lookup field. After the index is created, performance improves. Which statement best explains why the index improved performance?

A.The index reduces the number of records that need to be queried by pre-filtering the Log__c object.
B.The index allows the related list query to use an index to quickly retrieve Log__c records for the specific Case.
C.The index enables the related list to be loaded from a cache instead of the database.
D.The index sorts the Log__c records by Case, which speeds up the display order in the related list.
AnswerB

When viewing a Case record, the related list of Log__c records is displayed by querying Log__c where Case__c equals the current Case Id. This filter is highly selective because it returns only the logs for one case. With a custom index on Case__c, the query optimizer can use the index to efficiently retrieve those records, significantly improving performance.

Why this answer

The related list on Case queries Log__c records filtered by the Case lookup. This filter is highly selective because it returns only records for one Case. Creating a custom index on the lookup field allows the query optimizer to use the index, avoiding a full table scan.

Thus, the index directly improves the performance of the related list query.

Exam trap

The trap here is assuming that indexes improve performance by sorting or caching, when in fact their primary benefit is enabling fast, selective filtering.

17
Multi-Selecthard

Which TWO of the following are consequences of having excessive indexes on a Salesforce object with large data volumes?

Select 2 answers
A.Improved insert and update performance.
B.Increased time to complete DML operations.
C.Faster SOQL query execution times.
D.Higher risk of row-level locking contention.
E.Automatic reduction in storage space usage.
AnswersB, D

Because each write requires an update to the corresponding index table, having many indexes adds cumulative overhead to every DML statement. This lengthens the time required to commit transactions, which can eventually lead to governor limit issues and performance bottlenecks in high-volume environments.

Why this answer

While indexes improve read performance, they impose a cost on write operations. Every time a record is inserted, updated, or deleted, Salesforce must update all associated index tables. With an excessive number of indexes, this overhead becomes significant, leading to slower transaction times, potential lock contention, and overall system instability during high-volume DML activities.

Balancing read performance needs with write performance costs is critical in data architecture.

Exam trap

Many candidates assume indexes only have positive impacts, forgetting that database maintenance of numerous indexes severely degrades DML performance.

18
MCQmedium

Why is it recommended to perform large data deletes using a soft-delete approach followed by a hard-delete during off-peak hours?

A.To keep the Recycle Bin empty at all times.
B.To allow for data recovery if a mistake is made.
C.To prevent record locking and system-wide performance degradation.
D.To ensure that all triggers are fired during the deletion.
AnswerC

Large-scale deletions cause intense row-level and table-level locking. Breaking this into stages—marking records first and deleting them in batches during off-peak windows—reduces the contention for system resources, ensuring that the database remains responsive for other users while the deletion operation progresses.

Why this answer

Large-scale deletions are resource-intensive and can trigger cascading deletes if records have child relationships. By first marking records for deletion (soft-delete), you can manage the process in controlled batches. This prevents the system from locking up during a massive operation.

Performing the actual hard-delete during off-peak hours minimizes the impact on concurrent user activity and reduces the risk of reaching governor limits during peak times.

Exam trap

Candidates often try to delete records in bulk without considering the impact of cascading deletes or the performance hit of immediate record removal, causing system timeouts.

19
Multi-Selecthard

Which TWO of the following are true regarding the impact of 'Formula Fields' on Large Data Volumes?

Select 2 answers
A.Formula fields are indexed by default.
B.Filtering on formula fields causes full table scans.
C.They improve query performance by pre-calculating data.
D.Values are stored in the database for fast retrieval.
E.Complex formulas can lead to 'CPU Time Exceeded' errors.
AnswersB, E

Because formula fields are not indexed and must be calculated on the fly during a query, the Salesforce database engine must perform a full table scan to evaluate every record against the filter criteria. For objects with millions of records, this is a major performance bottleneck.

Why this answer

Formula fields are calculated at runtime, which means their values are not stored in the database. When querying or filtering on a formula field, the system must compute the value for every record, which is computationally expensive. As the data volume grows, this causes significant performance issues because the database cannot utilize indexes on formula fields (unless they are deterministic and specifically indexed), leading to slow, resource-intensive queries.

Exam trap

Candidates often assume formula fields are stored in the database and indexed like standard fields. They fail to realize that calculations happen at runtime, forcing full table scans on large datasets.

20
MCQmedium

A company maintains a custom object Asset__c with 5 million records. They need to archive records older than 7 years to a Big Object for compliance. After archiving, the records must be queryable via SOQL for audits, but not editable. Which approach should an architect recommend?

A.Create a Big Object with the same fields, use a batch Apex job to copy old records, and then delete them from Asset__c. Provide audit access via a Lightning component that uses Async SOQL to query the Big Object.
B.Create a Big Object with the same fields as Asset__c, use a batch Apex job to copy old records, and then delete them from Asset__c. Provide audit access via a Visualforce page that queries the Big Object using SOQL.
C.Create a custom object Asset_Archive__c to store old records, and use a batch Apex job to move them. Provide audit access via standard reports and list views.
D.Use Salesforce Connect to create an external object that points to an external database containing the archived records, and provide audit access via external object list views.
AnswerA

Big Objects are designed for large-scale archival and are queried using Async SOQL, which runs asynchronously and can handle massive datasets. Copying old records to a Big Object and deleting them from Asset__c reduces the size of the transactional object, improving performance. Async SOQL is the correct mechanism for querying Big Objects and can be invoked from a Lightning component, satisfying the audit requirement.

Why this answer

Big Objects are the Salesforce-native solution for archiving massive datasets. They support high-volume storage and are queried asynchronously using Async SOQL, which is suitable for audit scenarios where real-time access is not required. Copying old records to a Big Object and deleting them from the transactional object reduces the object's size, improving query and DML performance.

A Lightning component can invoke Async SOQL to retrieve archived records for audits.

Exam trap

The trap here is assuming that Big Objects can be queried with standard SOQL like other objects, when they actually require Async SOQL or a custom index-based query.

21
MCQmedium

When designing a system for LDV, what is the role of an 'Indexed Field'?

A.To increase storage capacity in the Salesforce database.
B.To speed up the retrieval of records by narrowing search space.
C.To enforce uniqueness constraints on every custom field.
D.To automate the conversion of records to external objects.
AnswerB

Indexes function like a book index, allowing the database to jump straight to the data rows that match the filter criteria. This eliminates the need for full table scans, which are prohibitively slow on tables with millions of records, and is vital for maintaining responsive application performance.

Why this answer

Indexes are structures that the database uses to find records quickly without scanning the entire table. In an LDV environment, they are the primary defense against query timeouts. By ensuring that frequently used filters in SOQL queries are indexed, architects can significantly reduce the amount of data the database engine must process, keeping performance within acceptable levels regardless of the total record count.

Exam trap

Test-takers frequently assume that creating an index will automatically speed up every type of query, ignoring that indexes only help specific filter conditions and can actually degrade write performance.

22
MCQhard

A Salesforce org has a custom object Shipment__c with 12 million records. Users frequently run a SOQL query that filters on Status__c and Order_Date__c, and sorts by CreatedDate. The query is timing out. The org has an index on Status__c and a custom index on Order_Date__c. What is the most likely reason the query is still slow?

A.The query is slow because it sorts by CreatedDate, which is not indexed, and Salesforce cannot sort large result sets without an index on the sort field.
B.The query is non-selective because the combination of filters does not reduce the result set below the selectivity threshold, and sorting by CreatedDate requires additional processing.
C.The custom index on Order_Date__c is not used because Salesforce does not support custom indexes on date fields.
D.The query is slow because Status__c is a picklist field, and picklist fields cannot be indexed.
AnswerB

Salesforce selectivity thresholds depend on the object size; for 12 million records, a query must filter to a small percentage to use an index. Even with indexes on Status__c and Order_Date__c, if the filters are not restrictive enough, the optimizer may choose a full scan. Sorting by CreatedDate adds a sort operation that can further degrade performance when the result set is large. This combination explains the timeout.

Why this answer

For a 12-million-record object, selectivity is the critical factor. Salesforce uses a threshold based on object size to decide whether to use an index. Even with indexes on Status__c and Order_Date__c, if the filters return too many rows, the optimizer performs a full scan.

Sorting by CreatedDate adds overhead. The correct diagnosis is that the query is non-selective and the sort compounds the cost.

Exam trap

The trap here is assuming that having indexes on individual fields guarantees query performance, when the real determinant is whether the combined filters meet the selectivity threshold for that object's size.

23
MCQmedium

Universal Containers has 50 million records in a custom object. A report summarizing these records times out. Which strategy is most effective for improving report performance?

A.Increase the report timeout limit in Setup.
B.Implement a custom lightning web component for reporting.
C.Create a summary object to store pre-aggregated data.
D.Add more indexes to the custom object fields.
AnswerC

Pre-aggregation shifts the computational load from read-time to write-time. By calculating metrics asynchronously and storing them in a summary object, the report engine only processes a small number of aggregate records rather than scanning the entire base table, significantly reducing execution time and resource consumption.

Why this answer

Summarizing 50 million records directly in a standard report is inefficient due to the overhead of the reporting engine. Using a Summary Object allows Salesforce to pre-aggregate data asynchronously, moving the heavy lifting from query time to insert time. This approach ensures reports operate on a significantly smaller dataset, preventing timeout errors and improving user experience for large-scale data analysis scenarios.

Exam trap

Candidates often suggest using standard Salesforce reports or roll-up summary fields on large objects, forgetting that standard reports time out on millions of rows and roll-ups have hard limits.

24
MCQmedium

A company is importing 50 million records into a custom object. Which strategy should be used to minimize record locking contention during the high-volume insert operation?

A.Perform the load using the standard Salesforce UI 'Import Wizard'.
B.Increase the batch size to the maximum allowed limit of 10,000.
C.Sort the data by OwnerID and process records in serial mode.
D.Disable all Validation Rules and Apex Triggers permanently.
AnswerC

Sorting by OwnerID or a parent lookup field allows the system to process related records in a predictable order. By avoiding concurrent updates to the same parent or owner records, you reduce the likelihood of row-level lock contention, ensuring that the database engine can process the transaction blocks without waiting for locks.

Why this answer

Minimizing locking contention requires optimizing the database interaction patterns within Salesforce. Serializing record processing and organizing batches by OwnerID or AccountID prevents multiple threads from attempting to lock the same parent records simultaneously. This approach ensures that the database index updates are serialized per specific record groups, drastically reducing the incidence of 'UNABLE_TO_LOCK_ROW' errors during large volume data loads in a multi-tenant environment.

Exam trap

Candidates often recommend parallel processing to speed up imports, completely ignoring how parallel threads exacerbate row-level locking on shared records.

25
MCQmedium

A Salesforce org has a custom object Invoice__c with 8 million records. The business requires a dashboard that shows the total invoice amount grouped by Account and by Fiscal Year. The dashboard must refresh quickly, even during peak usage. The Invoice__c object has a lookup to Account and a formula field Fiscal_Year__c that extracts the year from Invoice_Date__c. What is the most appropriate design to support this dashboard efficiently?

A.Enable Big Object indexing on Invoice__c and point the dashboard to a Big Object-backed report.
B.Build a custom Lightning Web Component that calls Apex to aggregate Invoice__c records on demand and display the results on the dashboard.
C.Create a summary report grouped by Account and Fiscal_Year__c, and add it as a dashboard component with a scheduled refresh.
D.Create a custom object Invoice_Summary__c that stores pre-aggregated totals per Account and Fiscal Year, and update it via batch Apex or scheduled flow when invoices change.
AnswerD

Pre-aggregating into a summary object reduces the dashboard query to a small number of rows, so it remains fast regardless of Invoice__c volume. Batch Apex or scheduled flow can maintain the summary asynchronously, avoiding user-facing delays. This pattern is a standard LDV optimization because it decouples heavy aggregation from interactive reporting and keeps dashboard components selective and lightweight.

Why this answer

Pre-aggregating data into a summary object is the most reliable way to keep a dashboard fast when the source object has millions of records. By storing totals per Account and Fiscal Year, the dashboard queries a small, indexed dataset instead of scanning Invoice__c. Asynchronous maintenance via batch Apex or scheduled flow avoids impacting user transactions and keeps the summary current enough for reporting.

Exam trap

The trap here is assuming that a summary report or on-demand Apex aggregation can scale to millions of records without pre-aggregation, when in fact formula-field grouping forces full scans and governor limits apply.

26
MCQhard

You are auditing a Salesforce environment and discover a custom object with 20 million records. Users report that searching for records by a custom field 'External_ID__c' is extremely slow. What is the most appropriate action to take?

A.Create a custom report type to better filter the data.
B.Mark the field as an External ID to trigger an automatic index.
C.Increase the batch size of the user's list view.
D.Convert the custom object into a Big Object.
AnswerB

Marking a field as an External ID (or unique) creates a database index. This is a standard Salesforce optimization for large objects, ensuring that queries filtering on that field become highly performant by avoiding full table scans, which are the primary bottleneck for large-scale record retrieval.

Why this answer

When an object has large data volumes, non-indexed fields cause full table scans, which are highly inefficient and slow. By marking the 'External_ID__c' field as an External ID or unique field, Salesforce automatically creates a database index. This allows the query engine to perform a targeted lookup rather than scanning all 20 million records, providing an immediate and significant performance improvement for queries and searches involving that specific field.

Exam trap

Candidates often assume that standard fields like Name or Id are the only ones automatically indexed, forgetting that custom fields require explicit configuration like marking them as External ID or Unique.

27
MCQmedium

A custom object 'Invoice__c' contains 8 million records and has a lookup to Account. A nightly batch job deletes approximately 2 million old invoices using Database.delete() in batches of 200. Users report that the batch sometimes fails with 'UNABLE_TO_LOCK_ROW' errors when running concurrently with account updates. What is the most likely cause of these lock contention errors?

A.The nightly batch job is running in parallel mode, which causes multiple batch threads to compete for the same parent Account locks.
B.The delete operation is acquiring exclusive locks on the parent Account records referenced by the deleted invoices.
C.The Invoice__c records have a master-detail relationship to Account, so deleting them cascades and locks all child records simultaneously.
D.The batch size of 200 is too large, causing the database to lock the entire Invoice__c table during each batch execution.
AnswerB

When deleting child records, Salesforce locks the parent records referenced by lookup relationships to maintain referential integrity. This blocks concurrent updates on those Accounts, causing UNABLE_TO_LOCK_ROW errors. The 2 million deletes touch many parent Accounts, increasing collision probability with the nightly account updates.

Why this answer

Deleting child records in a lookup relationship acquires exclusive locks on the parent records to maintain referential integrity. When concurrent updates target those same parent Accounts, lock contention triggers UNABLE_TO_LOCK_ROW errors. This is a common issue with high-volume deletes on objects with lookups to frequently updated parents.

Exam trap

The trap here is assuming that batch size or batch mode is the root cause, when the real issue is parent record locking inherent to deleting child records with lookup relationships.

28
MCQmedium

A large custom object has 20 million records. A SOQL query is taking too long. What should the architect evaluate first?

A.Check if the object has too many fields.
B.Verify if the query filters are selective and indexed.
C.Increase the organization's total storage limit.
D.Rewrite the query to use the Data Loader API.
AnswerB

Selectivity is the most critical factor for SOQL performance on large datasets. If the filter is not selective, the database must scan all 20 million records. Adding an index on the filter field allows the engine to jump directly to the relevant records, restoring query performance.

Why this answer

The architect should first inspect the query's filter criteria to determine if it is selective. If the filter is not indexed or is too broad, the system will perform a full table scan. Verifying the selectivity of the query and ensuring an index exists is the fundamental first step in optimizing performance for large data volumes within the Salesforce platform.

Exam trap

Candidates often suggest adding more filters or changing the query syntax without checking if the fields are actually indexed, which is the root cause of performance issues.

29
MCQmedium

A Salesforce org has a custom object Case_Comment__c with 5 million records. The object has a lookup to Case. Users frequently run SOQL queries that filter by CaseId and OrderBy CreatedDate. The queries are slow. What should be done to improve performance?

A.Create a separate custom index on CreatedDate only.
B.Use a formula field to combine CaseId and CreatedDate.
C.Create a composite custom index on CaseId and CreatedDate.
D.Enable skinny table for Case_Comment__c.
AnswerC

A composite index on both the filter field (CaseId) and the sort field (CreatedDate) allows the database to efficiently satisfy the query's WHERE and ORDER BY clauses. This reduces the need for sorting and scanning, significantly improving performance for large data volumes.

Why this answer

When queries filter on one field and sort by another, a composite index covering both fields allows the database to use the index for both operations, avoiding expensive sorting and full scans. This is the most effective optimization for the given query pattern.

Exam trap

The trap here is assuming that indexing only the sort field or using a formula field will help, but the filter field must also be indexed, and composite indexes are key for combined filter-sort queries.

30
MCQhard

A data architect is designing a solution for a custom object Order__c that will contain 30 million records. Users need to frequently query orders by Customer__c (a lookup to Account) and Order_Date__c. The architect plans to create a composite custom index on Customer__c and Order_Date__c. Which consideration is most critical for the index to be used by the query optimizer?

A.The composite index must include at least three fields to be effective for large data volumes.
B.The composite index must be created on fields that are not already indexed by default.
C.The fields in the composite index must be in the same order as they appear in the SOQL WHERE clause.
D.The leading field in the composite index must be highly selective for the queries being executed.
AnswerD

For a composite index to be used, the leading field must be selective. In this scenario, Customer__c is likely to be selective if queries filter by a specific customer. If the leading field is not selective, the optimizer may ignore the index entirely. Thus, ensuring the leading field's selectivity is critical.

Why this answer

For a composite index to be used by the query optimizer, the leading field must be selective. In this scenario, queries filter by Customer__c and Order_Date__c. If Customer__c is highly selective (e.g., a specific customer), the index will be used.

If Customer__c is not selective, the optimizer may perform a full scan. Therefore, the most critical consideration is the selectivity of the leading field.

Exam trap

The trap here is focusing on the order of fields in the WHERE clause or the number of fields, rather than the selectivity of the leading field in the composite index.

31
MCQmedium

What is the recommended approach to manage 'Skinny Tables' in an environment with Large Data Volumes?

A.Enable skinny tables for every object with more than 1 million records.
B.Request them only for high-read-volume objects where performance is critical.
C.Manually update skinny tables using DML statements in Apex.
D.Always include every field of the object in the skinny table.
AnswerB

Skinny tables are designed to solve specific performance issues by reducing join complexity. Because they are maintained by Salesforce Support, they add administrative overhead. They are best reserved for core objects where read performance is the primary bottleneck and standard indexing cannot provide the necessary speed.

Why this answer

Skinny tables are a specialized feature used to improve performance by flattening data structures to avoid joins. They should only be used when standard indexing is insufficient for performance. Because they require Salesforce Support to maintain and are automatically synchronized, they are a powerful but rigid tool that should be applied sparingly to the most critical, high-read-volume scenarios only, to avoid unnecessary maintenance overhead.

Exam trap

Candidates frequently view Skinny Tables as a general performance optimization for all objects. They fail to consider the high maintenance cost and the restriction that they are only for specific scenarios.

32
MCQhard

Which TWO strategies help manage row locking when inserting records into an object with multiple lookup relationships?

A.Sort the input data by the lookup field values.
B.Use the Bulk API 2.0 in serial mode.
C.Disable unnecessary triggers on the parent objects.
D.Increase the batch size to 10,000.
E.Convert all lookups to master-detail.
AnswerA, C

Sorting the input file by the parent ID ensures that all children associated with the same parent are grouped together. This serializes access to the parent records, preventing multiple batch threads from attempting to lock the same parent concurrently, which is the primary cause of row contention.

Why this answer

High-volume loads often fail due to locking on parent records referenced by lookups. Sorting the input data by the lookup ID ensures that all children for a single parent are processed together, reducing the time a parent record remains locked. Additionally, avoiding unnecessary triggers on the parent object prevents downstream locking, significantly improving throughput for bulk data migrations.

Exam trap

Candidates often focus solely on chunking batch sizes during data loads while ignoring the record ordering strategy that prevents concurrent database locks on shared parent records.

33
MCQmedium

A custom object named Invoice__c contains 15 million records. Reports and list views frequently filter on a custom date field, Invoice_Date__c, and are timing out. The field is not indexed. Which action should a data architect take to improve query performance while keeping the field available for filtering?

A.Add a formula field that returns the value of Invoice_Date__c and filter on it instead.
B.Convert Invoice_Date__c to an external lookup field using Salesforce Connect.
C.Enable Divisions on the Invoice__c object to partition the data.
D.Create a custom index on Invoice_Date__c because it is used as a filter condition.
AnswerD

A custom index can be requested on a custom field that is frequently used in filter conditions. For an object with millions of records, indexing Invoice_Date__c allows the query optimizer to avoid a full table scan, reducing report and list view timeouts. This is the appropriate declarative performance tuning step for LDV scenarios.

Why this answer

For large data volumes, selective queries require indexed filter fields. Requesting a custom index on a frequently filtered custom date field enables the query optimizer to use the index and avoid full table scans. This directly addresses the report and list view timeouts without changing the data model or introducing external dependencies.

Exam trap

The trap here is assuming that any frequently filtered field is automatically indexed or that a formula field can substitute for an index.

34
MCQmedium

What is the primary function of a skinny table in Salesforce?

A.To provide more storage for large objects.
B.To improve performance by avoiding table joins.
C.To automatically archive old records.
D.To encrypt data for security compliance.
AnswerB

By denormalizing data into a single table, the database no longer needs to perform costly join operations during query execution. This drastically reduces the time required to retrieve large sets of data, providing a significant performance boost for reports and SOQL queries in high-volume environments.

Why this answer

Skinny tables are used to optimize performance by flattening data from multiple related objects into a single table. This removes the need for joins when querying across related records, which significantly speeds up SOQL queries and reporting. They are a powerful tool for large-scale data environments where complex relationship traversal is the main bottleneck for query latency.

Exam trap

Candidates often confuse skinny tables with indexes, assuming they are the same, whereas skinny tables specifically target the performance bottleneck of joining multiple tables together.

35
MCQhard

Refer to the exhibit. What is the most likely reason this query fails?

A.The LIMIT clause is too high for the object size.
B.The field 'Status__c' is not indexed.
C.The query includes too many fields in the SELECT clause.
D.The record count exceeds the object storage limit.
AnswerB

Without an index on the 'Status__c' field, the database must perform a full table scan to evaluate every record. Since the object contains millions of records, this exceeds the system's threshold for selectivity, causing the query to fail to protect overall system performance and availability.

Why this answer

The query fails because 'Status__c' is likely not indexed, or the 'Pending' value is too common, exceeding the selectivity threshold. On objects with millions of records, Salesforce requires that filters target a small enough subset to be efficient. When a filter is too broad or lacks an index, the system rejects the query to prevent performance degradation for the entire multi-tenant environment.

Exam trap

Candidates frequently assume a field is indexed by default, failing to account for the selectivity threshold, where even an indexed field can cause failure if too many records match.

36
MCQmedium

What is the primary architectural benefit of using 'External Objects' (Salesforce Connect) for large volumes of historical data?

A.To speed up the performance of local object triggers.
B.To keep the local Salesforce database lean and performant.
C.To bypass the need for API callout security.
D.To allow for complex joins between local and external objects.
AnswerB

By offloading data to an external system and accessing it via OData, you maintain a smaller record count in the local Salesforce instance. This ensures that indexes remain efficient and queries on local objects stay fast, which is critical for maintaining performance as the organization scales.

Why this answer

External objects allow Salesforce to access data stored in an external system in real-time without importing it into the local database. This keeps the Salesforce storage usage low and avoids the performance impacts associated with hosting millions of records locally. It is an ideal solution for historical or secondary data that needs to be visible in the UI but does not require being in the same transactional database.

Exam trap

Candidates confuse Salesforce Connect with a data replication tool. They incorrectly believe it stores data locally, missing the architectural purpose of keeping the Salesforce database lean to maintain high performance.

Ready to test yourself?

Try a timed practice session using only Large Data Volume Considerations questions.