A developer is creating a new DynamoDB table to store order data. The orders have a unique order ID and are retrieved by order ID. Occasionally, the developer needs to query orders by customer ID. Which design approach would minimize costs and provide the fastest queries?
This design effectively supports two distinct access patterns: retrieving a specific order by its unique order ID using a highly efficient GetItem operation, and querying all orders associated with a particular customer ID. By establishing a Global Secondary Index (GSI) with customer ID as its partition key, DynamoDB can efficiently retrieve all items matching that customer, optimizing performance and minimizing read capacity unit consumption for both primary and secondary query types.
Why this answer
Using the order ID as the partition key ensures the most efficient primary key access for the primary query pattern (retrieving by order ID). Creating a Global Secondary Index (GSI) on customer ID allows efficient querying by customer ID without scanning the base table, and GSIs have separate read/write capacity from the base table, so you only pay for the index when it is used. This design minimizes costs by avoiding unnecessary scans and provides the fastest queries for both access patterns.
Exam trap
The trap here is that candidates often choose Option B (customer ID as partition key) thinking it naturally supports both access patterns, but they overlook the hot partition problem and the fact that retrieving a single order by order ID would require a scan or a query with a known customer ID, which is not always available.
How to eliminate wrong answers
Option B is wrong because using customer ID as the partition key would cause all orders for the same customer to be stored in the same partition, leading to hot partitions and potential throttling, and it does not provide efficient retrieval by order ID (which would require a scan or a query with a known customer ID). Option C is wrong because scanning the entire table to find orders by customer ID is extremely inefficient and costly, as it reads every item in the table and incurs read capacity for all items, even those not matching the query. Option D is wrong because a Local Secondary Index (LSI) requires the same partition key as the base table (customer ID), which would still cause hot partitions for high-volume customers, and LSIs share the base table's read/write capacity, so they do not provide the same cost flexibility as a GSI.