DEA-C01 Data Store Management Practice Question
A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?
⚠ Common exam trap
Watch out — candidates often think concatenating the row key into a single partition key (Option C) preserves the query pattern, but DynamoDB requires the partition key to be known exactly for queries, making it impossible to query by customer_id alone without a full scan.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a table with partition key = customer_id and sort key = timestamp.
DynamoDB's partition key (customer_id) evenly distributes data across partitions, while the sort key (timestamp) enables efficient range queries using Query with ScanIndexForward=false to retrieve the latest entries. This design directly maps the HBase composite row key pattern to DynamoDB's primary key structure, optimizing for the described access pattern.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a table with partition key = customer_id and sort key = timestamp.
Why this is correct
Splitting the composite row key into partition key customer_id and sort key timestamp lets DynamoDB distribute items across partitions by customer, while the sort key orders entries chronologically. Querying a single customer_id with ScanIndexForward=false returns the latest entries efficiently, satisfying the query-by-customer, retrieve-latest access pattern without full table scans.
- ✗
Use Amazon S3 with customer_id as prefix and timestamp as object name.
Why it's wrong here
S3 is object storage, not a key-value database: it cannot serve DynamoDB's query-by-partition-key access pattern, and listing objects by prefix returns lexicographic, not timestamp-descending, order, so latest-entry retrieval needs full scans. It is tempting as cheap durable storage for HBase exports, and would suit archival or batch analytics, not low-latency point queries.
- ✗
Create a table with partition key = concatenated customer_id and timestamp.
Why it's wrong here
Concatenating customer_id and timestamp into the partition key scatters each customer's entries across partitions, so retrieving the latest entries requires scanning multiple partitions rather than one Query. It is tempting because it mirrors the HBase row key. This design suits point lookups by full composite key, not latest-per-customer retrieval.
- ✗
Create a table with partition key = timestamp and sort key = customer_id.
Why it's wrong here
Partitioning by timestamp spreads a single customer's rows across many partitions, so latest-entry queries cannot target one partition and hot partitions arise from write-time clustering. It is tempting because timestamps order data. This design suits time-range scans across all customers, not per-customer latest retrieval.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.