MLS-C01 Data Engineering Practice Question
A company runs a real-time analytics platform that ingests IoT sensor data from millions of devices. The data is sent to Amazon Kinesis Data Streams with 16 shards. A custom Java application using the Kinesis Client Library (KCL) processes the data and writes aggregated results to Amazon DynamoDB. The application runs on a fleet of EC2 instances in an Auto Scaling group. Recently, the team noticed that some records are being processed multiple times, resulting in duplicate entries in DynamoDB. The application uses the DynamoDB PutItem API to write records. The team needs to eliminate duplicates without significantly increasing latency. Which solution should the team implement?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use DynamoDB TransactWriteItems with a condition check that the record's Kinesis sequence number does not already exist in the table.
Using a DynamoDB transaction with a condition check on the Kinesis sequence number ensures that each record is written only once. Option A is wrong because increasing write capacity does not address duplicate processing; duplicates arise from the KCL consumer processing records multiple times, not from throttling. Option C is wrong because while SQS FIFO provides deduplication at the queue level, it does not guarantee exactly-once processing downstream in DynamoDB; the consumer could still write duplicates if it fails after writing but before deleting the message. Additionally, adding an extra queue increases latency and complexity. Option D is wrong because BatchWriteItem does not decrease duplicates; it only batches multiple put requests into one API call and still requires idempotency measures.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable DynamoDB auto scaling to increase write capacity and reduce throttling, which causes retries and duplicates.
Why it's wrong here
Idempotent writes would require a unique identifier; using PutItem with a condition expression on a unique attribute (like the sequence number) is effectively the same as option B but transactions provide atomicity.
- ✓
Use DynamoDB TransactWriteItems with a condition check that the record's Kinesis sequence number does not already exist in the table.
Why this is correct
Using a DynamoDB transaction with a condition check on the Kinesis sequence number ensures that each record is written only once.
- ✗
Place an Amazon SQS FIFO queue between the KCL application and DynamoDB to deduplicate messages.
Why it's wrong here
Placing an SQS FIFO queue adds latency and complexity without solving the duplicate write issue. The KCL already provides at-least-once delivery, so the consumer must handle deduplication using conditional writes based on the Kinesis sequence number.
- ✗
Modify the application to use DynamoDB BatchWriteItem instead of PutItem to reduce the number of write requests.
Why it's wrong here
Adding a FIFO queue adds latency and complexity without guaranteeing exactly-once processing in the consumer.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLS-C01 question from scratch — 1,672 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.