DEA-C01 Data Operations and Support Practice Question
A data engineer is designing a data pipeline that processes streaming data. The pipeline must be able to handle duplicate records and ensure exactly-once processing semantics. Which THREE AWS services or features should the engineer consider? (Choose three.)
⚠ Common exam trap
DEA-C01 often tests the misconception that Firehose retries provide exactly-once, but retries can cause duplicates; candidates must recognize that exactly-once requires specific features like Flink checkpointing and idempotent sinks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon EMR with Apache Flink for exactly-once semantics.
Option A is correct because Amazon EMR running Apache Flink supports exactly-once state consistency via its checkpointing mechanism (distributed snapshots aligned with barriers), so a streaming pipeline built on Flink can guarantee exactly-once processing even with duplicate or replayed records. Option C is correct because Kinesis Data Streams assigns each record a unique sequence number within a shard, and consumers can use those sequence numbers (or the KCL's checkpointing) to detect and skip already-processed records, enabling deduplication and effectively-once consumption. Option E is correct because Kinesis Data Analytics for Apache Flink inherits Flink's exactly-once checkpointing and, when combined with idempotent sinks (e.g., writing to DynamoDB with conditional writes or to S3 with deterministic keys), prevents duplicate records from producing duplicate downstream effects. Option B is not appropriate because Kinesis Data Firehose delivers at-least-once and its automatic retries can actually introduce duplicates rather than deduplicate them. Option D is not appropriate because DynamoDB Streams provides change data capture with at-least-once delivery and no built-in deduplication or exactly-once guarantee for a streaming pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Amazon EMR with Apache Flink for exactly-once semantics.
Why this is correct
Apache Flink on Amazon EMR provides checkpointing and two-phase commit sinks, giving exactly-once state consistency across operator failures. This satisfies the stem's requirement to tolerate duplicate records while guaranteeing each record affects downstream state only once.
- ✗
Amazon Kinesis Data Firehose with automatic retries.
Why it's wrong here
Firehose retries redeliver batches after failures, which can duplicate records rather than guarantee exactly-once semantics. It suits loading streaming data into S3, Redshift or OpenSearch without custom consumers, but its at-least-once delivery model cannot satisfy the deduplication requirement here.
- ✓
Amazon Kinesis Data Streams with sequence numbers for deduplication.
Why this is correct
Each Kinesis Data Streams record carries a sequence number unique within its shard, letting consumers detect and discard reprocessed records after retries. This provides the deduplication mechanism needed to approximate exactly-once delivery in the pipeline.
- ✗
Amazon DynamoDB Streams for change data capture.
Why it's wrong here
DynamoDB Streams captures item-level changes for event-driven replication, not stream deduplication or checkpointed exactly-once delivery; it would be the right pick for syncing table changes into a downstream store. The pipeline needs Kinesis Client Library checkpointing and idempotent writes instead.
- ✓
Amazon Kinesis Data Analytics for Apache Flink with idempotent sinks.
Why this is correct
Apache Flink on Kinesis Data Analytics provides checkpointing and exactly-once state consistency, so duplicate records are deduplicated within the application state. Writing to idempotent sinks ensures replayed records after recovery do not create duplicates downstream, directly satisfying the pipeline's exactly-once requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.