DEA-C01 Data Operations and Support Practice Question
A company runs a data processing pipeline on Amazon EMR. The pipeline reads data from S3, processes it with Spark, and writes results back to S3. The engineer notices that the cluster is underutilized and wants to reduce costs. Which TWO actions should the engineer take? (Choose TWO.)
⚠ Common exam trap
Watch out — candidates often confuse cost optimization features like Spot instances and auto-termination with performance improvements or data consistency settings, leading them to select options that increase resources or enable features unrelated to cost reduction.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Spot instances for task nodes.
Option A is correct because using Spot instances for task nodes is a standard EMR cost-optimization technique: task nodes perform only HDFS-non-persistent work (Spark executors), so they can be interrupted without data loss, and Spot capacity typically costs significantly less than On-Demand. Option B is correct because configuring the cluster to terminate after the job completes (a transient cluster, e.g., via --auto-terminate or a termination-protected=false setting) stops paying for idle EC2 instances once the batch pipeline finishes, directly addressing the underutilization. Option C is wrong because enlarging the master node does not improve processing throughput and increases cost, since the master only manages the cluster. Option D is wrong because EMRFS consistent view is a data-consistency feature for S3 reads/writes, not a cost-reduction mechanism. Option E is wrong because adding core nodes increases cost and capacity rather than reducing spend, and the cluster is already underutilized.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Spot instances for task nodes.
Why this is correct
Spot instances suit task nodes because EMR task nodes hold no HDFS data, so interruption only loses in-flight Spark tasks, which rerun. This directly addresses the underutilised cluster's cost problem by cutting compute spend on the elastic portion of the cluster.
- ✓
Configure the cluster to terminate after the job completes.
Why this is correct
Transient clusters terminate automatically once all steps finish, so no idle core or task nodes keep billing after the Spark job writes results to S3. This removes the waste caused by the underutilisation the engineer observed.
- ✗
Change the master node to a larger instance type.
Why it's wrong here
A larger master node adds no Spark executor capacity; the master only runs the resource manager and NameNode, so underutilised worker capacity stays unchanged while cost rises. Larger master instances suit clusters whose master is the bottleneck through many concurrent jobs or heavy HDFS metadata operations.
- ✗
Enable EMRFS consistent view.
Why it's wrong here
EMRFS consistent view adds DynamoDB-backed metadata tracking for S3 read-after-write consistency, consuming extra resources without reclaiming idle worker capacity. It suits pipelines needing immediate consistency when listing or reading newly written S3 objects, not cost reduction on an underutilised cluster.
- ✗
Increase the number of core nodes to improve parallelism.
Why it's wrong here
Adding core nodes increases worker capacity, worsening underutilisation and raising cost rather than reducing it. Scaling core nodes suits clusters genuinely short of HDFS storage or Spark executors; here the cluster already has spare capacity, so the fix is shrinking or removing nodes instead.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.