Question 1,312 of 1,660
SAP-C02 Practice Question: Accelerate Workload Migration and Modernization
A company is migrating a large-scale data analytics workload from on-premises to AWS. The workload uses Apache Spark to process terabytes of data daily. The company wants to use Amazon EMR for the migration. The current on-premises cluster has 20 nodes, each with 64 vCPUs and 256 GB of RAM. The data is stored in HDFS on the cluster. The company wants to minimize costs while maintaining performance. The data sources are in Amazon S3 and on-premises. The company has set up a dedicated AWS Direct Connect connection. Which EMR configuration should the company use?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use EMR with EC2 instances of similar size (e.g., r5.8xlarge) and store data in Amazon S3 using EMRFS.
It directly maps the on-premises cluster capacity to equivalent EC2 instances (r5.8xlarge provides 32 vCPUs and 256 GB RAM, so two per node would match the 64 vCPUs and 256 GB RAM). Storing data in S3 via EMRFS eliminates HDFS management and leverages S3 durability and scalability. This configuration minimizes costs by avoiding over-provisioning and using S3 for cost-effective storage, while maintaining performance through Direct Connect for data transfer. Using EMR with similar instance sizes ensures the Spark jobs run efficiently without reconfiguration. Other options introduce unnecessary complexity or higher costs: Option B uses Graviton-based instances which may require code changes, and storing intermediate data on EBS volumes incurs additional costs and management overhead. Option C mixes On-Demand and Spot Instances, which can reduce costs but still requires cluster management and does not align with the goal of minimizing costs while maintaining performance as effectively as option A. Option D uses AWS Glue, which is a fully managed service but may not provide the same level of control or performance for large-scale workloads; it may also be more expensive for sustained high-volume processing compared to EMR with reserved capacity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use EMR with EC2 instances of similar size (e.g., r5.8xlarge) and store data in Amazon S3 using EMRFS.
Why this is correct
Correct. Using EMR with r5.8xlarge instances closely matches the on-premises resources (64 vCPUs, 256 GB RAM per node) and using S3 with EMRFS eliminates HDFS overhead, reducing costs and management effort while maintaining performance via Direct Connect.
- ✗
Use EMR with Graviton-based instances and store intermediate data in HDFS on EBS volumes.
Why it's wrong here
Incorrect. Graviton-based instances may not be fully compatible with all Apache Spark workloads and could require code changes. Additionally, storing intermediate data in HDFS on EBS volumes adds cost and complexity, and does not leverage S3's durability or cost advantages.
- ✗
Use EMR with a mix of On-Demand and Spot Instances, and use S3 for all data storage.
Why it's wrong here
Using S3 for all data storage is a sound cloud-native approach, and mixing On-Demand and Spot Instances effectively minimises costs, aligning with the company's objective. However, this option fails to address the immediate need to process existing terabytes of data stored in on-premises HDFS. Migrating all this data to S3 *before* processing would introduce significant initial data transfer costs and delays, despite the Direct Connect. This configuration would be correct for a fully migrated, cloud-native workload where all data already resides in S3, leveraging its scalability and cost-effectiveness for decoupled storage.
- ✗
Use AWS Glue to run the Spark jobs with the same resource configuration.
Why it's wrong here
Incorrect. AWS Glue is a serverless ETL service, but for a large-scale daily workload with terabytes of data, Glue may incur higher costs than EMR, and it offers less control over cluster configuration and performance tuning. The question specifically asks for an EMR configuration, so Glue is not an appropriate choice.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
About these practice questions
Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →
Last reviewed: Jun 20, 2026
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.
Question Discussion
Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.
Sign in to join the discussion.