SAP-C02 Practice Question: Accelerate Workload Migration and Modernization
A company is planning to migrate a large-scale Hadoop cluster to Amazon EMR. The cluster currently processes batch jobs using a mix of MapReduce and Spark. The company wants to minimize changes to the existing code and operational processes. Which migration approach should the architect recommend?
⚠ Common exam trap
SAP-C02 often tests migration strategy selection — the trap is choosing 'refactor' or 'replatform' when the scenario explicitly says 'minimize changes to existing code,' which points to rehost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Rehost the cluster on Amazon EMR using the same MapReduce and Spark configurations
Rehosting the Hadoop cluster on Amazon EMR with the same MapReduce and Spark configurations minimizes code changes and operational disruption, which is exactly what the company wants. EMR supports both MapReduce and Spark natively, so existing jobs can run with minimal modification. This is the classic lift-and-shift approach for Hadoop-to-EMR migrations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Refactor all jobs to use only Apache Spark on Amazon EMR
Why it's wrong here
Refactoring all jobs to Spark rewrites existing MapReduce code, directly violating the minimise-changes requirement. It is tempting because Spark is the modern EMR engine and would be correct for a greenfield build, but the stem's mixed MapReduce and Spark estate must run largely unmodified, which EMR supports natively.
- ✗
Retire the cluster and use Amazon Athena for ad-hoc queries
Why it's wrong here
Athena queries S3 data with SQL and cannot execute existing MapReduce or Spark code, so the company's jobs would need rewriting. It is tempting because Athena suits serverless ad-hoc querying over data lake files, which would be the right choice if the requirement were interactive SQL analysis rather than running existing batch applications.
- ✗
Replatform the data processing to use Amazon Redshift Spectrum
Why it's wrong here
Redshift Spectrum queries S3 data through Redshift, requiring the Hadoop jobs and their code to be rewritten as SQL, breaking the minimise-changes requirement. It is tempting because Spectrum avoids cluster management and suits analytical querying, but it is not a drop-in execution engine for existing MapReduce or Spark workloads.
- ✓
Rehost the cluster on Amazon EMR using the same MapReduce and Spark configurations
Why this is correct
Rehosting on Amazon EMR preserves the existing MapReduce and Spark APIs, so job code and operational tooling transfer with minimal modification. This directly satisfies the stem's constraint of minimising code and process changes, unlike replatforming to serverless or rewriting jobs for a different engine.
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.