Courseiva

SAP-C02 Practice Question: Accelerate Workload Migration and Modernization

A company is planning to migrate a large-scale Hadoop cluster to Amazon EMR. The cluster currently processes batch jobs using a mix of MapReduce and Spark. The company wants to minimize changes to the existing code and operational processes. Which migration approach should the architect recommend?

⚠ Common exam trap

SAP-C02 often tests migration strategy selection — the trap is choosing 'refactor' or 'replatform' when the scenario explicitly says 'minimize changes to existing code,' which points to rehost.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Rehost the cluster on Amazon EMR using the same MapReduce and Spark configurations

Rehosting the Hadoop cluster on Amazon EMR with the same MapReduce and Spark configurations minimizes code changes and operational disruption, which is exactly what the company wants. EMR supports both MapReduce and Spark natively, so existing jobs can run with minimal modification. This is the classic lift-and-shift approach for Hadoop-to-EMR migrations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Refactor all jobs to use only Apache Spark on Amazon EMR

    Why it's wrong here

    Refactoring all jobs to Spark rewrites existing MapReduce code, directly violating the minimise-changes requirement. It is tempting because Spark is the modern EMR engine and would be correct for a greenfield build, but the stem's mixed MapReduce and Spark estate must run largely unmodified, which EMR supports natively.

  • ✗

    Retire the cluster and use Amazon Athena for ad-hoc queries

    Why it's wrong here

    Athena queries S3 data with SQL and cannot execute existing MapReduce or Spark code, so the company's jobs would need rewriting. It is tempting because Athena suits serverless ad-hoc querying over data lake files, which would be the right choice if the requirement were interactive SQL analysis rather than running existing batch applications.

  • ✗

    Replatform the data processing to use Amazon Redshift Spectrum

    Why it's wrong here

    Redshift Spectrum queries S3 data through Redshift, requiring the Hadoop jobs and their code to be rewritten as SQL, breaking the minimise-changes requirement. It is tempting because Spectrum avoids cluster management and suits analytical querying, but it is not a drop-in execution engine for existing MapReduce or Spark workloads.

  • ✓

    Rehost the cluster on Amazon EMR using the same MapReduce and Spark configurations

    Why this is correct

    Rehosting on Amazon EMR preserves the existing MapReduce and Spark APIs, so job code and operational tooling transfer with minimal modification. This directly satisfies the stem's constraint of minimising code and process changes, unlike replatforming to serverless or rewriting jobs for a different engine.

About these practice questions

Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.