Courseiva

Migrate Hadoop to Amazon EMR: Decouple Compute and Storage Using S3 and EMRFS

A company is migrating a large Hadoop cluster to Amazon EMR. The cluster uses HDFS for storage. The company wants to decouple compute and storage to reduce costs. Which approach should the company take?

Quick Answer

The correct approach is to use Amazon S3 as the data store with EMRFS. This decouples compute and storage because S3 is a separate, durable object storage layer, while EMRFS provides the Hadoop-compatible file system interface that allows EMR clusters to read and write data directly to S3 as if it were HDFS. By eliminating the need to replicate data across cluster nodes, you can spin down compute when not in use and only pay for storage separately, drastically reducing costs. On the AWS Certified Solutions Architect Professional SAP-C02 exam, this scenario tests your understanding of the Hadoop to Amazon EMR migration pattern and the core principle of separating compute from storage to optimize for variable workloads. A common trap is choosing EBS-backed instances, which keep storage tied to the cluster’s lifecycle, or FSx for Lustre, which is designed for high-performance computing but not for true decoupling. Memory tip: think “S3 + EMRFS = Storage Freedom” — the FS stands for “file system,” but the key is that S3 lives outside the cluster.

⚠ Common exam trap

SAP-C02 often tests whether candidates understand the difference between decoupling storage and compute versus using a high-performance file system — the trap is choosing FSx for Lustre when the primary goal is cost reduction through S3 decoupling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Amazon S3 as the data store and EMRFS

Using Amazon S3 as the data store with EMRFS is correct because it decouples compute and storage — EMR clusters can read/write directly to S3, and the cluster can be terminated without data loss. This reduces costs by allowing transient clusters and eliminating HDFS replication overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Amazon S3 as the data store and EMRFS

    Why this is correct

    EMRFS lets EMR read and write HDFS-compatible data directly in Amazon S3, so the cluster's storage layer persists independently of transient EC2 nodes. This decouples compute from storage, allowing clusters to be terminated between jobs and cutting the cost of maintaining replicated HDFS volumes.

  • ✗

    Use Amazon FSx for Lustre

    Why it's wrong here

    FSx for Lustre provides a high-performance scratch or persistent filesystem for HPC and EMR processing, but it is not an HDFS-compatible storage layer that EMR can use as a drop-in replacement for decoupling. It suits temporary scratch data alongside S3, not primary cluster storage.

  • ✗

    Use Amazon EFS for HDFS

    Why it's wrong here

    EMR's HDFS is not replaced by EFS; EMRFS is the S3-backed filesystem that decouples compute from storage, and EFS lacks the HDFS API and NameNode semantics EMR expects. EFS suits shared POSIX file access across EC2 instances, not Hadoop storage.

  • ✗

    Use EBS volumes for HDFS

    Why it's wrong here

    EBS volumes attach to individual EC2 nodes, so HDFS data stays tied to the cluster and compute cannot scale independently of storage. It appeals because EBS gives durable block storage for HDFS on self-managed clusters, which suits lift-and-shift migrations where decoupling is not a requirement.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on SAP-C02

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is migrating an on-premises Hadoop cluster to Amazon EMR. The cluster processes large datasets that are stored on HDFS. The company wants to minimize migration effort and cost. Which storage option should the company use with Amazon EMR?

medium
  • A.Amazon EBS volumes attached to the core nodes
  • ✓ B.Amazon S3 with EMRFS
  • C.Amazon FSx for HDFS
  • D.Amazon Elastic File System (EFS)

Why B: Amazon S3 with EMRFS allows EMR to use S3 as a scalable, durable, and cost-effective storage layer, eliminating the need to manage HDFS. Option A is incorrect because EBS volumes require provisioning and management, and do not provide the same scalability and cost benefits. Option C is incorrect because Amazon FSx for HDFS is suitable for workloads requiring HDFS compatibility but is more complex and expensive than using S3. Option D is incorrect because Amazon EFS is a file system, not optimized for Hadoop workloads and typically used for shared file storage.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.