Migrate Hadoop to Amazon EMR: Decouple Compute and Storage Using S3 and EMRFS
A company is migrating a large Hadoop cluster to Amazon EMR. The cluster uses HDFS for storage. The company wants to decouple compute and storage to reduce costs. Which approach should the company take?
Quick Answer
The correct approach is to use Amazon S3 as the data store with EMRFS. This decouples compute and storage because S3 is a separate, durable object storage layer, while EMRFS provides the Hadoop-compatible file system interface that allows EMR clusters to read and write data directly to S3 as if it were HDFS. By eliminating the need to replicate data across cluster nodes, you can spin down compute when not in use and only pay for storage separately, drastically reducing costs. On the AWS Certified Solutions Architect Professional SAP-C02 exam, this scenario tests your understanding of the Hadoop to Amazon EMR migration pattern and the core principle of separating compute from storage to optimize for variable workloads. A common trap is choosing EBS-backed instances, which keep storage tied to the cluster’s lifecycle, or FSx for Lustre, which is designed for high-performance computing but not for true decoupling. Memory tip: think “S3 + EMRFS = Storage Freedom” — the FS stands for “file system,” but the key is that S3 lives outside the cluster.
⚠ Common exam trap
SAP-C02 often tests whether candidates understand the difference between decoupling storage and compute versus using a high-performance file system — the trap is choosing FSx for Lustre when the primary goal is cost reduction through S3 decoupling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Amazon S3 as the data store and EMRFS
Using Amazon S3 as the data store with EMRFS is correct because it decouples compute and storage — EMR clusters can read/write directly to S3, and the cluster can be terminated without data loss. This reduces costs by allowing transient clusters and eliminating HDFS replication overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Amazon S3 as the data store and EMRFS
Why this is correct
EMRFS lets EMR read and write HDFS-compatible data directly in Amazon S3, so the cluster's storage layer persists independently of transient EC2 nodes. This decouples compute from storage, allowing clusters to be terminated between jobs and cutting the cost of maintaining replicated HDFS volumes.
- ✗
Use Amazon FSx for Lustre
Why it's wrong here
FSx for Lustre provides a high-performance scratch or persistent filesystem for HPC and EMR processing, but it is not an HDFS-compatible storage layer that EMR can use as a drop-in replacement for decoupling. It suits temporary scratch data alongside S3, not primary cluster storage.
- ✗
Use Amazon EFS for HDFS
Why it's wrong here
EMR's HDFS is not replaced by EFS; EMRFS is the S3-backed filesystem that decouples compute from storage, and EFS lacks the HDFS API and NameNode semantics EMR expects. EFS suits shared POSIX file access across EC2 instances, not Hadoop storage.
- ✗
Use EBS volumes for HDFS
Why it's wrong here
EBS volumes attach to individual EC2 nodes, so HDFS data stays tied to the cluster and compute cannot scale independently of storage. It appeals because EBS gives durable block storage for HDFS on self-managed clusters, which suits lift-and-shift migrations where decoupling is not a requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on SAP-C02
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company is migrating an on-premises Hadoop cluster to Amazon EMR. The cluster processes large datasets that are stored on HDFS. The company wants to minimize migration effort and cost. Which storage option should the company use with Amazon EMR?
medium- A.Amazon EBS volumes attached to the core nodes
- ✓ B.Amazon S3 with EMRFS
- C.Amazon FSx for HDFS
- D.Amazon Elastic File System (EFS)
Why B: Amazon S3 with EMRFS allows EMR to use S3 as a scalable, durable, and cost-effective storage layer, eliminating the need to manage HDFS. Option A is incorrect because EBS volumes require provisioning and management, and do not provide the same scalability and cost benefits. Option C is incorrect because Amazon FSx for HDFS is suitable for workloads requiring HDFS compatibility but is more complex and expensive than using S3. Option D is incorrect because Amazon EFS is a file system, not optimized for Hadoop workloads and typically used for shared file storage.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.