A data engineer needs to migrate an on-premises Apache Hadoop cluster to AWS. The cluster stores data in HDFS and runs MapReduce jobs. The company wants to minimize operational overhead and leverage serverless technologies where possible. Which AWS service should the data engineer use to replace HDFS storage?
Amazon S3 provides durable, serverless object storage that replaces HDFS without cluster management, satisfying the minimise-operational-overhead constraint. MapReduce workloads migrate to Amazon EMR or Athena, while S3 becomes the decoupled storage layer, eliminating NameNode and DataNode administration entirely.
Why this answer
Amazon S3 is the correct replacement for HDFS because it provides highly durable, scalable, and serverless object storage that can be used as the primary storage layer for Amazon EMR. Unlike HDFS, S3 decouples storage from compute, eliminating the need to manage cluster storage and allowing jobs to run on ephemeral clusters, which minimizes operational overhead. S3 integrates with EMR via the EMR File System (EMRFS), enabling MapReduce jobs to read/write data directly from S3 as if it were HDFS.
Exam trap
The trap here is that candidates confuse Amazon EMR (a compute service) with a storage service, assuming it replaces HDFS, when in fact EMR can use either HDFS or S3 for storage, and the question explicitly asks for the storage replacement.
How to eliminate wrong answers
Option A is wrong because Amazon EBS provides block-level storage volumes attached to EC2 instances, which is not serverless and requires manual management of volume size, snapshots, and replication; it also ties storage to a specific compute instance, defeating the purpose of decoupling storage from compute for a Hadoop migration. Option B is wrong because Amazon EMR is a managed big data platform that runs MapReduce jobs, not a storage service; it can use HDFS or S3 for storage, but the question specifically asks for a replacement of HDFS storage, not the compute framework. Option D is wrong because Amazon Redshift is a fully managed data warehouse optimized for SQL-based analytics and structured data, not a general-purpose distributed file system for Hadoop workloads; it does not support HDFS semantics or MapReduce jobs natively.