Courseiva
Data Store Management →mediumMultiple Choice

DEA-C01 Data Store Management Practice Question

A company is migrating an on-premises Hadoop cluster to AWS. The cluster processes large files in CSV format using Apache Spark. Which data store should be used as the primary storage for the data lake to optimize cost and performance?

⚠ Common exam trap

DEA-C01 often tests the misconception that HDFS or EBS is suitable for a data lake, but they are not cost-effective or scalable for long-term storage. Candidates might choose EMRFS with HDFS due to familiarity with Hadoop.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon S3

Amazon S3 is the optimal primary storage for a data lake on AWS because it offers high durability, scalability, and cost-effectiveness. It integrates seamlessly with Apache Spark on Amazon EMR, allowing direct access to data without moving it. S3 also supports various file formats and decouples storage from compute, enabling independent scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon EMR File System (EMRFS) backed by HDFS

    Why it's wrong here

    EMRFS backed by HDFS keeps data on the cluster's local disks, so storage scales with compute and vanishes when the cluster terminates; it cannot serve as a durable, decoupled data lake. It suits transient scratch space for a running Spark job, not persistent primary storage.

  • ✗

    Amazon RDS for MySQL

    Why it's wrong here

    RDS for MySQL is a relational OLTP database with row-oriented storage and limited capacity, unsuited to large CSV files processed by Spark. A relational database would be correct for transactional application data, not a data lake.

  • ✗

    Amazon EBS volumes attached to the EMR cluster

    Why it's wrong here

    EBS volumes are block storage tied to a single Availability Zone and cannot be mounted concurrently by multiple EMR nodes, so they cannot serve as a shared data lake. They suit single-instance scratch or database workloads, not multi-node Spark storage.

  • ✓

    Amazon S3

    Why this is correct

    Amazon S3 provides durable, virtually unlimited object storage with separate compute and storage scaling, so Spark reads CSV files directly and cost stays low. HDFS on Amazon EMR couples storage to cluster lifetime, raising cost and complicating elasticity.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.