Courseiva
Data EngineeringhardMultiple ChoiceObjective-mapped

MLS-C01 Data Engineering Practice Question

A research lab stores large genomic datasets in Amazon S3 Glacier Deep Archive. They need to run a one-time analysis on a subset of 10 PB of data. The analysis will use an Amazon EMR cluster with Amazon S3 as the data source. What is the MOST cost-effective and performant way to make the data available for the EMR cluster?

⚠ Common exam trap

A common mix-up: candidates assume Expedited retrieval is always the fastest and thus best for performance, ignoring the massive cost difference at petabyte scale and the fact that Bulk retrieval's 48-hour window is acceptable for a one-time analysis.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Initiate a Bulk retrieval request and restore the data to S3 Standard for the duration of the analysis

Bulk retrieval is the most cost-effective retrieval tier for large, non-urgent data from S3 Glacier Deep Archive, completing within 48 hours. Restoring to S3 Standard provides direct, high-throughput access for the EMR cluster, and deleting the data after analysis avoids ongoing storage costs. This approach balances performance (EMR reads from S3 Standard) with minimal cost (Bulk retrieval is the cheapest retrieval option).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Restore the data to S3 Standard-IA and delete after the analysis

    Why it's wrong here

    Standard-IA has a minimum storage charge and retrieval cost that may be higher than Standard for single access.

  • Configure the EMR cluster to read directly from Glacier Deep Archive using S3 Console

    Why it's wrong here

    EMR cannot read directly from Glacier Deep Archive; data must be restored first.

  • Initiate a Bulk retrieval request and restore the data to S3 Standard for the duration of the analysis

    Why this is correct

    Bulk retrieval is the lowest cost tier, and restoring to Standard avoids IA minimum charges.

  • Initiate an Expedited retrieval request and use the temporary copy for the EMR cluster

    Why it's wrong here

    Expedited is more expensive than Bulk and not necessary for a batch job.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.