Courseiva
Data EngineeringhardMultiple SelectObjective-mapped

MLS-C01 Data Engineering Practice Question

A data engineer needs to set up a data lake on S3 that supports both batch and streaming ingestion. The data must be queryable by Athena, Redshift Spectrum, and EMR. Which TWO configurations are essential? (Choose two.)

⚠ Common exam trap

It's easy for candidates to confuse the ingestion mechanism (e.g., Kinesis Data Firehose) with the essential data lake configuration, or assume that S3 Select is required for queryability, when in fact the core requirements are a unified metadata catalog and an efficient storage format.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Store data in columnar formats like Parquet or ORC.

Columnar formats like Parquet and ORC are optimized for analytical queries, reducing I/O by reading only the necessary columns. This is essential for Athena, Redshift Spectrum, and EMR, which all benefit from the efficient compression and predicate pushdown capabilities of these formats, enabling faster query performance and lower costs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Store data in columnar formats like Parquet or ORC.

    Why this is correct

    Columnar formats improve query performance and reduce scan costs for Athena and Redshift Spectrum.

  • Use the AWS Glue Data Catalog as a central metadata repository.

    Why this is correct

    Athena, Redshift Spectrum, and EMR all integrate with the Glue Data Catalog.

  • Enable S3 Select on the target buckets.

    Why it's wrong here

    S3 Select is a feature for filtering data, not a configuration required for querying.

  • Enable S3 versioning on all buckets.

    Why it's wrong here

    Versioning is for data protection, not necessary for querying.

  • Set up Kinesis Data Firehose for streaming ingestion.

    Why it's wrong here

    Streaming ingestion is optional; the question asks for essential configurations.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.