Courseiva
Data EngineeringhardMultiple SelectObjective-mapped

MLS-C01 Data Engineering Practice Question

A company uses Amazon S3 to store historical transaction data in CSV format. The data is partitioned by transaction_date. A data analyst runs Amazon Athena queries that frequently filter on customer_id and transaction_date. The queries are slow and expensive. The team needs to improve query performance and reduce cost. Which combination of actions should the team take? (Choose TWO.)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Convert the data from CSV to Parquet format.

Converting CSV to Parquet (option C) improves performance because Parquet is a columnar storage format that reduces the amount of data scanned by Athena, especially when queries only select a subset of columns. It also uses efficient compression, reducing storage and data scanned. Reorganizing the partition order (option D) to have customer_id first (the most frequently filtered column) improves partition pruning, reducing the amount of data read. Option A (S3 Select pushdown) is not fully supported by Athena; Athena already uses S3 Select for certain formats, but it doesn't guarantee significant improvement and may not be applicable. Option B (JSON) is worse than CSV because JSON is typically larger and not columnar. Option E (increasing workers) is not applicable as Athena is serverless and automatically scales.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Enable S3 Select pushdown in Athena to reduce data transfer.

    Why it's wrong here

    Athena does not use S3 Select pushdown for query optimization.

  • Convert the data to JSON format for better query performance.

    Why it's wrong here

    JSON is row-based and larger, increasing scan size.

  • Convert the data from CSV to Parquet format.

    Why this is correct

    Parquet is columnar and compressed, reducing data scanned.

  • Reorganize the data by partitioning on customer_id first, then transaction_date.

    Why this is correct

    Partition pruning on customer_id reduces the data scanned.

  • Increase the number of Athena query workers.

    Why it's wrong here

    Athena automatically scales; there is no worker configuration.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.