MLS-C01 Data Engineering Practice Question
A company uses Amazon S3 to store historical transaction data in CSV format. The data is partitioned by transaction_date. A data analyst runs Amazon Athena queries that frequently filter on customer_id and transaction_date. The queries are slow and expensive. The team needs to improve query performance and reduce cost. Which combination of actions should the team take? (Choose TWO.)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the data from CSV to Parquet format.
Converting CSV to Parquet (option C) improves performance because Parquet is a columnar storage format that reduces the amount of data scanned by Athena, especially when queries only select a subset of columns. It also uses efficient compression, reducing storage and data scanned. Reorganizing the partition order (option D) to have customer_id first (the most frequently filtered column) improves partition pruning, reducing the amount of data read. Option A (S3 Select pushdown) is not fully supported by Athena; Athena already uses S3 Select for certain formats, but it doesn't guarantee significant improvement and may not be applicable. Option B (JSON) is worse than CSV because JSON is typically larger and not columnar. Option E (increasing workers) is not applicable as Athena is serverless and automatically scales.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable S3 Select pushdown in Athena to reduce data transfer.
Why it's wrong here
Athena does not use S3 Select pushdown for query optimization.
- ✗
Convert the data to JSON format for better query performance.
Why it's wrong here
JSON is row-based and larger, increasing scan size.
- ✓
Convert the data from CSV to Parquet format.
Why this is correct
Parquet is columnar and compressed, reducing data scanned.
- ✓
Reorganize the data by partitioning on customer_id first, then transaction_date.
Why this is correct
Partition pruning on customer_id reduces the data scanned.
- ✗
Increase the number of Athena query workers.
Why it's wrong here
Athena automatically scales; there is no worker configuration.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.