easyMultiple Choice
MLA-C01 Practice Question: A data engineer is setting up a Glue ETL job to…
A data engineer is setting up a Glue ETL job to process a large dataset stored in Amazon S3. The job needs to read data in Parquet format, apply a filter, and write the results back to S3 in Parquet. The engineer wants to minimize the cost and runtime. Which optimization technique is MOST effective?
⚠ Common exam trap
MLA-C01 often tests the misconception that throwing more DPUs at a Glue job is the universal performance fix, when the real lever is reducing the data read via column pruning and predicate pushdown.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use column pruning to read only the columns needed in the transformation.
Column pruning is the most effective optimization because AWS Glue with the Spark engine can push column selection down to the Parquet reader, so only the required columns are read from S3 and processed. Parquet is a columnar format, meaning each column is stored separately, so skipping unneeded columns dramatically reduces I/O, memory footprint, and shuffle volume. This directly lowers DPU-hour consumption and runtime, which are the two cost drivers for Glue ETL jobs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of DPUs to the maximum allowed.
Why it's wrong here
Maximising DPUs scales cluster capacity, raising cost linearly while doing nothing to reduce the volume of data scanned. It is tempting because more compute speeds some jobs, but the filter and Parquet predicate pushdown already limit work; partition pruning and pushdown deliver the runtime and cost saving here.
- ✓
Use column pruning to read only the columns needed in the transformation.
Why this is correct
Parquet is columnar, so column pruning reads only the columns the filter and transformation reference, skipping irrelevant data on disk. This cuts bytes scanned and I/O, reducing both runtime and cost more effectively than row-based filtering or partition changes alone.
- ✗
Convert the data to JSON format for faster read performance.
Why it's wrong here
Converting Parquet to JSON discards columnar storage and predicate pushdown, forcing full scans of larger, text-based files and increasing both runtime and cost. It is tempting because JSON is familiar and flexible, but Glue reads Parquet natively; the optimisation lies in partition pruning and predicate pushdown on the existing columnar format.
- ✗
Use a single large file instead of partitioning.
Why it's wrong here
A single large file prevents partition pruning, so Glue must read the entire dataset even though the filter touches few partitions, increasing runtime and DPU cost. It is tempting because fewer files appear tidy, but partitioning by filter columns lets Glue skip irrelevant S3 prefixes, which is the effective optimisation.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.