MLA-C01 Data Preparation for Machine Learning Practice Question
An organization stores raw data in Amazon S3 as CSV files. They need to perform serverless data transformation and convert the data to Parquet format for efficient ML training. Which AWS service is most appropriate?
⚠ Common exam trap
Test-takers frequently confuse Amazon Athena's ability to query Parquet data with the ability to transform data into Parquet, but Athena is a query engine, not an ETL transformation service.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue
AWS Glue is the most appropriate service because it is a fully managed, serverless ETL service designed specifically for data transformation tasks like converting CSV to Parquet. It automatically handles schema inference, data partitioning, and optimization for ML training workloads without requiring infrastructure management.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
AWS Glue
Why this is correct
AWS Glue provides serverless Spark-based ETL that reads CSV from Amazon S3 and writes Parquet, satisfying the stem's transformation and format-conversion requirement. Its crawler and job model needs no cluster management, and Parquet's columnar layout accelerates downstream ML training.
- ✗
Amazon EMR
Why it's wrong here
Amazon EMR provisions clusters of EC2 instances, so it cannot deliver the serverless transformation the scenario requires. It is tempting because EMR genuinely handles large-scale Apache Spark and Hadoop processing, and would be the right choice for petabyte-scale batch jobs needing fine-grained cluster control, custom libraries, or persistent compute across many interdependent steps.
- ✗
Amazon Athena
Why it's wrong here
Athena queries S3 data and can create Parquet tables via CTAS, but it is a query engine, not a transformation pipeline with scheduling and dependencies. It is tempting because it is serverless and reads CSV directly, and would be correct for ad-hoc SQL analysis rather than repeatable ETL.
- ✗
Amazon Redshift
Why it's wrong here
Redshift is a provisioned data warehouse requiring cluster management, so it does not deliver the serverless transformation the scenario specifies. It is tempting because it can run SQL over S3 data via Spectrum and write Parquet, and would be correct for analytical warehousing rather than ML feature preparation.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.