DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer needs to transform CSV files in S3 to Parquet format using a serverless solution. The files are large (up to 5 GB each) and arrive irregularly. Which TWO services can accomplish this with minimal operational overhead? (Choose TWO.)
⚠ Common exam trap
Watch out — candidates often confuse query engines (like Athena or Redshift Spectrum) with transformation services, or assume that any AWS service with 'serverless' in its name can perform ETL, when in fact Athena CTAS requires Step Functions orchestration and is limited to SQL-based transformations, not direct file format conversion.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue ETL job
AWS Glue ETL job is correct because it is a fully managed, serverless service that can automatically convert CSV to Parquet without provisioning infrastructure. It handles large files (up to 5 GB) by scaling Spark executors dynamically, and can be triggered by S3 events for irregular arrivals, minimizing operational overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
AWS Glue ETL job
Why this is correct
Glue is serverless and can convert large CSV to Parquet efficiently.
- ✓
AWS Step Functions with Athena CTAS queries
Why this is correct
Athena CTAS can convert to Parquet; Step Functions can orchestrate triggered by S3 events.
- ✗
Amazon EC2 with a script
Why it's wrong here
EC2 requires maintenance and is not serverless.
- ✗
Amazon EMR cluster
Why it's wrong here
EMR requires cluster provisioning and management, not serverless.
- ✗
Amazon Redshift Spectrum
Why it's wrong here
Redshift Spectrum is for querying external data, not converting.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,711 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.