DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer must transform data in Amazon S3 using Apache Spark. The transformation logic needs to be reused across multiple AWS Glue jobs, and the engineer wants to version-control the code and run it in a serverless environment without managing clusters. Which approach should the engineer take?
⚠ Common exam trap
The trap here is assuming that AWS Glue cannot import external code and that all transformation logic must reside in a single job script.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Package the transformation logic as a Python wheel file stored in Amazon S3, reference it in AWS Glue job parameters, and import it as a module in the job script.
AWS Glue supports referencing external Python libraries through the --extra-py-files job parameter. Packaging transformation logic as a wheel file in Amazon S3 allows the code to be versioned, reused across multiple Glue jobs, and executed in a serverless Spark environment. This satisfies both the reusability and serverless requirements without duplicating code or introducing cluster management overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create an AWS Lambda function containing the transformation logic and invoke it from each AWS Glue job using boto3.
Why it's wrong here
AWS Lambda has a 15-minute maximum execution timeout and limited memory and disk space, making it unsuitable for large-scale Spark transformations. Invoking Lambda from Glue adds latency and complexity, and Lambda cannot directly participate in a distributed Spark execution plan. This approach does not satisfy the need for reusable Spark transformation logic running serverlessly within Glue.
- ✓
Package the transformation logic as a Python wheel file stored in Amazon S3, reference it in AWS Glue job parameters, and import it as a module in the job script.
Why this is correct
AWS Glue supports referencing additional Python modules via the --extra-py-files job parameter, which can point to a wheel file in Amazon S3. This enables code reuse across jobs, allows version control of the wheel artifact, and keeps the execution serverless. The engineer writes the transformation once, packages it, and imports it in any job that needs it.
- ✗
Provision an Amazon EMR cluster, store the transformation code in a Git repository, and run Spark jobs on the cluster.
Why it's wrong here
Amazon EMR requires the engineer to manage cluster provisioning, scaling, and termination, which contradicts the requirement for a serverless environment without cluster management. While EMR supports version-controlled code, it is not the serverless option described. The scenario specifically asks for a serverless approach that reuses code, and EMR introduces operational overhead that AWS Glue avoids.
- ✗
Create an AWS Glue job with a script that contains all transformation logic, and copy-paste the code into each job that needs it.
Why it's wrong here
Copy-pasting transformation logic across multiple AWS Glue jobs creates maintenance overhead and risks divergence when logic changes. It does not provide true code reuse or centralized version control, and it duplicates effort. While AWS Glue jobs can run Spark code serverlessly, this approach fails the requirement for reusable, version-controlled logic because each job maintains its own isolated copy.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.