DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is designing an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structures and write the output to Amazon Redshift. The engineer needs to ensure the job can handle schema evolution and efficiently process only new data on subsequent runs. (Choose two.)
⚠ Common exam trap
It's easy for candidates to confuse transforms that adjust schema (ApplyMapping, ResolveChoice) with transforms that flatten nested data (Relationalize).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the AWS Glue Relationalize transform to flatten nested JSON structures.
Enabling job bookmarks ensures that only new data is processed on subsequent runs, which is essential for incremental ETL. The Relationalize transform flattens nested JSON into relational tables, making the data compatible with Amazon Redshift. Together, these two features address the requirements for incremental processing and flattening nested structures.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the AWS Glue Relationalize transform to flatten nested JSON structures.
Why this is correct
The Relationalize transform in AWS Glue is specifically designed to flatten nested JSON and semi-structured data into relational tables. It produces multiple tables that can be joined, which is necessary when writing to Amazon Redshift. This transform handles arrays and nested objects, making it the appropriate choice for converting nested JSON into a format suitable for a relational data warehouse.
- ✗
Configure the job to use the 'ApplyMapping' transform to rename nested fields.
Why it's wrong here
ApplyMapping is used to rename, cast, or drop fields in a DynamicFrame, but it does not flatten nested structures. It operates on existing fields and cannot expand nested objects or arrays into separate columns or tables. While useful for schema adjustments, it does not address the core requirement of flattening nested JSON for Redshift, so it is not one of the required choices.
- ✓
Enable bookmarking in the AWS Glue job to track previously processed data.
Why this is correct
AWS Glue job bookmarks track the state of data that has already been processed. When bookmarking is enabled, subsequent job runs process only new or changed data since the last run, which directly satisfies the requirement to efficiently process only new data. Bookmarks work with S3 sources by tracking object creation and modification times, reducing redundant processing.
- ✗
Use the 'ResolveChoice' transform to handle data type conflicts in nested fields.
Why it's wrong here
ResolveChoice handles data type conflicts in a DynamicFrame by specifying how to resolve ambiguous types, such as casting to a specific type or retaining both. It is useful for schema inconsistencies but does not flatten nested structures or enable incremental processing. It is not one of the two required choices for this scenario, which are bookmarking and flattening nested JSON.
- ✗
Set the job's maximum concurrency to 1 to prevent schema conflicts.
Why it's wrong here
Maximum concurrency controls how many instances of a job can run simultaneously. Setting it to 1 prevents parallel runs but does not help with schema evolution or processing only new data. Schema evolution is handled by the Glue Data Catalog and DynamicFrame schema resolution, while incremental processing is handled by job bookmarks. This setting does not satisfy either requirement.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.