Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is configuring an AWS Glue ETL job to read data from an Amazon S3 bucket that contains nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer wants to optimize the job for performance and cost. Which two actions should the engineer take? (Choose two.)

⚠ Common exam trap

The trap here is assuming that a Glue crawler can convert file formats or that simply adding more DPUs is always the best way to optimize performance without considering cost.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable job bookmarks to track processed data and avoid reprocessing.

The Relationalize transform is purpose-built to flatten nested JSON, simplifying the ETL code and improving performance. Enabling job bookmarks ensures that only new or changed data is processed on subsequent runs, reducing cost and runtime. Together, these actions optimize the job for both performance and cost. The other options either do not address flattening, are not cost-effective, or misunderstand service capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of DPUs to the maximum allowed to ensure faster execution.

    Why it's wrong here

    Increasing DPUs can improve performance but also increases cost. The maximum DPU limit is 100 for Glue ETL jobs, but using the maximum without analyzing the workload can lead to unnecessary expenses. Performance tuning should be based on data size and complexity. Simply maxing out DPUs is not a cost-optimization action and may not be necessary if the job is already efficient.

  • ✓

    Enable job bookmarks to track processed data and avoid reprocessing.

    Why this is correct

    Job bookmarks in AWS Glue keep track of data that has already been processed in previous job runs. When reading from S3, bookmarks use the last modified time of objects to determine new data. Enabling bookmarks prevents reprocessing of unchanged data, reducing runtime and cost, especially for incremental loads. This is a best practice for ETL jobs that run repeatedly.

  • ✗

    Convert the JSON files to Parquet using an AWS Glue crawler before running the ETL job.

    Why it's wrong here

    An AWS Glue crawler catalogs data and infers schema; it does not convert file formats. Converting JSON to Parquet requires an ETL job or another tool like Amazon Athena CTAS. While Parquet is more efficient for querying, the crawler itself cannot perform the conversion. This option misrepresents the crawler's capabilities and does not flatten nested JSON.

  • ✓

    Use the AWS Glue DynamicFrame Relationalize transform to flatten nested JSON.

    Why this is correct

    The Relationalize transform in AWS Glue is specifically designed to flatten nested JSON structures into multiple relational tables. It automatically unnest arrays and structs, producing a set of DynamicFrames that can be joined or written separately. This reduces the need for custom code and improves performance by leveraging Glue's optimized execution engine.

  • ✗

    Use the AWS Glue ResolveChoice transform to handle data type conflicts.

    Why it's wrong here

    The ResolveChoice transform is used to handle ambiguous data types or conflicting schema in DynamicFrames, such as when a column has mixed types. While useful for schema resolution, it does not flatten nested JSON structures. The requirement is to flatten nested JSON, so this transform does not address the primary need and would not optimize for that purpose.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.