Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON and writes to a partitioned Parquet table. The engineer wants to reduce job cost and improve read performance. (Choose two.)

⚠ Common exam trap

The trap here is assuming that adding workers or running a crawler reduces cost, when bookmarks and output partitioning are the levers that actually cut DPU time and scan volume.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable job bookmarks to avoid reprocessing previously processed files.

Job bookmarks cut cost by skipping already-processed S3 objects, which is important for incremental pipelines. Partitioning the output Parquet table by low-cardinality filter columns enables partition pruning, reducing I/O and improving read performance for downstream queries. Crawlers, worker scaling, and pre-conversion do not directly deliver both cost reduction and read performance for this job configuration.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable job bookmarks to avoid reprocessing previously processed files.

    Why this is correct

    Job bookmarks track which S3 objects have already been processed and skip them on subsequent runs. For incremental data landing in S3, this avoids re-reading and re-transforming old files, directly reducing DPU hours and cost. It is a supported Glue feature and a common cost-control measure for recurring jobs.

  • ✗

    Convert the source JSON to Parquet before the Glue job runs.

    Why it's wrong here

    Converting source JSON to Parquet ahead of the job adds an extra pipeline stage and its own compute cost. While Parquet is more efficient to read, the question is about configuring the Glue job, and pre-converting does not reduce the job's cost directly. It also complicates the architecture by introducing another transformation step.

  • ✗

    Increase the number of workers to the maximum allowed for the job.

    Why it's wrong here

    Adding workers increases parallelism but also increases the hourly DPU cost. It can shorten runtime, but the total cost may rise if the job is not compute-bound. The question asks for cost reduction and read performance, and simply scaling out does not guarantee either without addressing data layout or reprocessing.

  • ✗

    Use the Glue crawler to infer the schema and store it in the Data Catalog.

    Why it's wrong here

    A crawler populates the Data Catalog with table metadata, which helps query engines and job authoring. It does not reduce the cost of an ETL run or improve read performance of the job itself. The crawler runs separately and charges its own DPU time, so it is not a cost-reduction measure for the ETL job.

  • ✓

    Partition the output Parquet table by low-cardinality columns used in filters.

    Why this is correct

    Partitioning the output by columns frequently used in downstream filters enables partition pruning. Readers scan only relevant partitions, cutting I/O and improving query performance. This also reduces the volume of data processed in subsequent jobs, which lowers cost. The key is choosing low-cardinality columns to avoid creating too many small partitions.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.