Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is designing an AWS Glue ETL job that reads data from an Amazon S3 bucket containing many small JSON files and writes the output to Amazon S3 in Parquet format. The engineer wants to improve read performance and reduce the number of output files. Which two actions should the engineer take? (Choose two.)

⚠ Common exam trap

The trap here is thinking that adding more compute or enabling bookmarks will solve the small-file problem, when the real fixes are read-side file grouping and write-side partition control.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call repartition or coalesce on the DynamicFrame before writing to control the number of output files.

Grouping small files during the read reduces file-open overhead and improves read performance, while repartitioning or coalescing before writing controls the number of output files. Job bookmarks only help across runs, adding DPUs does not fix per-file overhead or output count, and pre-converting files with Lambda does not address the small-file or output-file problems within the Glue job.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of DPUs to allow more executors to read the small files in parallel.

    Why it's wrong here

    Adding DPUs increases parallel capacity, but reading many small files is often limited by per-file overhead and metadata operations rather than raw compute. More executors may not resolve the inefficiency of opening thousands of small files. This action also does not reduce the number of output files written to S3, so it does not meet either goal.

  • ✓

    Call repartition or coalesce on the DynamicFrame before writing to control the number of output files.

    Why this is correct

    Repartitioning or coalescing before the write step controls how many partitions exist and therefore how many output files are produced. Coalesce reduces the number of partitions without a full shuffle, which is efficient when reducing file count. This action directly reduces the number of output files written to S3, complementing the read-side grouping.

  • ✗

    Enable job bookmarks to skip previously processed files and reduce the number of files read.

    Why it's wrong here

    Job bookmarks track previously processed data to avoid reprocessing on subsequent runs, which reduces input volume over time but does not address the small-file problem within a single run. They do not group files or change output file count. While bookmarks are useful for incremental processing, they do not improve read performance for the many small files in the current run or reduce the number of output files.

  • ✗

    Convert the source JSON files to Parquet before the Glue job runs using an AWS Lambda function triggered on each S3 object.

    Why it's wrong here

    Converting source files to Parquet ahead of the job changes the input format but does not address the number of small files or the output file count within the Glue job. A Lambda function per object would also add complexity and may not scale well for many objects. This approach does not use Glue's built-in small-file handling and does not control output partitioning.

  • ✓

    Use the AWS Glue groupFiles option to group multiple small files into a single partition during the read.

    Why this is correct

    The groupFiles option in Glue ETL allows the reader to coalesce multiple small files into a single partition, reducing the overhead of opening many small files. This improves read performance by lowering the number of tasks and file-open operations. It directly addresses the many-small-files problem and is a documented Glue feature for S3 sources.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.