DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is designing an AWS Glue ETL job that reads data from an Amazon S3 bucket containing many small JSON files and writes the output to Amazon S3 in Parquet format. The engineer wants to improve read performance and reduce the number of output files. Which two actions should the engineer take? (Choose two.)
⚠ Common exam trap
The trap here is thinking that adding more compute or enabling bookmarks will solve the small-file problem, when the real fixes are read-side file grouping and write-side partition control.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call repartition or coalesce on the DynamicFrame before writing to control the number of output files.
Grouping small files during the read reduces file-open overhead and improves read performance, while repartitioning or coalescing before writing controls the number of output files. Job bookmarks only help across runs, adding DPUs does not fix per-file overhead or output count, and pre-converting files with Lambda does not address the small-file or output-file problems within the Glue job.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of DPUs to allow more executors to read the small files in parallel.
Why it's wrong here
Adding DPUs increases parallel capacity, but reading many small files is often limited by per-file overhead and metadata operations rather than raw compute. More executors may not resolve the inefficiency of opening thousands of small files. This action also does not reduce the number of output files written to S3, so it does not meet either goal.
- ✓
Call repartition or coalesce on the DynamicFrame before writing to control the number of output files.
Why this is correct
Repartitioning or coalescing before the write step controls how many partitions exist and therefore how many output files are produced. Coalesce reduces the number of partitions without a full shuffle, which is efficient when reducing file count. This action directly reduces the number of output files written to S3, complementing the read-side grouping.
- ✗
Enable job bookmarks to skip previously processed files and reduce the number of files read.
Why it's wrong here
Job bookmarks track previously processed data to avoid reprocessing on subsequent runs, which reduces input volume over time but does not address the small-file problem within a single run. They do not group files or change output file count. While bookmarks are useful for incremental processing, they do not improve read performance for the many small files in the current run or reduce the number of output files.
- ✗
Convert the source JSON files to Parquet before the Glue job runs using an AWS Lambda function triggered on each S3 object.
Why it's wrong here
Converting source files to Parquet ahead of the job changes the input format but does not address the number of small files or the output file count within the Glue job. A Lambda function per object would also add complexity and may not scale well for many objects. This approach does not use Glue's built-in small-file handling and does not control output partitioning.
- ✓
Use the AWS Glue groupFiles option to group multiple small files into a single partition during the read.
Why this is correct
The groupFiles option in Glue ETL allows the reader to coalesce multiple small files into a single partition, reducing the overhead of opening many small files. This improves read performance by lowering the number of tasks and file-open operations. It directly addresses the many-small-files problem and is a documented Glue feature for S3 sources.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.