Courseiva

DEA-C01 Data Operations and Support Practice Question

A company uses AWS Glue to run ETL jobs that process data from an Amazon RDS for MySQL database and load it into an Amazon S3 data lake. The Glue job runs daily and processes incremental data. Recently, the job has been taking longer than expected. The engineer checks the CloudWatch logs and sees that the job is spending most of its time on the 'Reading from JDBC' phase. The MySQL table has 10 million rows and is indexed on the primary key. The Glue job uses a 'job bookmark' to track processed data. The engineer wants to improve the performance of the read phase. Which action is most likely to help?

⚠ Common exam trap

The trap is thinking that adding more DPUs or increasing fetchSize will solve the problem; the real issue is reading too much data, so the fix is to reduce the data volume at the source with a filtered query.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Modify the job to use a 'query' parameter that selects only the new or modified rows based on a timestamp column.

Using a 'query' parameter with a timestamp filter allows the Glue job to push down the filtering to the JDBC source, so only new or modified rows are read from MySQL. This drastically reduces the volume of data transferred and the time spent in the 'Reading from JDBC' phase, especially when combined with an index on the timestamp column.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the JDBC 'fetchSize' parameter to 10000.

    Why it's wrong here

    Raising fetchSize to 10000 makes the MySQL driver buffer the entire result set in executor memory, causing spills or out-of-memory errors rather than faster reads. It tempts because larger fetch batches do reduce round trips, and a moderate value such as 1000 would be the right tuning for this scenario.

  • ✗

    Disable job bookmark and perform a full refresh each time.

    Why it's wrong here

    Disabling the job bookmark forces every run to re-read all 10 million rows, increasing the JDBC phase rather than shortening it. It tempts because bookmarks can misbehave after schema changes, and a full refresh would be correct when bookmark state is corrupt or data must be reprocessed.

  • ✗

    Increase the number of DPUs for the Glue job.

    Why it's wrong here

    Adding DPUs scales Spark executor parallelism, but the bottleneck is the single JDBC connection pulling rows sequentially from RDS, so extra workers sit idle. It tempts because DPU increases usually fix slow Glue jobs, and they would be correct if the delay were in transformation or write stages.

  • ✓

    Modify the job to use a 'query' parameter that selects only the new or modified rows based on a timestamp column.

    Why this is correct

    A timestamp-based query parameter pushes filtering to MySQL, so only new or modified rows cross JDBC rather than the full 10 million rows. This directly reduces time spent in the Reading from JDBC phase, satisfying the stem's constraint of incremental daily loads tracked by a job bookmark.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.