Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A company runs a nightly batch ETL job using AWS Glue to transform data from Amazon RDS for MySQL to Amazon S3. The job reads 100 tables and writes Parquet files partitioned by date. Recently, the job started failing with 'ThrottlingException' from the RDS database. The data volume has increased, and the Glue job is reading large tables without any filtering. The job uses a single Glue job with multiple Spark executors. The engineer needs to reduce the load on the RDS database while maintaining the same processing time. What should the engineer do?

⚠ Common exam trap

The trap is assuming 'more DPUs = faster and less load'; in JDBC-based Glue jobs, more parallelism increases concurrent connections and worsens source-side throttling, so the correct answer is to reduce rows read, not add compute.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Modify the Glue job to use a JDBC connection with a WHERE clause to read only the latest partition.

The ThrottlingException originates from RDS because the Glue job issues full-table JDBC reads across 100 tables with no predicate pushdown. Adding a WHERE clause to read only the latest partition (D) reduces the number of rows scanned and transferred per run, directly lowering the query load on RDS while keeping the same nightly processing window. This is the least disruptive change that preserves the existing architecture and processing time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AWS DMS to continuously replicate data to S3 and then run Glue on the S3 data.

    Why it's wrong here

    Continuous DMS replication adds ongoing cost and change-data-capture complexity for a nightly batch job, and full-load replication still reads every row from RDS. DMS fits continuous migration or near-real-time replication, not a scheduled ETL that only needs filtered incremental extracts.

  • ✗

    Change the job to read all data from RDS into a staging table in S3 first, then transform.

    Why it's wrong here

    Staging all 100 tables into S3 first still requires the same unfiltered reads from RDS, so throttling persists and an extra copy lengthens the pipeline. Staging suits decoupling transformations from source availability, not reducing concurrent read pressure on a throttled database.

  • ✗

    Increase the number of DPUs to process data faster.

    Why it's wrong here

    More DPUs increase parallel Spark connections to RDS, worsening ThrottlingException rather than relieving it, and processing time is unchanged because the bottleneck is the source database. Scaling DPUs suits compute-bound transformations, not source-side read contention.

  • ✓

    Modify the Glue job to use a JDBC connection with a WHERE clause to read only the latest partition.

    Why this is correct

    Reading all rows from every table drives concurrent JDBC connections that overwhelm RDS, triggering ThrottlingException. Adding a WHERE clause restricts each read to the latest partition, sharply reducing rows scanned and database load while preserving the same overall processing time.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.