Courseiva
mediumMultiple Choice

PDE Practice Question: A data engineering team uses Cloud Data Fusion to…

A data engineering team uses Cloud Data Fusion to build ETL pipelines. They have a pipeline that reads from Cloud SQL, transforms data using Wrangler, and writes to BigQuery. The pipeline fails intermittently with a 'connection timeout' error from Cloud SQL. What is the best way to handle this?

⚠ Common exam trap

It's easy for candidates to assume connectivity issues require network-level fixes (like static IPs or NAT) or scaling, rather than recognizing that transient timeouts are best handled by application-level retry and timeout configuration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the Cloud SQL connector in Data Fusion to use retry logic and increase the connection timeout.

Cloud Data Fusion's Cloud SQL connector can be configured with retry logic and an increased connection timeout to handle transient network issues. This directly addresses the intermittent 'connection timeout' error without requiring architectural changes, as the error is likely due to brief network latency or resource contention, not a persistent connectivity problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Cloud NAT to provide a static IP for Data Fusion to whitelist.

    Why it's wrong here

    Cloud NAT gives outbound egress a static IP, but the timeout stems from Cloud SQL connection limits or network path, not IP allowlisting; Data Fusion's private IP already reaches the instance. It is tempting because whitelisting is a common fix, yet Cloud NAT is correct only when an external service requires a fixed egress address.

  • ✓

    Configure the Cloud SQL connector in Data Fusion to use retry logic and increase the connection timeout.

    Why this is correct

    Intermittent connection timeouts are transient network or database availability faults, not pipeline logic errors. Configuring the Cloud SQL connector with retry logic and a longer connection timeout lets Data Fusion reattempt failed connections automatically, resolving the intermittent failure without redesigning the pipeline.

  • ✗

    Increase the number of Data Fusion nodes to distribute the load.

    Why it's wrong here

    Adding nodes scales pipeline parallelism, not the Cloud SQL connection pool; more concurrent workers worsen timeout errors by exhausting available connections. It is tempting because throughput problems often respond to more workers, but node scaling is correct only when the bottleneck is Data Fusion compute, not the source database.

  • ✗

    Migrate Cloud SQL to Cloud Spanner to handle higher concurrency.

    Why it's wrong here

    Spanner is a globally distributed database requiring schema and application redesign; it does not address intermittent connection timeouts, which stem from Cloud SQL connectivity or connection limits. It is tempting because Spanner handles high concurrency, but migration is correct only when the workload genuinely outgrows Cloud SQL's scaling ceiling.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.