mediumMultiple Choice
PDE Practice Question: Your company uses Vertex AI Pipelines to automate…
Your company uses Vertex AI Pipelines to automate model retraining. The pipeline has three steps: data extraction from BigQuery, feature engineering using Dataflow, and model training using a custom container on Vertex AI Training. Recently, the pipeline has been failing intermittently at the Dataflow step with a 'The job encountered a transient error. Please retry.' message. You have enabled pipeline retries with 3 attempts. However, the pipeline still fails after 3 retries. You check the logs and find that the Dataflow job requires more resources than the default worker configuration provides. Which change should you make to reduce the failure rate?
⚠ Common exam trap
Google Cloud often tests the misconception that increasing parallelism (more workers) or retries will fix resource exhaustion errors, when the actual fix is to increase per-worker resources by selecting a larger machine type.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the Dataflow worker machine type to have more memory and CPU in the pipeline step configuration
The pipeline fails due to insufficient resources (memory and CPU) in the default Dataflow worker configuration. By increasing the worker machine type (e.g., using a custom machine type with more vCPUs and memory), the Dataflow job can handle the feature engineering workload without hitting resource limits, reducing transient failures. This directly addresses the root cause identified in the logs, unlike retries or parallelism changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of Dataflow workers to improve parallelism
Why it's wrong here
Adding workers raises parallelism, but the logged constraint is per-worker resource capacity, so each worker still exhausts memory or disk and the job fails. Scaling worker count is tempting because it speeds throughput, and would be correct when the job is slow due to insufficient parallelism rather than undersized workers.
- ✗
Increase the number of retries in the pipeline to 5
Why it's wrong here
Raising attempts to five repeats the same deterministic resource exhaustion; retries only help transient faults, and the logs show a capacity problem. More retries are tempting because the error message says 'transient', and would be correct when failures genuinely stem from intermittent infrastructure blips.
- ✗
Replace Dataflow with Dataproc to run the feature engineering step
Why it's wrong here
Migrating to Dataproc changes the execution engine but leaves the same undersized cluster configuration, so the resource shortfall persists. Dataproc is tempting for Spark or Hadoop workloads needing fine-grained cluster control, and would be correct when feature engineering requires those frameworks rather than Beam.
- ✓
Increase the Dataflow worker machine type to have more memory and CPU in the pipeline step configuration
Why this is correct
The Dataflow job fails because default workers lack memory and CPU, and retries cannot fix a resource shortfall. Increasing the worker machine type gives the step sufficient capacity, addressing the logged root cause and reducing the intermittent failures.
Visual reference
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.