hardMultiple Choice
PMLE Practice Question: A large e-commerce company uses Vertex AI…
A large e-commerce company uses Vertex AI Pipelines to orchestrate its recommendation model training. The pipeline has several parallel components: feature engineering, model training, and model evaluation. Recently, they noticed that the pipeline often fails due to resource exhaustion in the Vertex AI custom training job for the model training component. The training job consumes significant memory and occasionally exceeds the allocated memory limit, causing the pod to be OOMKilled. The team has already increased the memory to the maximum allowed for the chosen machine type. They need to prevent the pipeline from failing while still using the same machine type. Which approach should they take?
⚠ Common exam trap
Google Cloud often tests the misconception that retry policies or pre-checks can solve resource exhaustion, but the correct approach is to redesign the component to reduce peak memory usage, as retries do not fix the underlying OOM condition.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the training component into multiple smaller steps that process data in chunks to reduce peak memory usage.
Splitting the training component into smaller steps that process data in chunks directly addresses the root cause of OOMKilled failures—peak memory usage exceeding the allocated limit. By reducing the memory footprint per step, the pipeline can stay within the maximum memory of the existing machine type without requiring a larger instance. This approach aligns with best practices for Vertex AI custom training jobs, where resource limits are fixed per machine type and cannot be exceeded.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Split the training component into multiple smaller steps that process data in chunks to reduce peak memory usage.
Why this is correct
Chunked processing bounds peak memory by loading and training on data subsets sequentially, so the custom training job stays within the machine type's memory ceiling. This satisfies the stem's constraint of preventing OOMKilled pod failures while retaining the same machine type, rather than exceeding its maximum.
- ✗
Use a larger machine type with more memory to accommodate the peaks.
Why it's wrong here
The stem fixes the machine type, so scaling hardware contradicts the stated constraint. Larger machine types exist precisely to absorb memory peaks when the constraint is cost or availability rather than machine type; here that lever is explicitly removed, leaving only pipeline-level mitigation of the OOMKilled training component.
- ✗
Add a memory check step before training that estimates memory usage and skips training if it exceeds the limit.
Why it's wrong here
Skipping training when estimated memory exceeds the limit prevents the OOMKill but abandons model training entirely, so the pipeline completes without producing a model. It is tempting because pre-flight checks genuinely avoid wasted compute, which would suit scenarios where training is optional or can be deferred rather than required.
- ✗
Implement a retry policy with exponential backoff for the training component, so it automatically retries on failure.
Why it's wrong here
Retrying the component restarts the same job on the same machine type, so the OOMKilled pod recurs and the pipeline still fails. It is tempting because retry policies genuinely handle transient failures, which would suit scenarios involving intermittent infrastructure errors rather than deterministic memory exhaustion.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.