hardMultiple Choice
PMLE Practice Question: A company has a large-scale ML system that uses…
A company has a large-scale ML system that uses Vertex AI Pipelines to retrain models weekly. The pipeline includes a custom training job and a batch prediction step. After moving to production, they observe that batch prediction jobs often fail with 'Quota exceeded' errors. The project has sufficient CPU quota. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the misconception that all quota errors are related to CPU or memory, but the trap here is that accelerator types (GPUs/TPUs) have their own independent quota limits that are easily overlooked when CPU quota appears sufficient.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The batch prediction job is requesting a specific accelerator type that has a separate quota limit.
The most likely cause is that the batch prediction job is requesting a specific accelerator type (e.g., GPU or TPU) that has a separate quota limit from CPU quota. In Vertex AI, accelerator quotas are distinct from general compute (CPU) quotas, and even if the project has sufficient CPU quota, the accelerator quota may be exhausted, causing 'Quota exceeded' errors.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The pipeline is exceeding the maximum number of concurrent pipeline runs.
Why it's wrong here
Vertex AI Pipelines concurrency limits produce pipeline-level errors, not batch prediction 'Quota exceeded' failures, and weekly retraining rarely saturates concurrent runs. It is tempting because pipeline quotas exist, and would be correct if many pipelines ran simultaneously and failed at submission rather than during prediction.
- ✓
The batch prediction job is requesting a specific accelerator type that has a separate quota limit.
Why this is correct
Accelerator quotas are tracked separately from CPU quota in Google Cloud. A batch prediction job requesting GPUs or TPUs can therefore hit 'Quota exceeded' despite ample CPU quota, making the accelerator request the most likely cause.
- ✗
The batch prediction job is using a machine type that is not available in the region.
Why it's wrong here
An unavailable machine type in the region yields an invalid resource or location error, not 'Quota exceeded'. It is tempting because region and machine-type mismatches do cause job failures, and would be correct if the error named an unsupported machine type or region rather than a quota limit.
- ✗
The custom training job is consuming all available quota before the batch prediction job starts.
Why it's wrong here
Training and batch prediction consume separate Vertex AI resource quotas; a custom training job cannot exhaust the prediction quota that batch prediction requires. It is tempting because both steps share the project, and would apply if they drew on one pooled quota, but the failure is prediction-specific.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.