mediumMultiple Choice
PMLE Practice Question: An ML engineer notices that predictions are…
Exhibit
Refer to the exhibit.
gcloud ai endpoints describe projects/my-project/locations/us-central1/endpoints/456
...
deployedModels:
- id: 'bert-model-1'
model: projects/my-project/locations/us-central1/models/bert
displayName: bert
automaticResources:
minReplicaCount: 1
maxReplicaCount: 10
machineType: n1-standard-4
accelerator:
count: 0
enableAccessLogging: true
...
disableContainerLogging: true
...An ML engineer notices that predictions are taking longer than expected under moderate traffic. Reviewing the endpoint configuration, what is the most likely cause of the high latency?
⚠ Common exam trap
Google Cloud often tests the misconception that CPU machine type is the primary cause of latency, when in fact the accelerator count being zero is the more direct and common misconfiguration for deep learning models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The accelerator count is 0, meaning no GPU is used.
When the accelerator count is set to 0, the endpoint runs inference on the CPU only, which is significantly slower than GPU-accelerated inference for deep learning models. This is the most direct cause of high latency under moderate traffic, as the model's compute demands exceed CPU throughput.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Container logging is disabled, slowing down request processing.
Why it's wrong here
Disabled container logging removes diagnostic output; it does not add request-processing delay, since logging is asynchronous and off the inference path. This option tempts because logging feels tied to performance, but it would be the answer only if the symptom were missing telemetry rather than slow predictions.
- ✓
The accelerator count is 0, meaning no GPU is used.
Why this is correct
An accelerator count of 0 means the deployment runs on CPU only, so inference executes without GPU parallelism and latency rises under moderate load. This directly satisfies the stem's observed slowdown, since GPU-backed endpoints process matrix operations far faster.
- ✗
The machine type n1-standard-4 is underpowered for the model's compute needs.
Why it's wrong here
Underpowered compute would show sustained CPU or memory saturation, but moderate traffic alone rarely saturates an n1-standard-4; the stem points to configuration, not capacity. This option is tempting because undersized machines do raise latency, yet it misreads the scenario's moderate-load clue.
- ✗
Automatic scaling is set with a maxReplicaCount of 10, which creates overhead.
Why it's wrong here
A higher maxReplicaCount does not itself add per-request latency; it only caps how far the deployment may scale out under load. The overhead stems from insufficient provisioned replicas or a low request timeout, not the ceiling. Raising maxReplicaCount is the right move when sustained traffic exceeds the current replica limit.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.