hardMultiple Choice
PMLE Practice Question: A travel booking company has a real-time…
A travel booking company has a real-time recommendation system that suggests hotels and flights to users. The model is served using TensorFlow Serving on a Google Kubernetes Engine (GKE) cluster with auto-scaling enabled. The cluster uses n1-standard-4 machine types. The team has set up Cloud Monitoring dashboards and alerts. Last week, during a major holiday promotion, the team noticed that the model's inference latency P99 increased from 150 ms to 450 ms over a 30-minute period, while the request throughput increased from 500 to 1,200 requests per second. CPU utilization across the cluster rose to 95%, but memory utilization remained at 60%. The model version and the serving infrastructure configuration have not changed since the last deployment. Which action should the team take to mitigate the latency issue?
⚠ Common exam trap
Google Cloud often tests the misconception that reducing per-pod CPU requests (Option C) is a valid scaling strategy, but in reality this increases overcommitment and can worsen latency under high load, whereas adding nodes (Option D) provides dedicated resources without contention.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add more nodes to the GKE cluster to increase the total CPU resources available for serving.
The latency spike is caused by CPU saturation (95% utilization) under increased load (500 to 1,200 RPS). Adding more nodes to the GKE cluster directly increases the total CPU resources available, allowing the existing TensorFlow Serving pods to handle the higher throughput without contention. This is the most immediate and infrastructure-appropriate fix because the model version and serving configuration have not changed, ruling out model-level or code-level optimizations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Implement a feature engineering pipeline that compresses the input features to reduce data size and inference time.
Why it's wrong here
Feature compression reduces payload size and pre-processing time, but the stem shows CPU saturation at 95% during a throughput surge with an unchanged model, so inference compute, not input volume, drives the latency. It is tempting because feature engineering genuinely cuts inference cost when inputs are large, and would be correct for a data-transfer-bound pipeline, not a CPU-bound serving tier.
- ✗
Deploy a newer version of the model that uses a more efficient architecture to reduce computational complexity.
Why it's wrong here
CPU saturation at 95% with unchanged memory points to insufficient compute capacity, so swapping in a lighter model does not address the immediate scaling shortfall. It tempts because reducing computational complexity lowers per-request cost, which would be right if the model itself were the bottleneck rather than replica count.
- ✗
Increase the number of TensorFlow Serving instances by reducing the CPU request per pod in GKE to allow more pods per node.
Why it's wrong here
Lowering the CPU request does not add physical cores; n1-standard-4 nodes are already at 95% CPU, so packing more pods per node increases contention and worsens P99 latency. It is tempting because reducing requests raises pod density and triggers scale-out, which suits memory-bound or over-provisioned workloads, but here the bottleneck is genuine CPU saturation.
- ✓
Add more nodes to the GKE cluster to increase the total CPU resources available for serving.
Why this is correct
Adding nodes directly addresses the CPU saturation driving the latency spike: throughput tripled while CPU hit 95%, yet memory stayed at 60%, so the bottleneck is compute, not capacity per pod. Horizontal node scaling gives TensorFlow Serving more cores to parallelise inference across, restoring P99 latency without altering the model or serving configuration.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.