Courseiva
mediumMultiple Choice

Reduce Vertex AI Endpoint Latency — CPU Preprocessing Bottleneck

You deployed a model on Vertex AI Endpoints using a custom container. The model serves predictions but the latency is higher than expected. You suspect the container is not making full use of the CPU resources. What should you do to reduce latency?

⚠ Common exam trap

Google Cloud often tests the misconception that scaling out (adding more nodes) or upgrading hardware (GPU) is the default fix for latency, when the real issue is often software-level concurrency configuration within the container.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Modify the container to use multi-threading or increase the number of workers in the prediction server (e.g., Gunicorn workers).

High latency in a CPU-based custom container often stems from underutilizing available CPU cores. By increasing the number of workers (e.g., Gunicorn workers) or enabling multi-threading, you allow the prediction server to handle multiple requests concurrently, reducing queue time and improving throughput. This directly addresses the symptom of the container not making full use of CPU resources.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Modify the container to use multi-threading or increase the number of workers in the prediction server (e.g., Gunicorn workers).

    Why this is correct

    A single-threaded prediction server serialises requests, leaving CPU cores idle and inflating latency. Enabling multi-threading or raising Gunicorn worker counts lets the container process concurrent inferences in parallel, directly addressing the underused CPU identified in the stem.

  • ✗

    Enable response caching on the endpoint.

    Why it's wrong here

    Response caching only returns previously computed answers for repeated identical requests; it cannot raise CPU utilisation for novel inputs, so latency on first-time predictions stays unchanged. It is tempting because caching genuinely cuts latency for high-repetition, read-heavy inference workloads, but this scenario requires tuning the container's CPU thread or worker settings instead.

  • ✗

    Change the machine type to a GPU-accelerated machine.

    Why it's wrong here

    A GPU accelerates parallel matrix computation, not the CPU-bound preprocessing or inference path the container is underusing. GPU machines are the right choice for large deep-learning models with heavy tensor operations; here the container's CPU utilisation is the stated bottleneck, so adding GPU capacity leaves it idle.

  • ✗

    Increase the number of nodes by adjusting autoscaling limits.

    Why it's wrong here

    Autoscaling adds replicas to serve more concurrent requests, but each replica still runs the same underutilised CPU configuration, so per-request latency is unchanged. Node scaling suits rising traffic volume; the stem identifies inefficient CPU use within the container, which more nodes do not address.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.