Courseiva
hardMultiple Choice

PDE Practice Question: Designing a system to serve predictions from a…

You are designing a system to serve predictions from a large language model (LLM) with a latency SLO of 500ms. The model does not fit on a single GPU and requires model parallelism. You are considering using Vertex AI Endpoints with a custom container. What additional setup is required to achieve the latency target?

⚠ Common exam trap

Candidates often incorrectly assume that distributing inference across multiple Vertex AI endpoints with a load balancer achieves model parallelism. However, this approach introduces network latency and cannot match the low-latency inter-GPU communication required for tensor parallelism within a single multi-GPU instance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a machine type with multiple GPUs and configure the container to use tensor parallelism.

The model does not fit on a single GPU and requires model parallelism. Using a machine type with multiple GPUs and configuring the container to use tensor parallelism allows the model to be split across GPUs within a single instance, enabling efficient parallel computation to meet the 500ms latency SLO. Tensor parallelism distributes individual tensor operations across GPUs, reducing communication overhead compared to pipeline parallelism and is a standard approach for large models on multi-GPU instances in Vertex AI.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Compile the model using TensorFlow XLA to optimize for single GPU execution.

    Why it's wrong here

    XLA compilation optimises graph execution on one GPU; it cannot place a model that exceeds single-GPU memory, so the deployment fails outright. It is tempting because XLA genuinely reduces latency for models that do fit. Model parallelism across multiple GPUs is what this workload requires.

  • ✗

    Deploy the model across multiple endpoints and use a load balancer to send requests to different parts of the model.

    Why it's wrong here

    Splitting one model across separate endpoints forces network hops between layers, adding latency that breaches the 500ms SLO. Load balancers distribute independent replicas, not shards of a single model. This suits horizontally scaled stateless services, not model-parallel inference needing intra-node GPU interconnect.

  • ✗

    Use Vertex AI Prediction as a service for LLMs, which automatically handles hardware selection.

    Why it's wrong here

    Vertex AI Prediction's managed LLM serving selects hardware automatically but does not shard a model too large for one GPU, so the memory constraint remains unmet. It is tempting because it removes infrastructure work. Custom containers with explicit multi-GPU model parallelism are needed here.

  • ✓

    Use a machine type with multiple GPUs and configure the container to use tensor parallelism.

    Why this is correct

    Tensor parallelism shards each model layer's weights across multiple GPUs, enabling the oversized LLM to fit and compute in parallel. A multi-GPU machine type with the container configured for tensor parallelism meets the 500ms latency SLO.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.