Courseiva
hardMultiple Choice

PMLE Practice Question: Your model serving endpoint on Vertex AI is…

Your model serving endpoint on Vertex AI is experiencing increased memory usage after a recent update. The model was converted from TensorFlow to TF Lite for faster inference. You notice that the endpoint's instances occasionally get killed due to out-of-memory (OOM) errors. What is the most likely cause?

⚠ Common exam trap

The trap is assuming that increased traffic or model size is the cause, but the question highlights the recent conversion to TF Lite, pointing to a configuration issue like thread count, which is a known memory multiplier.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The number of inference threads in the TF Lite runtime is set too high, causing memory consumption.

The most likely cause is that the number of inference threads in the TF Lite runtime is set too high. TF Lite uses multiple threads for inference, and each thread consumes additional memory for its own stack and intermediate tensors. If the thread count is set too high relative to the available memory, it can lead to out-of-memory errors. This is a common configuration issue when moving to TF Lite.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The TF Lite model is larger in size than the original model.

    Why it's wrong here

    TF Lite conversion typically shrinks the model, so a larger file is not the cause; the OOM stems from runtime memory behaviour such as per-instance tensor allocation. It is tempting because conversion changes artefacts, but size reduction is the expected outcome, not growth.

  • ✗

    The Vertex AI endpoint is not configured with enough CPU.

    Why it's wrong here

    CPU shortage wouldn't cause OOM; it's memory-related.

  • ✓

    The number of inference threads in the TF Lite runtime is set too high, causing memory consumption.

    Why this is correct

    TF Lite's interpreter allocates per-thread memory arenas, so a high inference thread count multiplies working memory across concurrent threads. This satisfies the stem's OOM scenario: the converted model's runtime configuration, not the model size alone, drives the increased memory consumption killing instances.

  • ✗

    The traffic to the endpoint has increased significantly.

    Why it's wrong here

    Increased traffic raises request volume but does not itself increase per-instance memory footprint; OOM after a model-format change points to the conversion. It is tempting because load commonly causes resource exhaustion, but the timing ties the fault to the TF Lite update.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.