Courseiva
mediumMultiple Select

PMLE Practice Question: A model serving team is experiencing high latency…

A model serving team is experiencing high latency in production. Which TWO actions should they take to diagnose the root cause? (Choose 2.)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable Cloud Trace to analyze request latency across services.

The correct actions to diagnose root cause of high latency are enabling Cloud Trace (B) and checking autoscaling metrics and cold start frequency (C). Cloud Trace provides detailed latency breakdown across services, helping identify bottlenecks. Checking autoscaling and cold start metrics reveals if latency is due to scaling delays or initialization overhead. Option A (converting framework) is not a diagnostic step and may introduce risk. Option D (increasing replicas) may temporarily reduce load but does not diagnose the cause. Option E (DEBUG logging) adds overhead without providing latency analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the model to a different framework that is faster.

    Why it's wrong here

    Framework conversion alters the model itself rather than measuring where latency originates, so it cannot diagnose the root cause. It is tempting because a faster runtime can reduce inference time, but that is a remediation step taken after profiling identifies the bottleneck.

  • ✓

    Enable Cloud Trace to analyze request latency across services.

    Why this is correct

    Cloud Trace captures distributed traces across your serving path, breaking end-to-end request latency into per-service and per-span timings. This directly satisfies the stem's need to diagnose where latency originates, revealing whether delay sits in preprocessing, model inference, or downstream calls rather than merely confirming that latency exists.

  • ✓

    Check the endpoint's autoscaling metrics and cold start frequency.

    Why this is correct

    Autoscaling metrics and cold start frequency expose whether latency stems from insufficient replicas, scale-up lag or cold-start overhead on the serving endpoint. These are direct infrastructure causes of high latency, making this a targeted diagnostic action.

  • ✗

    Increase the number of replicas to reduce load per replica.

    Why it's wrong here

    Adding replicas scales capacity to absorb concurrent requests; it does not reveal why individual requests are slow, so it masks rather than diagnoses the root cause. It is tempting because horizontal scaling is the standard remedy for load-induced latency, and would be correct if profiling showed replicas saturated at their throughput ceiling.

  • ✗

    Set the logging verbosity to DEBUG in the container.

    Why it's wrong here

    DEBUG verbosity floods logs with detail unrelated to latency, obscuring the timing signals needed to locate the bottleneck. It is tempting because more logging feels diagnostic, yet latency diagnosis requires request tracing and metrics, not container log level changes.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.