Courseiva
hardMultiple Choice

PMLE Practice Question: A company uses Vertex AI Predictions with a…

A company uses Vertex AI Predictions with a custom container that invokes an external API for feature enrichment. The prediction response time is highly variable. The engineer wants to monitor the external API's contribution to latency. What should the engineer do?

⚠ Common exam trap

Google Cloud often tests the distinction between monitoring (custom metrics) and tracing (Cloud Trace) — the trap here is that candidates assume Cloud Trace automatically captures all downstream calls, but it requires explicit instrumentation of the external API call to record its duration, whereas custom metrics can be emitted directly from the container code without needing distributed tracing context.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Instrument the prediction container to emit custom metrics for the time spent in each prediction step, including the external API call.

Instrumenting the custom container to emit custom metrics (e.g., using OpenTelemetry or a Prometheus client library) allows the engineer to directly measure the time spent in each prediction step, isolating the external API call's contribution to latency. This provides granular, real-time visibility into the specific bottleneck, which is essential when the response time is highly variable and the external API is a known dependency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Instrument the prediction container to emit custom metrics for the time spent in each prediction step, including the external API call.

    Why this is correct

    Instrumenting the container to emit custom metrics isolates the external API call's duration from total prediction latency, satisfying the requirement to monitor that specific contribution. Vertex AI's built-in metrics only capture overall request latency, so per-step timing must be emitted by the container itself.

  • ✗

    Add a timeout setting to the endpoint's request to limit the external API call duration.

    Why it's wrong here

    A timeout bounds the external call but does not measure its latency contribution; the engineer needs per-request latency logging or custom metrics from the container. It tempts because timeouts protect the endpoint from slow dependencies, and they are correct when the goal is preventing hung external calls from exhausting prediction capacity.

  • ✗

    Monitor the Vertex AI endpoint latency metric and correlate with system metrics like CPU and memory.

    Why it's wrong here

    Endpoint latency aggregates the whole request path, so it cannot isolate the external API's share; CPU and memory metrics describe the container, not the outbound call. This approach suits capacity planning for the model server itself, where resource saturation drives variability, rather than attributing latency to a downstream dependency.

  • ✗

    Use Cloud Trace to trace the prediction request end-to-end, including the external API call.

    Why it's wrong here

    Cloud Trace can trace the prediction request end-to-end, but it cannot isolate the external API’s contribution to latency unless the custom container explicitly creates a child span for that call. Without manual instrumentation of the external API invocation, Cloud Trace merges the API latency into the container’s own span, obscuring the specific delay. It is tempting because Cloud Trace is designed for distributed tracing across services, and would be correct if the external API were instrumented to propagate trace context.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.