PMLE Serving and Scaling Models Practice Question
A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?
⚠ Common exam trap
The trap here is conflating operational metrics with payload logging and assuming both must be turned off together to protect sensitive data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Disable request-response logging on the deployed model while leaving endpoint monitoring and Cloud Monitoring metrics enabled.
Payload logging in Vertex AI is controlled by an explicit setting on each deployed model and is off unless enabled. Disabling it prevents request and response content from reaching Cloud Logging, while the endpoint continues to publish aggregate metrics such as request counts, error rates, and latency to Cloud Monitoring, which is exactly the separation this compliance scenario requires.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Delete the endpoint's service account permissions on Cloud Logging so the platform cannot write any log entries, including metrics.
Why it's wrong here
Revoking logging permissions breaks operational logging broadly and does not selectively suppress payload capture. It can also cause deployment or health-check failures, and Cloud Monitoring metrics are delivered through a separate pipeline, so this approach harms observability without cleanly meeting the requirement.
- ✓
Disable request-response logging on the deployed model while leaving endpoint monitoring and Cloud Monitoring metrics enabled.
Why this is correct
Request-response logging is a separate, opt-in setting on the deployed model. Leaving it disabled means payloads are never captured, while the endpoint still emits standard metrics like request count, error rate, and latency percentiles to Cloud Monitoring, satisfying both the compliance constraint and the observability requirement.
- ✗
Set the endpoint's traffic split to route all requests to a canary deployment that has logging disabled, while the primary deployment keeps logging enabled.
Why it's wrong here
Traffic splitting governs which model version serves requests; it does not disable logging on the version that receives traffic. If the canary has logging off but the primary still logs, any request routed to the primary is captured, so the compliance guarantee is not met and the split adds needless complexity.
- ✗
Enable request-response logging but set the sampling rate to a very small value so almost no payloads are captured.
Why it's wrong here
Any nonzero sampling still writes some prediction payloads to Cloud Logging, which violates the requirement that no request or response content ever be logged. Compliance controls must be deterministic, not probabilistic, so a low sampling rate is unacceptable even if it reduces the volume of captured data.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.