Courseiva

PMLE Serving and Scaling Models Practice Question

A financial services firm serves a fraud-detection model on a Vertex AI endpoint that consumes features from a Vertex AI Feature Store online store. During a load test, prediction latency is acceptable, but the firm discovers that the model's feature values in production drift from the values used at training time because the training pipeline read from a BigQuery table with different transformation logic. The team wants the serving path to use the same feature definitions as training so online and offline values match. Which approach should they take?

⚠ Common exam trap

The trap here is treating a semantic mismatch between two transformation implementations as a latency or freshness problem, and reaching for storage scaling or caching instead of unifying the feature definition.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define the features once as Feature Store feature views with a shared transformation and serve the model from the online store while the training pipeline reads the same definitions from the offline store.

Training-serving skew from divergent transformation logic is solved by unifying feature definitions, not by tuning storage or adding caching. Vertex AI Feature Store lets a feature view define a transformation once, then materialize the result to the online store for low-latency serving and to the offline store for training. Because both paths derive from the same definition, the values the model sees at inference match the values it was trained on, eliminating the drift at its source.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Log all online feature values and retrain the model weekly on the logged production data.

    Why it's wrong here

    Retraining on logged online values would teach the model the skewed distribution rather than fix the discrepancy, and it leaves the two transformation paths in place. The underlying cause, two independent implementations of the same feature, would persist and could silently change again after any pipeline edit. This treats the symptom on a weekly cadence instead of eliminating the divergent logic.

  • ✗

    Increase the online store's node count and shorten the feature value TTL so fresher values are served.

    Why it's wrong here

    Scaling the online store improves throughput and freshness but does not reconcile differing transformation logic. The mismatch is semantic: the same raw input is being transformed two different ways. Faster delivery of a differently computed value would actually make the skew more visible, since production predictions would consistently disagree with the training distribution the model learned from.

  • ✗

    Route serving requests through the offline store and cache results in Memorystore for low latency.

    Why it's wrong here

    The offline store is optimized for large analytical reads, not per-request lookups, so using it on the serving path would introduce unacceptable latency even with a cache layer. It also does not guarantee identical transformation logic unless the definitions themselves are unified. Caching a differently computed value simply delivers the skew faster and adds an operational dependency that the scenario does not require.

  • ✓

    Define the features once as Feature Store feature views with a shared transformation and serve the model from the online store while the training pipeline reads the same definitions from the offline store.

    Why this is correct

    Feature Store provides a single definition for each feature, materialized to the online store for low-latency serving and to the offline store for training. When both paths derive from the same feature view and transformation, online and offline values are computed identically, which is the definition of training-serving skew prevention. This directly removes the divergent BigQuery logic that caused the drift.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.