Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A data science team is building a real-time feature engineering pipeline for ML model training and serving. They need to compute features from streaming data, store them for low-latency serving, and ensure consistency between training and serving. Which TWO Google Cloud services should they use?

⚠ Common exam trap

A common trap in Google PMLE exams is assuming BigQuery can serve as a low-latency online feature store for real-time inference, but it is designed for analytical queries with seconds-to-minutes latency, not sub-millisecond serving required for real-time ML inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Vertex AI Feature Store

Vertex AI Feature Store (A) is correct because it provides a centralized repository for storing, serving, and sharing feature data with low-latency online serving and batch serving for training, ensuring consistency between training and serving through point-in-time lookups and feature value time-stamping. Cloud Dataflow (D) is correct because it is a fully managed stream and batch processing service based on Apache Beam, enabling real-time feature engineering from streaming data with exactly-once processing semantics and automatic scaling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Vertex AI Feature Store

    Why this is correct

    Vertex AI Feature Store provides a centralised repository for feature values, enabling low-latency online serving alongside consistent offline retrieval for training. It directly satisfies the stem's requirement for training-serving consistency and low-latency access, ingesting features computed from streaming data without duplicating logic between pipelines.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery is an analytics warehouse optimised for batch SQL over large datasets, so it cannot compute features from streaming data or serve them at the low latency the pipeline requires. It is tempting because it stores training data and supports ML, but the correct pair is Dataflow plus Vertex AI Feature Store.

  • ✗

    Cloud Functions

    Why it's wrong here

    Cloud Functions runs event-driven code snippets, not stateful streaming feature computation or low-latency feature storage, so it cannot maintain training/serving consistency. It is tempting because it handles lightweight event triggers, and would suit glue logic such as invoking a pipeline when new data lands.

  • ✓

    Cloud Dataflow

    Why this is correct

    Cloud Dataflow provides the streaming compute engine, running Apache Beam pipelines that transform real-time events into features. It satisfies the constraint of computing features from streaming data, and the same pipeline code can be reused for batch training, supporting training-serving consistency.

  • ✗

    Cloud SQL

    Why it's wrong here

    Cloud SQL is a managed relational database for transactional workloads; it lacks the streaming ingestion, feature-store semantics and training-serving consistency the pipeline demands. It is tempting because it stores data with low-latency reads, but the correct choices are Dataflow for stream computation and Vertex AI Feature Store for serving.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

3 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A data science team needs to share features across multiple ML models while ensuring consistency between training and serving. Which approach best achieves this?

medium
  • A.Store features in a shared BigQuery dataset without versioning
  • B.Export features to CSV files shared via Cloud Storage
  • ✓ C.Use Vertex AI Feature Store to define and serve features for both training and online prediction
  • D.Each team maintains its own feature engineering code in separate pipelines

Why C: Vertex AI Feature Store provides a central repository where features are defined once and reused across models, reducing training-serving skew.

Variation 2. A machine learning team wants to share features across multiple models to reduce training-serving skew and ensure consistency. Which Vertex AI service should they use?

easy
  • A.Vertex AI Workbench
  • B.Vertex AI Model Registry
  • ✓ C.Vertex AI Feature Store
  • D.Vertex AI Experiments

Why C: Vertex AI Feature Store centralizes feature storage, ensuring the same features are used for training and serving, reducing training-serving skew.

Variation 3. A data science team wants to share engineered features across multiple projects while ensuring low-latency serving for online predictions. Which Google Cloud service should they use to store and serve these features?

easy
  • A.Vertex AI Model Registry
  • B.Cloud Storage
  • C.BigQuery
  • ✓ D.Vertex AI Feature Store

Why D: Vertex AI Feature Store is purpose-built to store, share, and serve machine learning features with low-latency online serving and consistent offline serving for training. It lets a data science team centralize engineered features so multiple projects reuse them, while providing an online serving endpoint that returns feature values in milliseconds for real-time predictions.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.