Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 601–675

775 questions total · 11pages · All types, answers revealed

Page 8

Page 9 of 11

Page 10
601
MCQeasy

A data analyst wants to train a linear regression model to predict house prices using only SQL queries on BigQuery. Which BigQuery ML model type should they use?

A.BOOSTED_TREE_REGRESSOR
B.LOGISTIC_REG
C.LINEAR_REG
D.DNN_REGRESSOR
AnswerC

LINEAR_REG is BigQuery ML's built-in linear regression model type, trained directly with CREATE MODEL on a numeric label such as house price. It satisfies the constraint of using only SQL, requiring no exported data or external framework.

Why this answer

The question specifies a linear regression model for predicting house prices, which is a regression task with a continuous target variable. BigQuery ML's LINEAR_REG model type is explicitly designed for linear regression, making it the correct choice for this use case.

Exam trap

Google often tests the distinction between regression and classification model types, and the trap here is that candidates might confuse LOGISTIC_REG (classification) with linear regression due to the word 'logistic' sounding similar to 'linear', or they might overcomplicate the solution by choosing a tree or neural network model when a simple linear model suffices.

How to eliminate wrong answers

Option A is wrong because BOOSTED_TREE_REGRESSOR is a tree-based ensemble method, not a linear model, and is overkill for a simple linear regression task. Option B is wrong because LOGISTIC_REG is used for binary classification, not regression (predicting continuous values like house prices). Option D is wrong because DNN_REGRESSOR is a deep neural network regressor, which is unnecessarily complex and not a linear model.

602
MCQmedium

A developer needs to transcribe phone calls with high accuracy for a call center analytics application. The audio is in English and has background noise. Which Speech-to-Text model should they choose?

A.telephony
B.latest_short
C.latest_long
D.Any model; they are all equivalent
AnswerA

The telephony model is trained on 8 kHz narrowband audio, matching the bandwidth of phone calls, and is optimised to suppress background noise. This directly satisfies the stem's call centre constraint: English telephone speech with noise, where the default model's wideband training would degrade accuracy.

Why this answer

The 'telephony' model is specifically optimized for transcribing audio from phone calls, which often includes background noise, low fidelity, and narrowband audio. It is designed to handle the unique characteristics of telephony audio, such as 8kHz sampling rate and compression artifacts, providing higher accuracy for call center analytics. Other models like 'latest_short' and 'latest_long' are general-purpose and may not perform as well on noisy phone calls.

Exam trap

PMLE often tests the misconception that any Speech-to-Text model can be used interchangeably, but the trap is failing to recognize that telephony audio requires a specialized model due to its unique acoustic properties.

How to eliminate wrong answers

Option B is wrong because 'latest_short' is optimized for short utterances (e.g., voice commands) and does not handle the background noise and telephony-specific audio characteristics as effectively. Option C is wrong because 'latest_long' is designed for long-form audio like interviews or lectures, but it is not specifically tuned for telephony noise and compression. Option D is wrong because models are not equivalent; each is optimized for different audio types, and using the wrong model can significantly reduce transcription accuracy.

603
MCQeasy

A company has deployed a fraud detection model on Vertex AI Prediction. After three months, the model's accuracy has degraded, and the business is losing money due to undetected fraud. What should the team implement to proactively detect such issues?

A.Enable Vertex AI Model Monitoring to track prediction drift and alert when metrics exceed thresholds.
B.Set up Cloud Logging to capture all prediction requests and responses for manual review.
C.Randomly shuffle the training data before retraining to improve robustness.
D.Schedule a monthly job to retrain the model with the latest data without monitoring.
AnswerA

Vertex AI Model Monitoring continuously compares incoming prediction requests against the training baseline, detecting feature skew and prediction drift. This satisfies the stem's requirement to proactively detect degradation before financial loss accumulates, triggering alerts when drift metrics breach configured thresholds so the team can retrain the fraud model.

Why this answer

Vertex AI Model Monitoring tracks prediction drift and alerts when metrics exceed thresholds, enabling proactive detection of model degradation. Option B is wrong because Cloud Logging captures requests and responses but does not automatically detect drift, requiring manual review. Option C is wrong because shuffling training data does not help detect drift.

Option D is wrong because scheduling retraining without monitoring cannot proactively detect issues before they cause loss.

Exam trap

A common trap is to confuse monitoring with logging or retraining. Logging provides data but not automated drift detection; retraining fixes drift but doesn't detect it proactively.

604
MCQmedium

Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?

A.Create two separate endpoints and use a load balancer to distribute traffic between them.
B.Deploy both model versions to the same endpoint and set a traffic split percentage.
C.Deploy the new model version as a separate endpoint and update the client to call both endpoints.
D.Use a Vertex AI batch prediction job to send a percentage of live traffic to the new model.
AnswerB

Vertex AI Endpoints support deploying multiple models to the same endpoint and assigning a traffic split percentage to each deployed model. This allows a gradual rollout to a new version while keeping the client application pointed at a single endpoint URL, which matches the requirement exactly.

Why this answer

Deploying multiple models to a single endpoint and configuring a traffic split is the native Vertex AI mechanism for canary or A/B testing. The endpoint continues to expose one URL, so clients remain unchanged, and the split percentage controls how much live traffic reaches each model version. This enables safe evaluation of the new version with minimal risk.

Exam trap

The trap here is believing that a separate endpoint plus an external load balancer is equivalent to an endpoint traffic split, when the native feature avoids client changes and additional infrastructure.

605
MCQhard

A company uses Vertex AI Pipelines to orchestrate ML workflows. After a pipeline run, they want to query the lineage of a particular model artifact to find out which dataset and hyperparameters were used to produce it. Which API method should they use?

A.projects.locations.metadataStores.artifacts.queryArtifactLineageSubgraph
B.projects.locations.metadataStores.artifacts.get
C.projects.locations.metadataStores.contexts.addContextArtifactsAndExecutions
D.projects.locations.metadataStores.executions.queryExecutionInputsAndOutputs
AnswerA

queryArtifactLineageSubgraph traverses the metadata store's lineage graph in both directions from a given artifact, returning the executions and artifacts that produced or consumed it. This directly satisfies the requirement to trace a model artifact back to its source dataset and hyperparameters.

Why this answer

The queryArtifactLineageSubgraph method returns the lineage subgraph for a given artifact, showing the executions, contexts, and other artifacts connected to it — exactly what is needed to trace which dataset and hyperparameters produced a model. It traverses both upstream (inputs) and downstream (outputs) relationships in the metadata store.

Exam trap

PMLE often tests the difference between a single-node lookup (artifacts.get) and a graph traversal (queryArtifactLineageSubgraph) — candidates pick the simpler get method and miss that lineage requires traversing relationships.

How to eliminate wrong answers

Option B is wrong because artifacts.get only retrieves the artifact's own metadata (name, URI, properties) and does not traverse lineage relationships. Option C is wrong because contexts.addContextArtifactsAndExecutions is a write operation that associates artifacts and executions with a context — it does not query lineage. Option D is wrong because executions.queryExecutionInputsAndOutputs returns the inputs and outputs of a single execution, which is only one hop and does not give the full lineage subgraph across the pipeline.

606
MCQmedium

An ML engineer is building a Vertex AI pipeline that includes a component to train a model. The component takes a long time to run and occasionally fails due to transient errors (e.g., network timeouts). The engineer wants to automatically retry the component a few times if it fails. How should the engineer configure this?

A.Use a Cloud Scheduler job to trigger the pipeline again if it fails, and rely on the pipeline to resume from the failed step.
B.Set the pipeline's execution timeout to a high value and hope that the component succeeds on its own.
C.Set the retry policy on the pipeline task using the 'retry' parameter in the component's task definition, specifying the number of retries and backoff.
D.Configure the component to catch exceptions and loop internally until it succeeds, with no limit on retries.
AnswerC

Vertex AI Pipelines allows you to set a retry policy on individual tasks. By specifying the retry parameter with max_retries and backoff settings, the pipeline will automatically retry the component if it fails, which is ideal for transient errors.

Why this answer

To automatically retry a component on failure, the engineer should set a retry policy on the task. This is done by specifying the retry parameter in the task definition, including max_retries and backoff settings. This leverages Vertex AI Pipelines' built-in retry mechanism, which is designed for transient failures.

Exam trap

The trap here is thinking that re-triggering the entire pipeline or increasing timeouts solves transient failures, when a task-level retry policy is the correct mechanism.

607
MCQeasy

You need to run a custom training job on Vertex AI using a pre-built container for scikit-learn. Which container image should you specify?

A.us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-9
B.us-docker.pkg.dev/vertex-ai/training/scikit-learn-cpu.0.23-0
C.us-docker.pkg.dev/vertex-ai/training/xgboost-cpu.1-3
D.us-docker.pkg.dev/vertex-ai/training/tf-cpu.2-6
AnswerB

Vertex AI provides pre-built training containers hosted in its Artifact Registry; the scikit-learn CPU image at that path bundles the framework and Python dependencies needed to run custom scikit-learn training code. Specifying it satisfies the stem's requirement to use a pre-built container rather than building your own.

Why this answer

For a scikit-learn custom training job on Vertex AI, you must specify a pre-built container image that includes scikit-learn. The correct image is us-docker.pkg.dev/vertex-ai/training/scikit-learn-cpu.0.23-0, which is the official Vertex AI pre-built training container for scikit-learn version 0.23 on CPU. Other images correspond to different frameworks (PyTorch, XGBoost, TensorFlow).

Exam trap

PMLE often tests the ability to match the correct pre-built container image to the framework, causing candidates to confuse scikit-learn with XGBoost or TensorFlow images.

How to eliminate wrong answers

Option A is wrong because pytorch-gpu.1-9 is a PyTorch GPU training container, not scikit-learn. Option C is wrong because xgboost-cpu.1-3 is an XGBoost training container, not scikit-learn. Option D is wrong because tf-cpu.2-6 is a TensorFlow CPU training container, not scikit-learn.

608
Matchingmedium

Match each ML acronym to its definition.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Area Under the ROC Curve

Mean Squared Error

Tensor Processing Unit

Support Vector Machine

Principal Component Analysis

Why these pairings

The correct matches are A (ML), B (AI), C (NLP), and D (CNN). Options E and F are incorrect because they swap definitions: E assigns the definition of AI to ML, and F assigns the definition of CNN to NLP.

609
Multi-Selectmedium

A machine learning team uses Vertex AI Pipelines for model training. They want to implement a conditional step that runs additional evaluation if the model accuracy exceeds 0.9, otherwise it runs a data augmentation component. Which two Kubeflow Pipelines SDK v2 constructs can they use to achieve this? (Choose two.)

Select 2 answers
A.dsl.ParallelFor
B.dsl.Else
C.dsl.ExitHandler
D.dsl.Collected
E.dsl.If
AnswersB, E

dsl.Else provides the alternative branch executed when the dsl.If condition evaluates false, so the data augmentation component runs whenever accuracy does not exceed 0.9. It is the required companion construct to dsl.If for building if/else conditionals in Kubeflow Pipelines SDK v2.

Why this answer

The team needs branching logic in a Vertex AI Pipelines (Kubeflow Pipelines SDK v2) graph, and dsl.If (option E) is the SDK v2 construct that creates a conditional branch whose condition is evaluated at runtime, so it correctly expresses 'if accuracy > 0.9, run the evaluation component.' dsl.Else (option B) is the companion construct used together with dsl.If to define the alternative branch, so it correctly expresses 'otherwise run the data augmentation component.' Together, dsl.If and dsl.Else form the if/else control-flow pattern required for this scenario. dsl.ParallelFor (option A) is for iterating over a list and launching parallel tasks, not for conditional branching. dsl.ExitHandler (option C) is for defining cleanup or exit tasks that run when a scope completes or fails, not for choosing between two branches. dsl.Collected (option D) is a type annotation used to collect outputs from loops, not a conditional construct.

Exam trap

The trap here is that candidates may confuse `dsl.ParallelFor` or `dsl.ExitHandler` with conditional constructs, but `dsl.If` and `dsl.Else` are the only SDK v2 constructs specifically designed for branching based on runtime conditions.

610
MCQhard

You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?

A.Deploy the model on GKE with Triton and expose via Istio.
B.Use the prebuilt Vertex AI PyTorch prediction container and set environment variables to enable Triton.
C.Use Vertex AI Model Optimization to automatically convert the model to TensorRT and deploy with built-in server.
D.Upload your Triton container to Container Registry and specify it as the prediction container in Vertex AI Model.
AnswerD

Vertex AI accepts a custom container as the prediction container, so pushing the Triton image to Container Registry and referencing it lets Vertex AI route prediction traffic through Triton's dynamic batching and concurrent model execution, satisfying the requirement to serve PyTorch via Triton.

Why this answer

To deploy a custom Triton Inference Server container on Vertex AI, you upload your container to Artifact Registry (or Container Registry) and specify it as the prediction container when creating the Vertex AI Model resource. Vertex AI supports custom containers for prediction, allowing you to run Triton with your model artifacts. This is the standard approach for custom serving frameworks.

Exam trap

PMLE often tests the misconception that prebuilt Vertex AI containers can be toggled to use Triton via environment variables, when in fact a custom container is required.

How to eliminate wrong answers

Option A is wrong because deploying on GKE with Istio is a manual, self-managed approach that bypasses Vertex AI's managed prediction service; the question asks for the Vertex AI serving configuration. Option B is wrong because the prebuilt Vertex AI PyTorch prediction container does not include Triton, and environment variables cannot enable a server that is not installed. Option C is wrong because Vertex AI Model Optimization converts models to TensorRT and deploys with the built-in server, but it does not use Triton and may not support all PyTorch models.

611
Multi-Selecthard

A team is operationalizing a machine learning pipeline using Vertex AI. They want to automatically track experiment runs, log model parameters and metrics, and store model artifacts for reproducibility. They also need to capture lineage between pipeline components (e.g., which dataset and hyperparameter tuning job produced a model). Which TWO services should they use together to achieve this? (Choose two.)

Select 2 answers
A.Vertex AI Model Registry
B.Vertex AI Feature Store
C.Vertex AI Metadata
D.Vertex AI Experiments
E.Vertex AI Workbench
AnswersC, D

Vertex AI Metadata records and stores lineage artefacts, capturing which dataset, pipeline component and hyperparameter tuning job produced each model. This satisfies the lineage requirement, letting teams trace model provenance across pipeline components for reproducibility and audit.

Why this answer

Vertex AI Experiments (D) is correct because it automatically tracks and compares experiment runs, logging parameters, metrics, and artifacts so results are reproducible and comparable across training runs. Vertex AI Metadata (C) is correct because it records lineage and context for ML artifacts, capturing relationships such as which dataset and hyperparameter tuning job produced a given model, which is exactly the lineage requirement. Together they cover both experiment tracking and artifact lineage.

Vertex AI Model Registry (A) manages model versions and deployment lifecycle but does not itself track experiment runs or component lineage. Vertex AI Feature Store (B) serves and manages feature values for training/serving, not experiment tracking or lineage. Vertex AI Workbench (E) is a notebook development environment and does not provide the required automatic tracking or lineage services.

Exam trap

PMLE often tests the overlap between Experiments and Metadata — candidates pick one service thinking it covers both experiment tracking and lineage, when the scenario explicitly requires both capabilities together.

612
Matchingmedium

Match each ML model interpretability method to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Game-theoretic approach to explain feature contributions

Local surrogate model to explain individual predictions

Ranking features by their impact on model output

Shows marginal effect of a feature on predictions

Measures decrease in performance when feature is shuffled

Why these pairings

LIME (Local Interpretable Model-agnostic Explanations) provides local explanations by approximating the model's behavior near a specific prediction using a simpler interpretable model. SHAP (SHapley Additive exPlanations) uses Shapley values from game theory to fairly distribute feature contributions for individual predictions. Partial Dependence Plots (PDP) show the average marginal effect of one or two features on the predicted outcome across the dataset, making it a global method.

Permutation Feature Importance measures the increase in prediction error when a feature's values are randomly shuffled, indicating feature importance globally. Common mistakes: confusing LIME with a global method (B) and misattributing SHAP to visualizing decision trees (D).

Exam trap

The most common trap is confusing local vs global interpretability methods. LIME and SHAP are local, while PDP and Permutation Importance are global.

613
MCQmedium

An ML engineer has a model trained in Vertex AI and wants to deploy it to an endpoint with autoscaling and traffic splitting for canary testing. They have the model artifact stored in Vertex AI Model Registry with alias 'champion'. What is the correct sequence of steps?

A.Upload model to registry, create endpoint, then deploy model to endpoint with traffic split.
B.Create endpoint, upload model to registry, then deploy model to endpoint with traffic split.
C.Create endpoint, deploy model directly from Cloud Storage, then add traffic split.
D.Upload model to registry, then create endpoint and deploy in one command using gcloud ai endpoints deploy-model.
AnswerA

Deployment requires the model in the registry first, then an endpoint, then deploying that model to the endpoint where traffic splits are configured. This sequence satisfies the canary traffic-splitting constraint, since traffic allocation is set at deploy time on the endpoint, not during upload.

Why this answer

The correct sequence is: upload the model to Vertex AI Model Registry (creating a model resource with the 'champion' alias), create an endpoint, then deploy the model to the endpoint with traffic split for canary testing. The model must exist in the registry before it can be deployed, and the endpoint must exist before deployment.

Exam trap

PMLE often tests the deployment sequence — candidates reverse the order (endpoint before model) or assume models can be deployed directly from GCS, missing the mandatory Model Registry upload step.

How to eliminate wrong answers

Option B is wrong because the model must be uploaded to the registry before it can be deployed — creating the endpoint first does not change the dependency that the model resource must exist. Option C is wrong because Vertex AI does not deploy models directly from Cloud Storage to an endpoint; the model must first be imported into Model Registry as a model resource. Option D is wrong because although gcloud can create an endpoint and deploy in sequence, the model must still be uploaded to the registry first, and the 'one command' framing omits the required upload step and the explicit traffic split configuration.

614
MCQhard

You are training a PyTorch model on Vertex AI using a custom container. The training script uses DistributedDataParallel (DDP) with NCCL backend across 4 nodes, each with 8 GPUs. You notice that training throughput is low and GPUs are underutilized. After profiling, you find that the data loading is the bottleneck. You need to improve throughput without changing the model. What should you do?

A.Increase the number of DataLoader worker processes and use a larger prefetch factor.
B.Reduce the batch size per GPU to decrease memory pressure and allow more concurrent kernels.
C.Switch from NCCL to Gloo backend for inter-node communication.
D.Use a larger machine type with more vCPUs to increase data loading throughput.
AnswerA

Increasing DataLoader workers and prefetch factor allows more parallel data loading and preloading of batches, reducing GPU idle time. This directly addresses the data loading bottleneck. It is a standard PyTorch optimization that requires minimal code changes and can significantly improve throughput in distributed training.

Why this answer

The data loading bottleneck in PyTorch distributed training is often alleviated by increasing the number of DataLoader worker processes and prefetch factor. This enables parallel data fetching and preloading of batches, keeping GPUs busy. Other options either do not address the bottleneck or could worsen performance.

Exam trap

The trap here is assuming that more vCPUs or a different communication backend will solve the issue, when the real fix is to tune the DataLoader to feed data faster.

615
MCQmedium

You have a prototype model trained on a single machine using scikit-learn. You now need to scale training to a larger dataset that does not fit in memory on one machine. You want to use Vertex AI training with minimal changes to your existing scikit-learn code. What should you do?

A.Use Vertex AI Training with a custom container that runs your scikit-learn script, and use a distributed framework such as Dask or Ray to partition the data across multiple workers.
B.Convert the scikit-learn model to a TensorFlow model and use MirroredStrategy to train across multiple GPUs.
C.Use Vertex AI Training with the pre-built scikit-learn container and set the worker count to more than one; Vertex AI will automatically distribute the scikit-learn training.
D.Use Vertex AI distributed training with a ParameterServerStrategy and wrap your scikit-learn estimator in a TensorFlow estimator.
AnswerA

Vertex AI custom training lets you bring your own container and code. By using a distributed framework like Dask or Ray inside the container, you can partition the dataset across workers and train a scikit-learn model in a distributed manner with relatively small changes to the original script. This directly addresses the out-of-memory issue while staying close to the existing code.

Why this answer

The best approach is to keep the scikit-learn code and run it in a custom container on Vertex AI, using a distributed framework like Dask or Ray to partition the data across workers. This scales beyond a single machine's memory while minimizing changes to the original training script. The other options either rely on incompatible distribution strategies, assume automatic distribution that does not exist, or require a full rewrite.

Exam trap

The trap here is assuming that Vertex AI automatically distributes any training code when you set multiple workers, or that TensorFlow distribution strategies apply to scikit-learn.

616
MCQmedium

A data scientist wants to use AutoML to classify images of retail products into categories. There are 50 categories and the dataset has 100,000 labelled images. Which Vertex AI AutoML service is most appropriate?

A.AutoML Tables
B.AutoML Video
C.AutoML Vision
D.AutoML NLP
AnswerC

AutoML Vision handles multi-class image classification directly, supporting up to 50 categories without custom architecture work. It ingests the 100,000 labelled images and trains a managed model, satisfying the stem's classification requirement. Vertex AI's image data type is the appropriate task-specific service here, unlike tabular or video alternatives.

Why this answer

AutoML Vision is Vertex AI's service for training custom image classification and object detection models. With 100,000 labeled images across 50 categories, AutoML Vision can train a multi-class image classifier that learns visual features for each product category. It is the appropriate choice when the input data is images and the task is classification.

Exam trap

PMLE often tests whether candidates match the data modality (image vs. tabular vs. text vs. video) to the correct AutoML service, so choosing AutoML Tables for image data is the common error.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for structured/tabular data (CSV, BigQuery), not images. Option B is wrong because AutoML Video is for video classification, object tracking, and action recognition, not still-image product classification. Option D is wrong because AutoML NLP handles text classification, entity extraction, and sentiment analysis, not image data.

617
MCQmedium

A team is building ML pipelines with Vertex AI. They want to reuse standard pipeline components across teams and enforce governance. What approach should they take?

A.Use Vertex AI Pipelines with pre-built and custom components organized in a component registry.
B.Store pipeline definitions in a shared Cloud Storage bucket and copy them manually.
C.Use Cloud Composer to orchestrate ad-hoc scripts.
D.Have each team build their own pipelines independently.
AnswerA

Vertex AI Pipelines executes containerised components, and storing pre-built and custom components in a component registry lets teams share and version them centrally, enforcing governance and reuse across teams. Components are defined by YAML specs, so standard interfaces are preserved.

Why this answer

Vertex AI Pipelines lets teams define ML workflows as DAGs of containerized components, and organizing those components in a component registry (e.g., Vertex AI's component registry or Artifact Registry) enables reuse across teams. This approach enforces governance through versioning, access control, and standardized interfaces, so teams share vetted components rather than duplicating pipeline logic.

Exam trap

PMLE often tests whether candidates choose ad-hoc or manual approaches (shared buckets, independent pipelines) over the standardized, governed Vertex AI Pipelines + component registry pattern, so picking 'shared Cloud Storage bucket' is the common trap.

How to eliminate wrong answers

Option B is wrong because storing pipeline definitions in a Cloud Storage bucket and copying them manually provides no versioning, access control, or discoverability — it is an anti-pattern that leads to drift and governance gaps. Option C is wrong because Cloud Composer orchestrates general-purpose workflows (often data/ETL) and is not the standard for reusable ML pipeline components with governance; ad-hoc scripts lack the structure and lineage of Vertex AI Pipelines. Option D is wrong because each team building pipelines independently defeats reuse and governance, causing duplicated effort and inconsistent standards.

618
MCQhard

A team deploys a real-time model using a custom container on Vertex AI Prediction. The container is large (5 GB) and cold starts are causing latency spikes. The endpoint is configured with `min_replica_count=0` to reduce cost. The team wants to keep the cost low while reducing cold starts. What is the best approach?

A.Set `min_replica_count=1` to keep at least one replica always warm.
B.Use a prebuilt container for the model framework to reduce image size.
C.Enable container memory optimization to reduce startup time.
D.Provision a Persistent Disk (SSD) for the container image to speed up download.
AnswerA

Setting min_replica_count to 1 keeps one replica permanently provisioned, eliminating cold-start latency from the 5 GB image while retaining autoscaling for spikes. It balances the stem's dual constraints: reducing cold starts without abandoning cost control entirely.

Why this answer

Setting min_replica_count=1 keeps at least one replica always provisioned and warm, eliminating cold starts for the first request after idle while still allowing autoscaling to add replicas under load. This directly addresses the latency spikes caused by the 5 GB container's slow initialization, and the cost of one always-on replica is typically far lower than the business impact of cold-start latency.

Exam trap

PMLE often tests the trade-off between cost and latency, tempting candidates to pick image-size or infrastructure tweaks when the direct, supported fix is simply raising min_replica_count to keep a warm replica.

How to eliminate wrong answers

Option B is wrong because switching to a prebuilt container may not be possible if the model requires custom dependencies, and it does not guarantee a smaller image or eliminate cold starts. Option C is wrong because Vertex AI does not offer a 'container memory optimization' feature that reduces startup time — this is a fabricated capability. Option D is wrong because Persistent Disk (SSD) is not a supported mechanism for accelerating container image download on Vertex AI Prediction; image pull performance is not configured via attached disks.

619
MCQeasy

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?

A.Use a fixed random seed in the training component and snapshot the training data.
B.Run the pipeline on a dedicated Vertex AI training cluster with the same machine type and accelerators.
C.Version all pipeline components and pin their dependencies to specific versions.
D.Store the pipeline definition in a Git repository and use the same pipeline template for all runs.
AnswerA

Reproducibility requires controlling both code and data. A fixed random seed ensures deterministic behavior in stochastic processes like weight initialization. Snapshotting the training data (e.g., using a BigQuery snapshot or copying to a versioned Cloud Storage bucket) ensures the same data is used. Together, these practices enable reproducible results even if the source data changes later.

Why this answer

Reproducibility in ML pipelines requires controlling both the code and the data. A fixed random seed makes training deterministic, and snapshotting the training data ensures the same input is used across runs. Without data versioning, changes in the source data would lead to different results.

Therefore, combining a fixed seed with data snapshots is the correct approach.

Exam trap

The trap here is focusing only on code versioning or hardware consistency, while overlooking the critical role of data versioning and random seed control in achieving reproducible ML results.

620
MCQhard

Your organization uses Vertex AI Pipelines for training. A compliance auditor asks you to prove which dataset version and which preprocessing code commit produced a model that is currently deployed. You need to retrieve this information programmatically for a specific model version. Which approach should you use?

A.List the endpoint's deployed models and read the model description field, which automatically records the dataset version and code commit.
B.Use Vertex ML Metadata to traverse the lineage from the model artifact to its parent execution and input artifacts.
C.Inspect the model's Cloud Storage directory for a metadata.json file that lists the dataset and code commit.
D.Query Vertex AI Experiments for the run that has the same display name as the model version.
AnswerB

Vertex ML Metadata stores artifacts, executions, and events, so you can start from the model artifact and walk backward to the training execution, then to the dataset artifact and code artifact that were inputs. This provides an auditable, programmatic chain of provenance for the deployed model version.

Why this answer

Vertex ML Metadata is the system of record for lineage, linking artifacts such as models and datasets through executions. Traversing lineages from the model artifact to its parent execution and input artifacts yields the dataset version and code commit programmatically, which is exactly what an auditor needs.

Exam trap

The trap here is trusting a human-readable description or display name instead of the machine-recorded lineage graph.

621
MCQeasy

An ML engineer needs to trigger a Vertex AI Pipeline on a recurring schedule, every 24 hours, to retrain a model with the latest data. Which approach should they use to set up this schedule?

A.Use Cloud Tasks to queue pipeline runs daily.
B.Create a Cloud Scheduler job that calls the Vertex AI API to create a pipeline job.
C.Set a cron expression in the pipeline definition file using the 'schedule' parameter.
D.Use the Vertex AI Pipelines UI to set a schedule directly on the pipeline.
AnswerB

Cloud Scheduler provides cron-based recurring triggers, satisfying the 24-hour retraining requirement. Its HTTP target can authenticate to the Vertex AI API and invoke pipelines.create, launching a pipeline run on each fire. This avoids maintaining external orchestration infrastructure, unlike Cloud Functions or manual invocation, and natively supports the fixed daily cadence the stem demands.

Why this answer

Cloud Scheduler is the native Google Cloud service for cron-based job scheduling. By configuring a Cloud Scheduler job to call the Vertex AI API (e.g., via a HTTP POST to the projects.locations.pipelineJobs.create endpoint), the engineer can trigger a pipeline run every 24 hours. This approach is reliable, supports authentication via OAuth, and integrates directly with Vertex AI Pipelines without requiring additional orchestration code.

Exam trap

The PMLE exam often tests the misconception that Vertex AI Pipelines has a built-in scheduling feature (like a cron parameter in the pipeline definition or a UI schedule button), when in fact scheduling must be implemented using Cloud Scheduler or similar external services.

How to eliminate wrong answers

Option A is wrong because Cloud Tasks is a distributed task queue designed for asynchronous message delivery and retries, not for recurring cron-based scheduling; it would require an additional scheduler to enqueue tasks daily. Option C is wrong because Vertex AI Pipeline definitions do not support a 'schedule' parameter; scheduling is handled externally, not within the pipeline YAML or JSON definition. Option D is wrong because the Vertex AI Pipelines UI does not provide a built-in recurring schedule feature; schedules must be created using Cloud Scheduler or other external tools.

622
MCQmedium

A retail company has deployed a demand forecasting model on a Vertex AI Endpoint. The model uses 20 numeric features. The MLOps team wants Vertex AI Model Monitoring to detect training-serving skew for each feature and receive alerts when skew exceeds a threshold. They have enabled Model Monitoring for the endpoint and configured a monitoring frequency of every 24 hours. However, after several days, no skew metrics appear in the Vertex AI console. What is the most likely cause?

A.The monitoring frequency must be set to every 1 hour for skew detection to work.
B.The endpoint must have at least 1000 prediction requests per day for skew metrics to appear.
C.The model's predictions must be logged to BigQuery before skew can be detected.
D.The training dataset used for the baseline was not provided or the baseline statistics were not generated.
AnswerD

Training-serving skew compares live prediction feature distributions against the training data distribution. If no baseline dataset was supplied when configuring the monitoring job, or if baseline statistics were not generated, Vertex AI cannot calculate skew. The console will show no skew metrics. Providing the training data or enabling baseline generation is required for skew detection.

Why this answer

Training-serving skew requires a baseline distribution from the training data. Without providing the training dataset or generating baseline statistics during monitoring configuration, Vertex AI cannot compute skew. The monitoring frequency, request volume, and prediction logging are not the cause.

The team must supply the training data or enable baseline generation to see skew metrics.

Exam trap

The trap here is assuming that enabling Model Monitoring automatically creates a baseline from the deployed model's training data without explicitly providing it.

623
MCQhard

A company trains a model using features from Vertex AI Feature Store. They notice training-serving skew because the feature values used at training time differ from those served online. How should they address this?

A.Use the same online store for both training and serving
B.Disable caching in the online store
C.Enable feature monitoring to detect drift
D.Use point-in-time correct retrieval from the offline store for training data
AnswerD

Point-in-time correct retrieval joins each training label to feature values as they existed at that timestamp, preventing future data leaking into training. This aligns offline training vectors with the online serving values, directly eliminating the training-serving skew caused by mismatched feature snapshots.

Why this answer

Training-serving skew occurs when the feature values used during training differ from those served at inference time. Point-in-time correct retrieval from the offline store ensures that, for each training example, the feature values used are exactly those that would have been available at that timestamp — matching what the online store would have served. This eliminates the temporal mismatch that causes skew.

Exam trap

PMLE often tests the misconception that training and serving can share the same online store for consistency — candidates pick 'use the same online store' thinking it guarantees identical values, when in fact the online store only holds the latest value and cannot provide historical point-in-time correctness.

How to eliminate wrong answers

Option A is wrong because using the online store for training is not how Vertex AI Feature Store is designed — the online store serves low-latency latest values, not historical point-in-time values, so it cannot reproduce training-time conditions. Option B is wrong because disabling caching only affects latency and freshness of online reads; it does nothing to reconcile the historical training data with serving-time values. Option C is wrong because feature monitoring detects drift after the fact but does not fix the root cause of skew — it is a detection tool, not a correction mechanism.

624
Multi-Selecteasy

A data analyst wants to use low-code ML to analyze text data. Which TWO Google Cloud services are appropriate?

Select 2 answers
A.Vertex AI Workbench
B.Document AI
C.Cloud Natural Language API
D.AutoML Natural Language
E.BigQuery ML for sentiment
AnswersC, D

Correct: Pre-trained sentiment and entity analysis via API.

Why this answer

Cloud Natural Language API is a low-code ML service that provides pre-trained models for analyzing text, including sentiment analysis, entity recognition, and syntax analysis, without requiring custom model training. It is appropriate for a data analyst who wants to quickly extract insights from text data using simple API calls.

Exam trap

The trap here is that candidates may confuse BigQuery ML's sentiment analysis feature (which is SQL-based and not a dedicated low-code service) with a standalone low-code ML service, or mistakenly think Vertex AI Workbench is low-code when it actually requires coding in Python or other languages.

625
MCQeasy

A machine learning engineer wants to manage multiple model versions and facilitate collaboration across teams. The goal is to track model lineage, versioning, and approvals. Which Vertex AI service should they use?

A.Vertex AI Model Registry
B.Vertex AI ML Metadata
C.Vertex AI Feature Store
D.Vertex AI Vizier
AnswerA

Vertex AI Model Registry provides centralised versioning, lineage tracking and approval workflows across teams, exactly the governance capabilities the stem requires. It organises model artefacts and their metadata rather than handling training or serving infrastructure.

Why this answer

Vertex AI Model Registry is the service for managing model versions, tracking lineage, and facilitating collaboration and approvals across teams. It provides a central repository where models are registered, versioned, and annotated with metadata, and it integrates with Vertex AI Pipelines and ML Metadata for lineage. This directly matches the requirement to track model lineage, versioning, and approvals.

Exam trap

PMLE often tests whether candidates confuse ML Metadata (the underlying lineage store) with Model Registry (the user-facing versioning and approval service), leading them to pick ML Metadata for versioning and approvals.

How to eliminate wrong answers

Option B is wrong because Vertex AI ML Metadata is a lower-level service that stores metadata about artifacts and executions; it underpins lineage but does not provide the user-facing model versioning and approval workflow that Model Registry offers. Option C is wrong because Vertex AI Feature Store manages feature definitions and serving for ML features, not model versions or approvals. Option D is wrong because Vertex AI Vizier is a hyperparameter tuning service that optimizes model performance, not a model registry or collaboration tool.

626
MCQmedium

A team is implementing CI/CD for ML using Cloud Build. They want to trigger a training pipeline in Vertex AI whenever a new model code is pushed to the main branch of the repository. Which Cloud Build configuration should they use to achieve this?

A.Set up a Cloud Build trigger that runs on push to any branch, and in the build step, use gcloud to submit a Vertex AI Pipeline job.
B.Use a Cloud Scheduler job to periodically check for new commits on main and trigger Cloud Build.
C.Use Cloud Functions to watch the repository and call Cloud Build on push to main.
D.Set up a Cloud Build trigger that runs on push to main branch, and in the build step, use gcloud to submit a Vertex AI Pipeline job.
AnswerD

A Cloud Build trigger scoped to pushes on the main branch fires the build, and a gcloud step submits the Vertex AI Pipeline job. This satisfies the stem's constraint that training be triggered specifically by new model code pushed to main.

Why this answer

Cloud Build triggers can be configured to fire specifically on pushes to the main branch. The build step then uses the gcloud command to submit a Vertex AI Pipeline job, which directly integrates the CI/CD pipeline with Vertex AI's orchestration. This approach is event-driven, immediate, and requires no additional services or polling.

Exam trap

Google often tests the candidate's understanding that Cloud Build triggers can be scoped to specific branches and that using gcloud directly in a build step is the simplest and most efficient way to invoke Vertex AI Pipelines, rather than introducing unnecessary intermediate services like Cloud Functions or Scheduler.

How to eliminate wrong answers

Option A is wrong because triggering on push to any branch would cause the pipeline to run on feature branches, pull requests, and other non-main branches, leading to unnecessary executions and potential conflicts. Option B is wrong because Cloud Scheduler polling is inefficient, introduces latency, and is not the intended event-driven mechanism; Cloud Build triggers are designed to react to repository events directly. Option C is wrong because using Cloud Functions as an intermediary adds unnecessary complexity and cost; Cloud Build natively supports repository event triggers without requiring a separate compute service.

627
Multi-Selectmedium

You are building a CI/CD pipeline for an ML model using Cloud Build. When code is pushed to the main branch, you want to automatically build a training image, run a Vertex AI pipeline, and if the model evaluation passes, deploy it to a staging endpoint. Which two components are essential for this CI/CD pipeline?

Select 2 answers
A.Cloud Scheduler to trigger the pipeline on a schedule.
B.Cloud Functions to deploy the model.
C.Vertex AI Pipelines to orchestrate training and evaluation.
D.Cloud Build trigger configured to respond to push events to the main branch.
E.Vertex AI Continuous Training service.
AnswersC, D

Vertex AI Pipelines orchestrates the training and evaluation steps as a managed DAG, executing each component container in sequence and surfacing the evaluation metrics that gate deployment. This satisfies the stem's requirement to run a pipeline and conditionally promote the model only when evaluation passes.

Why this answer

Option D is correct because a Cloud Build trigger is the native mechanism that responds to push events on the main branch, automatically starting the build of the training image and initiating the pipeline workflow. Option C is correct because Vertex AI Pipelines orchestrates the training and evaluation steps, and its evaluation component determines whether the model passes the quality gate before deployment to the staging endpoint. Option A is incorrect because Cloud Scheduler triggers on a time-based schedule, not on code push events, so it does not satisfy the push-to-main requirement.

Option B is incorrect because Cloud Functions is not the deployment mechanism for Vertex AI models; deployment is handled through Vertex AI endpoints or pipeline components. Option E is incorrect because Vertex AI Continuous Training is a managed retraining feature, not the CI/CD orchestration component needed to trigger and gate deployments on code changes.

Exam trap

Google often tests the distinction between event-driven triggers (Cloud Build trigger on push) and schedule-based triggers (Cloud Scheduler), so candidates mistakenly pick Cloud Scheduler when the requirement is for a code-push event.

628
MCQeasy

You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?

A.The model is too large for the machine type.
B.Cold start occurs because the endpoint scaled down to zero.
C.The VPC Service Controls are blocking the initial request.
D.The endpoint's autoscaling is misconfigured.
AnswerB

With minReplicas=0, Vertex AI scales the endpoint to zero replicas when idle, so the first request must provision a container and load the model — roughly 30 seconds. Once a replica is warm, subsequent requests hit it directly, which is why later latency drops sharply.

Why this answer

Vertex AI endpoints with minReplicas=0 scale down to zero when idle. The first request after a period of inactivity triggers a cold start, where the endpoint must provision a new VM instance and load the model, causing a ~30-second delay. Subsequent requests are fast because the instance remains warm and handles them without provisioning overhead.

Exam trap

Google often tests the distinction between cold start latency and persistent performance issues, so candidates may mistakenly attribute the initial delay to model size or network misconfiguration instead of recognizing the intentional scaling-to-zero behavior.

How to eliminate wrong answers

Option A is wrong because a model too large for the machine type would cause persistent latency or errors on every request, not just the first one after idle time. Option C is wrong because VPC Service Controls enforce network boundaries and would block all requests consistently, not just the initial one with a 30-second delay. Option D is wrong because the autoscaling configuration (minReplicas=0, maxReplicas=5) is correct for scaling to zero; the observed behavior is the expected cold start, not a misconfiguration.

629
Multi-Selecteasy

An ML engineer is monitoring a Vertex AI Endpoint and notices a spike in 5xx error rates. Which TWO metrics should they examine to diagnose the issue? (Choose 2)

Select 2 answers
A.Feature drift alert count
B.GPU utilization on the endpoint
C.CPU utilization on the endpoint
D.Vertex AI Model Monitoring skew score
E.Number of predictions per minute
AnswersB, C

GPU exhaustion can lead to prediction failures.

Why this answer

CPU/GPU utilization can indicate resource exhaustion causing errors. Prediction job failures metric directly shows failed predictions.

630
MCQmedium

Your team trains a model on a Vertex AI Workbench notebook and logs hyperparameters, metrics, and a confusion matrix. Your manager asks you to ensure that anyone in the organization can reproduce the exact training run and compare it with other runs without manually digging through notebook cells. Which Vertex AI component should you use to record this information?

A.Vertex AI Experiments
B.Vertex AI Pipelines
C.Vertex AI Model Registry
D.Vertex AI TensorBoard
AnswerA

Vertex AI Experiments records parameters, metrics, and artifacts from training runs, and each run is associated with a lineage context. It provides a searchable UI and API so teammates can compare runs and reproduce them without inspecting notebook code, which satisfies the requirement for organization-wide tracking.

Why this answer

A system of record for training runs must capture parameters, metrics, and artifacts in a searchable way so teammates can reproduce and compare results. Vertex AI Experiments provides exactly this by linking each run to metadata and lineage, while visualization, deployment, and orchestration tools address different stages of the lifecycle.

Exam trap

The trap here is confusing visualization with tracking, assuming that a dashboard of metric curves is sufficient to reproduce a run.

631
MCQhard

A company deploys a training pipeline on Vertex AI using custom containers. The pipeline includes a hyperparameter tuning job that uses Bayesian optimization. After several runs, they observe that the tuning job is not converging and the search space is large. They want to reduce the number of trials while still finding good hyperparameters. Which strategy should they use?

A.Increase the number of parallel trials to explore more points simultaneously.
B.Use Grid search instead of Bayesian optimization to systematically cover the search space.
C.Implement early stopping by using the 'early_stopping' flag in the hyperparameter tuning job.
D.Reduce the search space by applying feature selection and using prior knowledge.
AnswerD

Bayesian optimisation scales poorly with dimensionality, so pruning the search space via feature selection and encoding prior knowledge concentrates trials on promising regions. This reduces the number of trials needed while preserving solution quality, directly addressing the non-convergence and large search space.

Why this answer

Reducing the search space using prior knowledge directly decreases the number of trials needed. Option A is wrong because increasing parallel trials does not reduce the total number of trials. Option B is wrong because grid search generally requires more trials than Bayesian optimization.

Option C is wrong because early stopping reduces time per trial but does not reduce the number of trials.

632
Multi-Selecthard

Which THREE of the following are recommended practices for model governance and lineage in Vertex AI?

Select 3 answers
A.Enable Vertex AI ML Metadata to track artifacts, executions, and contexts.
B.Use Vertex AI Experiments to log parameters and metrics.
C.Store model artifacts in Cloud Storage with metadata in a database.
D.Manually record model lineage in a spreadsheet.
E.Use Vertex AI Model Registry to manage model versions and stages.
AnswersA, B, E

ML Metadata provides automated lineage tracking.

Why this answer

Vertex AI ML Metadata is a fully managed service that automatically tracks artifacts, executions, and contexts across the ML workflow. By enabling it, you create a lineage graph that records every step from data preparation to model deployment, which is essential for auditability and reproducibility. This is a core recommended practice for model governance because it provides an immutable, queryable history of all model-related activities.

Exam trap

Google Cloud often tests the distinction between using native Vertex AI services (like ML Metadata, Experiments, and Model Registry) versus ad-hoc or manual methods (like spreadsheets or custom databases) that lack automated governance and audit trails.

633
Multi-Selecteasy

Which TWO options are best practices for building ML pipelines on Vertex AI?

Select 2 answers
A.Use Cloud Functions to execute individual pipeline steps
B.Hardcode pipeline parameters in the component definitions
C.Use custom container components to encapsulate reusable logic
D.Always use the same compute environment for training and serving to ensure consistency
E.Leverage Vertex ML Metadata to track artifact lineage
AnswersC, E

Custom container components package code, dependencies and runtime into a portable image, letting teams reuse identical logic across pipelines without rewriting or re-resolving environments. This encapsulation satisfies the reusability best-practise requirement for Vertex AI Pipelines components.

Why this answer

Option C is correct because custom container components let you package arbitrary dependencies, libraries, and code into a portable, versioned image that Vertex AI Pipelines can reuse across multiple pipelines, which is a recommended way to encapsulate and share reusable logic. Option E is correct because Vertex ML Metadata automatically records artifacts, executions, and their lineage for Vertex AI Pipelines runs, enabling reproducibility, auditing, and comparison of models and datasets across experiments. Option A is not a best practice because Cloud Functions are event-driven, short-lived, and not designed to run heavy or stateful ML pipeline steps; Vertex AI Pipelines components (or Vertex AI Training jobs) should execute those steps.

Option B is wrong because hardcoding parameters in component definitions reduces reusability and makes experimentation and CI/CD harder; parameters should be passed at pipeline submission time via PipelineJob runtime parameters. Option D is wrong because training and serving often require different compute profiles (for example, GPUs for training versus optimized CPUs or different accelerators for serving), and consistency is achieved through containerized artifacts and pinned dependencies, not by forcing identical compute environments.

Exam trap

Google Cloud often tests the misconception that serverless functions like Cloud Functions are suitable for ML pipeline steps, but the trap is that ML steps require persistent state, longer timeouts, and specialized hardware, which Cloud Functions cannot provide.

634
MCQmedium

You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?

A.Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=5, and enable autoscaling based on CPU utilization.
B.Use a CPU-only machine type with minReplicaCount=5 and maxReplicaCount=5 to ensure high availability.
C.Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=1.
D.Use a GPU-enabled machine type with minReplicaCount=1 and maxReplicaCount=1.
AnswerA

This configuration uses cost-effective CPU instances and allows the endpoint to scale out when CPU utilization rises, ensuring latency remains low under load. Setting minReplicaCount=1 keeps idle cost low, while maxReplicaCount=5 provides headroom. Autoscaling based on CPU is appropriate for a CPU-bound model. This balances cost and performance effectively.

Why this answer

For a small CPU-bound model with variable traffic, the optimal cost-latency trade-off is a CPU-only machine type with a low minimum replica count and autoscaling enabled. This keeps baseline cost low while allowing the endpoint to add replicas when CPU utilization increases, preserving the 100 ms latency target. GPU instances and fixed high replica counts unnecessarily increase cost without improving latency.

Exam trap

The trap here is assuming that GPUs or a fixed number of replicas are needed for low latency, when a small CPU model can meet the SLO with autoscaling.

635
MCQeasy

A data scientist needs to train a time-series forecasting model on historical sales data stored in BigQuery to predict future demand. The data has strong seasonal patterns. Which BigQuery ML model type should they use?

A.MATRIX_FACTORIZATION
B.BOOSTED_TREE_REGRESSOR
C.ARIMA_PLUS
D.K_MEANS
AnswerC

ARIMA_PLUS handles seasonality natively through automatic seasonal decomposition and multiple seasonal period detection, satisfying the strong seasonal patterns constraint in the historical sales data. It also supports forecasting horizons directly in BigQuery ML, letting the data scientist train and predict demand without exporting data.

Why this answer

ARIMA_PLUS is the correct choice because it is specifically designed for time-series forecasting in BigQuery ML, handling seasonal patterns, trend decomposition, and automatic hyperparameter tuning. It models autoregressive (AR) and moving average (MA) components with seasonal differencing, making it ideal for historical sales data with strong seasonal cycles.

Exam trap

Google often tests the misconception that any regression model (like BOOSTED_TREE_REGRESSOR) can be naively applied to time-series data, ignoring the need for specialized models that handle temporal dependencies and seasonality natively.

How to eliminate wrong answers

Option A is wrong because MATRIX_FACTORIZATION is used for recommendation systems (e.g., collaborative filtering) and cannot model temporal dependencies or seasonality in time-series data. Option B is wrong because BOOSTED_TREE_REGRESSOR is a tree-based ensemble method for regression tasks but does not inherently capture time-series structures like seasonality, trend, or autocorrelation without extensive feature engineering. Option D is wrong because K_MEANS is an unsupervised clustering algorithm that groups data points by similarity and has no mechanism for forecasting future values or modeling sequential patterns.

636
MCQmedium

An ML engineer has set up Vertex AI Model Monitoring on an endpoint with a sampling rate of 0.1 (10%). They notice that the monitoring job runs hourly but the reported drift metrics seem inconsistent. What is the most likely cause?

A.The sampling rate is too low, leading to insufficient data for reliable drift statistics.
B.Prediction drift monitoring is not enabled; only feature drift is configured.
C.The drift detection algorithm is not suited for this model; try changing from JS divergence to L-infinity distance.
D.The monitoring frequency is too low; it should be set to every 5 minutes.
AnswerA

A 0.1 sampling rate means only 10% of prediction requests feed the drift calculation, so hourly windows may contain too few samples for statistically reliable distribution comparisons. Raising the sampling rate satisfies the need for sufficient data volume behind each reported drift metric.

Why this answer

A sampling rate of 0.1 means only 10% of prediction requests are logged and analyzed, which can produce statistically noisy or inconsistent drift metrics, especially for low-traffic endpoints. Drift detection relies on sufficient sample size to compute reliable distribution comparisons (e.g., JS divergence), so a low sampling rate is the most likely cause of the inconsistency.

Exam trap

The trap is assuming the monitoring configuration itself (frequency or metric choice) is broken, when the real issue is statistical — candidates overlook that sampling rate directly controls the sample size feeding drift calculations.

How to eliminate wrong answers

Option B is wrong because the question states drift metrics are being reported — if prediction drift were disabled, no prediction drift metrics would appear at all, rather than appearing inconsistent. Option C is wrong because changing the distance metric (JS divergence vs. L-infinity) is a tuning choice, not the cause of inconsistency from sparse sampling; both metrics are valid and the issue is data volume.

Option D is wrong because increasing frequency to every 5 minutes would not fix inconsistency caused by low sampling — it would just produce more noisy, small-sample results and increase cost.

637
MCQmedium

You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?

A.Store the model artifacts in a custom container image and push it to Artifact Registry.
B.Use a larger machine type with more vCPUs and memory for each replica.
C.Set the endpoint's minReplicaCount to a high value so that replicas are always warm.
D.Enable request-response logging on the endpoint to monitor startup latency.
AnswerA

Baking the model into a custom container image pulls the artifacts when the container image is downloaded, which happens as part of standard node provisioning. This reduces the additional model-download step at startup, cutting cold-start time and making scale-out more predictable.

Why this answer

Embedding the model artifacts in a custom container image eliminates the separate download from Cloud Storage during container startup. The image layers are pulled by the node as part of standard container initialization, which is typically faster and more predictable than fetching a large model from a GCS bucket after the container starts.

Exam trap

The trap here is assuming that raising the minimum replica count eliminates cold starts entirely, when it only masks them for steady-state traffic and does nothing for scale-out events.

638
Multi-Selecthard

Which TWO are best practices for implementing a low-code ML solution using Vertex AI AutoML? (Choose 2)

Select 2 answers
A.Use the AutoML recommended data split (train/validation/test) to avoid overfitting.
B.Impute missing values manually before uploading the dataset.
C.Normalize numerical features to zero mean and unit variance.
D.Enable automatic feature engineering by leaving feature columns as raw data.
E.Export the data and train a custom model with a different architecture.
AnswersA, D

Why A is correct: AutoML optimizes split for best performance.

Why this answer

AutoML's recommended data split (train/validation/test) is designed to prevent overfitting by ensuring the model is evaluated on unseen data. AutoML automatically handles the split ratio (e.g., 80/10/10) and stratification, which is a best practice for low-code ML solutions where manual split logic is error-prone.

Exam trap

Google Cloud often tests the misconception that manual preprocessing (like imputation or normalization) is required for AutoML, when in fact AutoML is designed to handle these steps automatically, and manual intervention can degrade performance or cause errors.

639
Multi-Selecthard

An e-commerce company uses a recommendation model that suggests products based on user browsing history. The model was trained on data from the past year and has high accuracy on the test set. However, after deployment, the click-through rate (CTR) on recommendations is much lower than expected. Which three steps should the data scientist take to diagnose and improve the model? (Choose THREE)

Select 3 answers
A.Run offline evaluation on a holdout dataset to confirm accuracy
B.Set up an A/B experiment comparing the model's recommendations against a baseline
C.Retrain the model on the most recent three months of data to capture recent trends
D.Check the distribution of predictions versus the training set to detect drift
E.Increase the training dataset size by including data from two years ago
AnswersB, C, D

An A/B experiment isolates whether the low CTR stems from the model itself or from external factors such as placement or latency, by comparing live recommendations against a baseline under identical conditions. This directly tests the deployment gap the stem describes, where offline accuracy failed to translate into online engagement.

Why this answer

Option B is correct because an A/B experiment comparing the deployed model against a baseline (e.g., current production recommender or popularity-based recommendations) directly measures real-world CTR impact and isolates whether the model itself is underperforming versus other factors like UI placement or latency. Option C is correct because retraining on the most recent three months of data addresses temporal concept drift — user browsing behavior and product trends change quickly in e-commerce, so a model trained on a year-old distribution can have stale item embeddings and relevance signals even with high offline accuracy. Option D is correct because comparing the distribution of predictions (and input feature distributions) against the training set detects data drift and covariate shift, which commonly explain why offline test accuracy fails to translate into online CTR.

Option A is not the best step because offline holdout evaluation was already performed and showed high accuracy, so repeating it won't reveal the deployment-time mismatch. Option E is not appropriate because adding two-year-old data would likely worsen drift by reinforcing outdated patterns rather than capturing recent trends.

Exam trap

Google Cloud often tests the misconception that high offline accuracy guarantees online success, ignoring that offline metrics can be misleading due to distribution shift, feedback loops, or mismatched optimization objectives (e.g., accuracy vs. CTR).

640
MCQhard

A team is training a large recommendation model on Vertex AI using a custom container. They need to log training metrics and visualize them in Vertex AI TensorBoard. The training code is written in PyTorch and runs on multiple worker nodes. Which of the following is the correct way to enable TensorBoard logging?

A.Install TensorBoard in the custom container and run it as a sidecar process on each worker node.
B.Use the Vertex AI SDK to create a TensorBoard instance and pass its resource name to the training job using the --tensorboard flag.
C.Write training metrics to a Cloud Storage bucket and then manually upload them to Vertex AI TensorBoard after training.
D.Use the torch.utils.tensorboard.SummaryWriter to write logs to a local directory, and Vertex AI will automatically sync it to TensorBoard.
AnswerB

The Vertex AI SDK allows creating a TensorBoard instance, and the training job can be configured with the tensorboard resource name. This enables automatic logging of metrics from the training container to the TensorBoard instance. The training code should write logs to the directory specified by the AIP_TENSORBOARD_LOG_DIR environment variable, which Vertex AI sets.

Why this answer

To integrate with Vertex AI TensorBoard, you must create a TensorBoard instance and pass its resource name to the training job. The training code should write logs to the directory specified by the AIP_TENSORBOARD_LOG_DIR environment variable. This allows Vertex AI to automatically collect and display metrics from all workers in the managed TensorBoard service.

Exam trap

The trap here is assuming that any local logging will be automatically synced to Vertex AI TensorBoard, but explicit configuration of the TensorBoard instance is required.

641
MCQhard

A company uses Vertex AI Matching Engine for a product recommendation system. They need to update the index with new product embeddings every hour, but the index is used for online queries with low latency. Which index update strategy should they use?

A.Use streaming updates to insert new embeddings incrementally
B.Use a hybrid approach with batch for daily full rebuild and streaming for hourly
C.Use batch updates to replace the index every hour
D.Recreate the index from scratch each hour
AnswerA

Streaming updates let Matching Engine insert or delete datapoints incrementally while the index stays queryable, so hourly embedding refreshes avoid the rebuild-and-redeploy cycle that batch updates require. This preserves the low-latency online serving the stem demands, since queries continue against the live index throughout.

Why this answer

Streaming updates in Vertex AI Matching Engine allow incremental insertion of new embeddings into an existing index without rebuilding it. This satisfies the requirement for hourly updates while maintaining low-latency online queries, as the index remains available and consistent during the update process.

Exam trap

Google often tests the misconception that batch updates are required for consistency or that streaming updates cannot handle frequent changes, leading candidates to choose hybrid or batch approaches when incremental streaming is both sufficient and optimal for low-latency online serving.

How to eliminate wrong answers

Option B is wrong because a hybrid approach with batch for daily full rebuild and streaming for hourly adds unnecessary complexity and cost; streaming updates alone suffice for hourly increments without needing a daily rebuild. Option C is wrong because batch updates replace the entire index, causing downtime or increased latency during the rebuild, which violates the low-latency online query requirement. Option D is wrong because recreating the index from scratch each hour is inefficient, time-consuming, and disrupts query availability, making it unsuitable for real-time serving.

642
MCQmedium

An organization wants to trigger a Vertex AI pipeline whenever a new commit is pushed to the main branch of their Cloud Source Repository. The pipeline should retrain and evaluate the model. Which service should they use to detect the push event and start the pipeline?

A.Vertex AI Pipeline schedule
B.Cloud Build trigger
C.Cloud Scheduler on a short interval
D.Pub/Sub with push subscription to a Cloud Function
AnswerB

Cloud Build triggers natively watch Cloud Source Repository branches and fire on push events, satisfying the commit-detection constraint. The trigger then invokes the Vertex AI pipeline, so no polling or custom webhook infrastructure is needed to bridge the repository event to pipeline execution.

Why this answer

Cloud Build triggers are designed to automatically invoke a build pipeline in response to events from Cloud Source Repository, such as a push to a specific branch. This allows the organization to directly start a Vertex AI pipeline for retraining and evaluation without additional infrastructure. Cloud Build triggers natively integrate with Cloud Source Repository, making them the simplest and most reliable choice for this event-driven workflow.

Exam trap

This question tests the distinction between event-driven triggers (Cloud Build) and time-based schedulers (Cloud Scheduler, Vertex AI Pipeline schedule), leading candidates to mistakenly choose a polling or cron-based option for an event-driven requirement.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipeline schedules are time-based (cron) triggers, not event-driven; they cannot detect a Git push event. Option C is wrong because Cloud Scheduler on a short interval would poll for changes, introducing latency and inefficiency, and it does not natively detect push events from Cloud Source Repository. Option D is wrong because while Pub/Sub with a push subscription to a Cloud Function could technically work, it adds unnecessary complexity and an extra compute layer; Cloud Build triggers provide a direct, managed integration without the need for custom code or additional services.

643
Multi-Selecthard

A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)

Select 2 answers
A.Set min_replicas = 0 to allow scale-to-zero and save costs.
B.Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.
C.Set min_replicas = 3 to keep a baseline of warm instances.
D.Enable Vertex AI Model Optimization for automatic quantization.
E.Use batch prediction instead of online prediction.
AnswersB, C

GPU acceleration (N1 with T4) cuts inference compute time, directly addressing the p99 < 100ms SLO that CPU inference on a TensorFlow fraud model would likely breach. It does not solve cold starts, so it pairs with warm replicas.

Why this answer

Option B is correct because a GPU-enabled machine type such as an N1 instance with an NVIDIA T4 accelerator provides the parallel compute throughput needed to keep TensorFlow inference within a p99 latency SLO under 100ms, which CPU-only serving often cannot guarantee for larger models. Option C is correct because setting min_replicas = 3 keeps a baseline of warm, already-loaded model instances, eliminating cold-start latency for the initial requests and giving the autoscaler headroom to absorb traffic spikes without waiting for new replicas to initialize. Option A is not appropriate because min_replicas = 0 enables scale-to-zero, which directly reintroduces cold-start latency and violates the strict p99 < 100ms SLO.

Option D is not selected because Vertex AI Model Optimization (quantization) is a model-compression technique that may reduce latency but is not a required configuration for meeting the SLO and can degrade accuracy. Option E is not selected because batch prediction processes data offline in bulk and cannot serve real-time fraud detection requests with sub-100ms latency.

Exam trap

A common misconception is that scale-to-zero (min_replicas = 0) is always cost-effective, but in latency-sensitive real-time inference, it introduces unacceptable cold-start delays, making baseline warm instances (min_replicas > 0) essential.

644
MCQhard

A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?

A.Use Vertex AI Hyperparameter Tuning to find optimal model partitioning
B.Use tf.distribute.MirroredStrategy across all GPUs
C.Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM
D.Use Vertex AI distributed training with TF_CONFIG to set up multi-worker mirrored strategy and rely on XLA to partition the model
AnswerC

Pipeline parallelism partitions consecutive model layers into stages across GPUs, so each device holds only a fraction of parameters. This directly addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

Why this answer

When a model does not fit into a single GPU's memory, model parallelism is required. Pipeline parallelism splits the model layers across multiple GPUs, and frameworks like PyTorch's RPC or Megatron-LM provide the necessary primitives to implement it. This is the correct strategy for fitting a large model like LLaMA 2 across multiple GPUs.

Exam trap

PMLE often tests the confusion between data parallelism (replicating the model) and model parallelism (splitting the model), causing candidates to choose MirroredStrategy when the model does not fit on a single GPU.

How to eliminate wrong answers

Option A is wrong because Vertex AI Hyperparameter Tuning is for tuning hyperparameters, not for partitioning a model across GPUs. Option B is wrong because tf.distribute.MirroredStrategy replicates the entire model on each GPU (data parallelism), which does not solve the memory issue when the model does not fit on one GPU. Option D is wrong because TF_CONFIG with multi-worker mirrored strategy also replicates the model and relies on XLA, which does not automatically partition a model across GPUs for model parallelism.

645
MCQmedium

A team deploys a model on Vertex AI that uses a custom prediction routine (CPR) with a dependency on a native library. The container crashes with 'ImportError: libcudart.so.11.0: cannot open shared object file'. How should they resolve this?

A.Build a custom container image that includes the CUDA runtime library.
B.Submit the model for batch prediction to avoid the error.
C.Request a GPU machine type for the endpoint.
D.Use a Vertex AI pre-built container for PyTorch instead.
AnswerA

The ImportError shows the CUDA runtime library is absent from the prediction environment. Vertex AI's prebuilt CPR containers do not bundle libcudart, so building a custom image with the CUDA runtime installed supplies the missing native dependency the model needs.

Why this answer

The error 'ImportError: libcudart.so.11.0: cannot open shared object file' indicates that the CUDA runtime library (version 11.0) is missing from the container environment. Since the custom prediction routine (CPR) depends on a native library that requires this CUDA runtime, the correct solution is to build a custom container image that includes the CUDA runtime library. This ensures the shared object is available at runtime, resolving the import error.

Exam trap

Google Cloud often tests the misconception that requesting a GPU machine type automatically provides the necessary CUDA libraries, but in reality, the CUDA runtime must be explicitly included in the container image, as the GPU machine type only provides the hardware and driver, not the user-space libraries.

How to eliminate wrong answers

Option B is wrong because submitting the model for batch prediction does not change the container environment; the same missing CUDA runtime library will cause the same ImportError during batch prediction. Option C is wrong because requesting a GPU machine type for the endpoint provides GPU hardware but does not install the CUDA runtime library into the container; the library must be present in the container image regardless of the underlying hardware. Option D is wrong because using a Vertex AI pre-built container for PyTorch does not guarantee inclusion of the specific CUDA runtime version 11.0 required by the native library; the pre-built container may have a different CUDA version or omit the library entirely.

646
MCQeasy

A data scientist runs a BigQuery ML prediction query and gets a region mismatch error. The model is in the US region, but the new_data table is in the EU region. What is the simplest way to resolve this?

A.Recreate the model in the EU region using the same training data
B.Copy the new_data table to the US region using the BigQuery UI or CLI
C.Enable cross-region query in BigQuery settings
D.Export the model from US and import it to EU
AnswerB

BigQuery ML requires the model and the input data to reside in the same region, so copying new_data into the US region places both objects together and the prediction query runs. This is the simplest fix, avoiding model retraining or dataset recreation.

Why this answer

The simplest fix is to move the new_data table to the same region as the model (US). BigQuery ML requires that the model and the data used for predictions reside in the same multi-region or regional location. Copying the table via the BigQuery UI or CLI (e.g., `bq cp`) is a straightforward, no-code operation that avoids retraining or exporting the model.

Exam trap

The trap here is that candidates may overthink the solution and choose to recreate the model or export/import it, not realizing that the simplest and most efficient fix is to copy the data table to the model's region.

How to eliminate wrong answers

Option A is wrong because recreating the model in the EU region would require retraining the model from scratch, which is unnecessary and time-consuming when a simple data copy resolves the mismatch. Option C is wrong because BigQuery does not support a 'cross-region query' setting; queries are always restricted to a single region or multi-region, and enabling such a feature is not possible. Option D is wrong because exporting and importing a model between regions is more complex and involves additional steps (e.g., using Cloud Storage as an intermediary), whereas copying the table is simpler and directly addresses the region mismatch.

647
MCQeasy

You are using Vertex AI Training to train a model and then automatically deploy the best candidate to a Vertex AI Prediction endpoint via the Vertex AI Model Registry. However, after deployment, you notice that the endpoint returns predictions for the new model, but they are significantly different from the evaluation metrics computed during training. The training scripts used TensorFlow with a serving input function. What is the most likely issue and how would you fix it?

A.The endpoint is using a different machine type affecting numerical precision; you should use the same machine type as training.
B.The serving input function's preprocessing steps do not match the training preprocessing; you should verify and align them.
C.The model registry deployed a different version; you should check the alias.
D.The model was saved with training-only metrics; you should retrain with evaluation metrics.
AnswerB

Training metrics are computed on preprocessed features, but the serving input function applies its own transformations. If those steps diverge, the endpoint receives differently scaled or encoded inputs, producing skewed predictions. Aligning the serving preprocessing with training preprocessing restores consistency.

Why this answer

The classic cause of a train/serve skew in Vertex AI is that the serving input function applies different preprocessing than the training pipeline — for example, different tokenization, normalization, or feature scaling. TensorFlow's serving_input_fn defines the graph that runs at prediction time, so any divergence from the training input_fn produces predictions that look correct structurally but are numerically off. Aligning the two preprocessing paths (ideally by sharing a single feature-engineering module) resolves the discrepancy.

Exam trap

PMLE often tests the train/serve skew concept by describing metrics that look good offline but bad online, tempting candidates to blame infrastructure (machine type) or versioning (alias) instead of the preprocessing mismatch.

How to eliminate wrong answers

Option A is wrong because machine type changes affect throughput and memory, not the mathematical precision of float32/float64 operations in a meaningful way — TensorFlow uses the same numeric types regardless of the underlying VM. Option C is wrong because a wrong model version would typically produce obviously wrong or shape-mismatched outputs, and the Model Registry alias is deterministic; it does not cause subtle metric drift. Option D is wrong because metrics are evaluation artifacts, not model weights — saving 'training-only metrics' is not a real failure mode and retraining would not fix a serving preprocessing mismatch.

648
MCQhard

Your team has deployed a text classification model on Vertex AI Endpoints. You notice that the model's latency has increased significantly over the last week, but the request rate has remained stable. Which of the following is the most likely cause?

A.A sudden increase in the number of prediction requests
B.The model was replaced with a larger version without updating the endpoint
C.A change in the preprocessing logic that now includes a computationally expensive step
D.A misconfiguration in the autoscaling policy
AnswerC

A change in preprocessing logic adds a computationally expensive step, directly increasing per-request processing time while request volume stays constant. This satisfies the stem's constraint: stable request rate with rising latency points to per-request cost, not load. Vertex AI Endpoints latency reflects the full prediction pipeline, including client-side preprocessing.

Why this answer

A computationally expensive preprocessing step directly increases per-request latency on the inference path, even when request rate is stable. Vertex AI Endpoints execute user-provided preprocessing code before model inference, so adding a heavy operation (e.g., large regex, image resizing, or external API call) will linearly increase response time for every prediction.

Exam trap

The trap here is that candidates confuse 'model latency' with 'request rate' and assume any latency increase must be due to scaling issues, ignoring that preprocessing logic changes can dramatically affect per-request performance without altering throughput.

How to eliminate wrong answers

Option A is wrong because a sudden increase in request rate would cause latency to rise, but the question explicitly states request rate has remained stable. Option B is wrong because replacing the model with a larger version requires deploying a new model to the endpoint or updating the endpoint's deployed model; simply replacing the model binary without updating the endpoint's deployment configuration would not change the model served, so latency would not increase. Option D is wrong because a misconfiguration in autoscaling policy (e.g., too few min replicas) would cause latency to increase only when request rate exceeds the current serving capacity, but request rate is stable and autoscaling would have already scaled to match the stable load.

649
MCQhard

You are collaborating on a Vertex AI Feature Store implementation. A data engineer updates a feature's values in the offline store, but the online store still serves the old values for several hours. The online store is configured with a feature value TTL of 24 hours and uses batch ingestion. What is the most likely cause of the stale online values?

A.The online store is configured to read from the offline store at query time, and the delay is due to eventual consistency between the two stores.
B.The feature values are ingested into the offline store only, and the online store is not being updated because batch ingestion does not automatically sync to the online store unless a separate ingestion job is run.
C.The online store's feature value TTL is set too high, causing it to serve cached values until the TTL expires.
D.The feature's online store TTL has expired, causing the online store to fall back to the offline store for the latest values.
AnswerB

In Vertex AI Feature Store, offline and online stores are separate. Batch ingestion writes to the offline store, and you must explicitly run an ingestion job to update the online store, or use streaming ingestion for real-time updates. If the data engineer only updated the offline store, the online store will not reflect the changes until a sync job is executed.

Why this answer

Vertex AI Feature Store maintains separate offline and online stores. Batch ingestion updates the offline store, but the online store requires its own ingestion job to sync values. Without running that job, the online store continues to serve the previous values.

The TTL affects validity, not propagation, so the missing sync job is the most likely cause.

Exam trap

The trap here is assuming that updating the offline store automatically propagates to the online store, or that TTL controls synchronization rather than validity.

650
MCQhard

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to build a custom model with minimal coding. They also need to integrate the model into an existing application that expects a REST API. Which Google Cloud service should they use to train and deploy the model with the least effort?

A.Vertex AI AutoML Tabular to train a model, then deploy it to a Vertex AI endpoint.
B.AI Platform Training with a custom TensorFlow model, then deploy to AI Platform Prediction.
C.Cloud Functions to call the BigQuery ML model directly via a REST API.
D.BigQuery ML to train a logistic regression model, then export the model to Vertex AI for deployment.
AnswerA

Vertex AI AutoML Tabular automates model training and tuning for tabular data, minimizing coding. Once trained, the model can be deployed to a Vertex AI endpoint, which provides a REST API for real-time predictions. This end-to-end integration requires minimal effort and meets the requirement for a custom model with low-code and REST API access.

Why this answer

Vertex AI AutoML Tabular provides a low-code solution for training custom models on tabular data. It automatically handles feature engineering and model selection, and the trained model can be deployed to a Vertex AI endpoint, which offers a REST API for real-time predictions. This integration minimizes effort and meets the need for a custom model with REST API access.

Other options either require more coding, add unnecessary complexity, or do not support real-time serving natively.

Exam trap

The trap here is assuming that BigQuery ML can directly serve real-time REST API predictions, but it is primarily for batch or SQL-based predictions, not low-latency online serving.

651
MCQhard

A company wants to build a recommendation system that suggests products to users based on their past interactions. They have user-item interaction data in BigQuery and want a low-code solution that can generate recommendations for all users. Which approach should they use?

A.Use BigQuery ML to train a k-means clustering model and assign users to clusters for recommendations.
B.Use Dataflow to preprocess data and train a custom recommendation model on Vertex AI.
C.Use Vertex AI AutoML Tables to train a model that predicts a rating for each user-item pair.
D.Use BigQuery ML to train a matrix factorization model and use ML.RECOMMEND to generate recommendations.
AnswerD

BigQuery ML's matrix factorization model is specifically designed for recommendation tasks. It can be trained directly on user-item interaction data using SQL, which is low-code. The ML.RECOMMEND function then generates top-N recommendations for all users efficiently. This approach meets the requirement for a low-code solution that scales to all users, making it the ideal choice for this scenario.

Why this answer

BigQuery ML's matrix factorization model is purpose-built for recommendations and can be trained with SQL on user-item interaction data. The ML.RECOMMEND function generates recommendations for all users in a scalable way. This low-code approach avoids custom coding and is optimized for the task, making it the best fit for the company's needs.

Exam trap

The trap here is assuming that any BigQuery ML model can generate recommendations, when only matrix factorization and a few other specialized types support ML.RECOMMEND.

652
Drag & Dropmedium

Drag and drop the steps to set up a BigQuery ML linear regression model for forecasting in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

For BigQuery ML linear regression, the correct order is: 1) Prepare training data by selecting and preprocessing features in a SQL query; 2) Create the model using `CREATE MODEL` with the training data; 3) Evaluate the model using `ML.EVALUATE` to check metrics like R²; 4) Use the model for predictions with `ML.PREDICT`. This sequence ensures the model is built on clean data, validated, and then applied.

653
Multi-Selectmedium

A company wants to implement a central model governance strategy using Vertex AI. They need to track model lineage, store evaluation metrics, and manage model versions across teams. Which THREE Vertex AI services should they use? (Choose 3)

Select 3 answers
A.Vertex AI Metadata
B.Vertex AI Model Registry
C.Vertex AI Workbench
D.Vertex AI Experiments
E.Vertex AI Feature Store
AnswersA, B, D

Vertex AI Metadata stores artefacts, executions and contexts in a managed ML metadata store, recording lineage as models train and deploy. It satisfies the lineage-tracking constraint directly, letting teams trace which datasets and training runs produced each model version. Evaluation metrics attach as metadata to those artefacts, supporting governance across teams.

Why this answer

Vertex AI Metadata (A) is correct because it provides the managed ML metadata store that automatically tracks artifacts, executions, and contexts, enabling model lineage tracking across teams. Vertex AI Model Registry (B) is correct because it is the central repository for managing model versions, including versioning, aliasing, and deployment tracking across teams. Vertex AI Experiments (D) is correct because it records and compares experiment runs, storing evaluation metrics and parameters so teams can track model performance over time.

Vertex AI Workbench (C) is not correct because it is an interactive notebook development environment, not a governance or lineage tracking service. Vertex AI Feature Store (E) is not correct because it manages feature storage and serving for training and prediction, not model lineage, metrics, or version management.

Exam trap

PMLE often tests which Vertex AI service does what; the trap is confusing Feature Store or Workbench with governance services, when the question specifically asks for lineage, versioning, and metrics.

654
MCQmedium

A media company wants to automatically transcribe and analyze customer support calls to identify common issues. They need a low-code solution that provides both transcription and sentiment analysis. Which Google Cloud service should they use?

A.Speech-to-Text API
B.Contact Center AI Insights
C.Dialogflow CX
D.Natural Language API
AnswerB

Contact Center AI Insights is designed to analyze customer interactions, providing transcription, sentiment analysis, and topic detection out of the box. It is a low-code solution that integrates with various contact center platforms. It can automatically process calls and surface insights without custom model development. This directly meets the requirement for both transcription and sentiment analysis in a low-code manner.

Why this answer

Contact Center AI Insights is a purpose-built, low-code solution for analyzing customer calls. It automatically transcribes audio, performs sentiment analysis, and extracts topics, providing actionable insights without custom development. Other options either lack sentiment analysis or require integration with multiple services, increasing complexity and coding effort.

Thus, Contact Center AI Insights is the correct choice.

Exam trap

The trap here is assuming that combining Speech-to-Text and Natural Language API is low-code; it actually requires integration work, whereas Contact Center AI Insights provides an all-in-one managed solution.

655
Multi-Selecthard

Which THREE factors are critical when designing a model serving architecture for a global user base with strict latency SLAs? (Choose 3.)

Select 3 answers
A.Use batch prediction to process requests in bulk for efficiency.
B.Deploy the model in a single region to avoid data sovereignty issues.
C.Enable autoscaling with request-based metrics to handle traffic spikes.
D.Implement request caching for idempotent predictions when appropriate.
E.Use multi-region deployment with Vertex AI Endpoints in multiple locations.
AnswersC, D, E

Request-based autoscaling adjusts serving replicas dynamically as traffic spikes, preventing queueing latency that would breach strict SLAs during demand surges. Scaling on CPU alone reacts too slowly for request-driven load, so request metrics directly protect the latency target.

Why this answer

Options C, D, and E are correct. Option A is wrong because batch prediction is designed for offline processing of large volumes, not for real-time serving with strict latency SLAs. Option B is wrong because a single-region deployment cannot provide low latency to a global user base and does not address data sovereignty issues appropriately; multi-region deployment with fine-grained controls is needed.

Option C is correct: autoscaling with request-based metrics ensures that resources scale dynamically to handle traffic spikes while maintaining latency. Option D is correct: caching idempotent predictions reduces latency for repeated queries and offloads the model server. Option E is correct: multi-region deployment with Vertex AI Endpoints in multiple locations minimizes network distance and latency for users worldwide.

656
MCQeasy

A data scientist has trained a scikit-learn model locally and wants to deploy it to Vertex AI for online predictions with low latency. The model is a small RandomForestClassifier (100 MB). What is the recommended way to deploy this model?

A.Deploy the model on a Kubernetes cluster with Istio.
B.Package the model as a Docker container with a custom prediction routine.
C.Upload the model to Vertex AI Model Registry using the pre-built scikit-learn serving container.
D.Export the model as a TensorFlow SavedModel and use the pre-built TF serving container.
AnswerC

Vertex AI's pre-built scikit-learn container already implements the prediction server and model-loading contract, so uploading the 100 MB artefact to the Model Registry avoids writing custom serving code and satisfies the low-latency online prediction requirement.

Why this answer

Vertex AI provides a pre-built container for scikit-learn that is optimized for serving predictions with low latency. For a small RandomForestClassifier (100 MB), this container handles model loading, request routing, and scaling automatically, eliminating the need for custom infrastructure. This is the recommended approach for deploying scikit-learn models to Vertex AI for online predictions.

Exam trap

Google Cloud often tests the misconception that any model must be containerized or converted to TensorFlow for deployment, but the correct answer leverages the platform's pre-built container for the specific framework, which is the simplest and most efficient path for small models.

How to eliminate wrong answers

Option A is wrong because deploying on a Kubernetes cluster with Istio adds unnecessary operational complexity and overhead for a small model that can be served directly via Vertex AI's managed infrastructure; it is not the recommended path for a simple scikit-learn model. Option B is wrong because packaging the model as a Docker container with a custom prediction routine is overkill when Vertex AI already offers a pre-built, optimized scikit-learn serving container that handles the prediction logic out of the box. Option D is wrong because exporting a scikit-learn model as a TensorFlow SavedModel is not a direct conversion; scikit-learn models are not natively compatible with TensorFlow Serving, and this would require significant re-engineering or use of ONNX, which is not the recommended path for a RandomForestClassifier.

657
MCQeasy

Refer to the exhibit. A data scientist notices that predictions from a deployed model are taking longer than expected. Which Cloud Monitoring metric should be inspected first to identify the bottleneck?

A.Vertex AI - Model - Compute utilization
B.Vertex AI - Endpoint - Prediction latency distribution
C.Vertex AI - Endpoint - Traffic
D.Vertex AI - Endpoint - Online prediction errors
AnswerB

Prediction latency distribution reports the spread of response times for online predictions at the endpoint, isolating whether slow inference is the bottleneck. It satisfies the stem's requirement by measuring the deployed model's serving latency directly, rather than CPU, memory or request-count metrics.

Why this answer

The data scientist is investigating slow predictions from a deployed model. The most direct metric to identify the latency bottleneck is the prediction latency distribution, which shows the distribution of response times for online prediction requests. This metric allows you to pinpoint whether the delay is due to model inference time, network overhead, or endpoint queuing, making it the first logical place to inspect.

Exam trap

Google Cloud often tests the distinction between metrics that measure performance (latency) versus metrics that measure capacity (utilization, traffic) or errors, leading candidates to mistakenly choose compute utilization or traffic when the question explicitly asks about prediction time.

How to eliminate wrong answers

Option A is wrong because Vertex AI - Model - Compute utilization measures the resource usage (CPU/memory) of the model's compute resources, which can indicate a resource bottleneck but does not directly show prediction latency; it is a secondary metric to investigate after latency is confirmed. Option C is wrong because Vertex AI - Endpoint - Traffic measures the number of requests per second (RPS) to the endpoint, which can indicate load but does not directly measure how long each prediction takes; high traffic can cause latency, but the metric itself is not a latency metric. Option D is wrong because Vertex AI - Endpoint - Online prediction errors tracks the count or rate of failed predictions (e.g., timeouts, invalid inputs), not the latency of successful predictions; errors may be a consequence of latency but are not the primary metric for identifying a latency bottleneck.

658
MCQhard

A team uses Vertex AI Metadata to track pipeline runs. They need to identify all artifacts that were generated by a particular pipeline execution. Which API method should they use?

A.List executions and then list artifacts separately
B.Use the lineage query API with the execution ID
C.Create a context and query executions
D.Query artifacts by filter on execution ID
AnswerB

The lineage query API accepts an execution ID and returns the subgraph of artifacts produced by that execution, directly answering which artifacts a given pipeline run generated. Listing executions alone returns run metadata without the produced artifact relationships.

Why this answer

The Vertex AI Metadata lineage query API accepts an execution ID and returns all artifacts, contexts, and events connected to that execution, giving the full provenance graph in one call. This is the purpose-built method for tracing which artifacts a pipeline run produced.

Exam trap

PMLE often tests the assumption that artifacts can be filtered directly by execution ID — candidates pick the 'filter artifacts' option because it sounds precise, missing that Vertex AI Metadata expresses execution-artifact relationships through lineage events, not a direct foreign key.

How to eliminate wrong answers

Option A is wrong because listing executions and artifacts separately returns flat, unlinked lists — you would have to manually correlate them, and there is no guarantee of a clean join without lineage data. Option C is wrong because creating a context and querying executions inverts the relationship — contexts group related executions, but the question asks for artifacts generated by a specific execution, not executions within a context. Option D is wrong because artifacts do not carry a direct 'execution ID' filter field in the Metadata API; the linkage is expressed through lineage events, so filtering artifacts by execution ID is not a supported query pattern.

659
MCQeasy

A company deploys a TensorFlow model on Vertex AI Prediction with a single node. During peak hours, inference latency increases. What should they do first to reduce latency?

A.Enable autoscaling for the deployment
B.Increase the machine type of the node
C.Decrease the min replicas to 0
D.Enable automatic batching of requests
AnswerA

Autoscaling adds replicas when traffic rises, spreading inference load across nodes so each handles fewer requests. This directly addresses the single-node bottleneck causing peak-hour latency, and it is the least disruptive first step before considering larger machines or GPUs.

Why this answer

Enabling autoscaling for the deployment is the correct first step because it allows Vertex AI Prediction to dynamically adjust the number of replicas based on incoming traffic. During peak hours, autoscaling can add more nodes to distribute the inference load, directly reducing latency without requiring manual intervention or over-provisioning.

Exam trap

The trap here is that candidates often confuse improving throughput (batching or bigger machines) with reducing latency under load, but the first action should always be to add more replicas via autoscaling to handle concurrent requests, not to optimize a single node's performance.

How to eliminate wrong answers

Option B is wrong because increasing the machine type of the node (e.g., moving to a larger VM) may improve per-node throughput but does not address the root cause of insufficient capacity during traffic spikes; it also increases cost without guaranteeing latency reduction if the single node is already saturated. Option C is wrong because decreasing the min replicas to 0 would cause the deployment to scale down to zero during idle periods, but during peak hours it would still need to scale up from zero, causing cold-start latency and potentially failing to handle the initial burst of requests. Option D is wrong because enabling automatic batching of requests can improve throughput by grouping multiple inference requests into a single batch, but it does not reduce latency for individual requests—in fact, it may increase latency as requests wait for a batch to fill.

660
MCQeasy

An ML engineer needs to monitor a deployed model for data drift. They want to compare the distribution of incoming predictions against a baseline distribution. Which Vertex AI service should they use?

A.Vertex AI Feature Store
B.Vertex AI Model Monitoring
C.Vertex AI Experiments
D.Vertex AI Explainable AI
AnswerB

Vertex AI Model Monitoring computes drift metrics by comparing incoming prediction request distributions against a baseline dataset, covering both feature and prediction drift. It directly satisfies the requirement to detect distribution shift on a deployed model.

Why this answer

Vertex AI Model Monitoring is the correct service because it is specifically designed to detect data drift and feature skew in deployed models. It continuously compares the distribution of incoming prediction requests against a baseline distribution (e.g., training data or a previous window) and alerts the engineer when statistically significant drift is detected, using metrics like Jensen-Shannon divergence or L-infinity distance.

Exam trap

Google Cloud often tests the distinction between monitoring (drift detection) and other MLOps components like feature stores or experiment tracking, so the trap here is that candidates may confuse 'monitoring' with 'storing features' or 'tracking experiments' because all are part of the ML lifecycle but serve different purposes.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store is a centralized repository for storing, managing, and serving feature values for training and serving, not for monitoring distributional shifts in predictions. Option C is wrong because Vertex AI Experiments is used for tracking and comparing machine learning experiments (e.g., hyperparameter tuning runs), not for real-time monitoring of deployed model predictions. Option D is wrong because Vertex AI Explainable AI provides feature attributions and explanations for model predictions, but does not perform statistical drift detection or baseline comparison.

661
MCQhard

A company runs a Vertex AI Pipeline that includes a hyperparameter tuning step followed by a training step. The tuning step outputs the best hyperparameters. The engineer wants the training step to use these hyperparameters and to ensure that the training step only runs if tuning succeeds. Which approach should the engineer take?

A.Use a conditional expression in the training step to check if the tuning step succeeded, and only then proceed.
B.Configure the tuning step to write hyperparameters to a Cloud Storage bucket, and have the training step read from that bucket at runtime.
C.Run the tuning and training steps in parallel to save time, and use a shared memory store to exchange hyperparameters.
D.Pass the output of the tuning step as an input parameter to the training step and set the `after` parameter of the training step to include the tuning step.
AnswerD

In Vertex AI Pipelines, data dependencies are established by passing outputs as inputs. The tuning step's output (best hyperparameters) can be passed as an input to the training step. Additionally, the `after` parameter ensures the training step runs only after the tuning step completes successfully. This combination enforces both data flow and execution order.

Why this answer

To pass data between tasks in Vertex AI Pipelines, the output of one task must be provided as an input to another. The tuning step's output (best hyperparameters) is passed as an input to the training step, creating a data dependency. Additionally, the `after` parameter ensures the training step runs only after the tuning step succeeds.

This dual approach guarantees both correct data flow and execution order, and if tuning fails, the pipeline stops before training.

Exam trap

The trap here is relying on external storage or parallel execution to share data between tasks, which does not enforce execution order or failure propagation; Vertex AI Pipelines requires explicit input/output connections and the `after` parameter for dependencies.

662
MCQmedium

Your team trains models in a shared Vertex AI project, and multiple engineers run pipelines against the same BigQuery training tables. A reviewer needs to reproduce the exact dataset used to train a model six weeks ago, but the source tables have been overwritten many times since. Which BigQuery capability should you have used to make each training snapshot reproducible?

A.BigQuery streaming inserts into a dedicated audit table
B.BigQuery time travel queries against the live table
C.BigQuery materialized views over the training tables
D.BigQuery table snapshots created before each training run
AnswerD

Table snapshots capture the table state at a point in time and are retained as independent, read-only copies that keep working even after the source table is overwritten. Taking a snapshot before each training run gives the reviewer a stable, addressable dataset to reproduce the exact training input without duplicating full storage cost.

Why this answer

Reproducing a past training run requires an immutable copy of the data as it existed then. Table snapshots provide exactly that: a point-in-time, read-only reference retained independently of later writes, so the reviewer can rebuild the training dataset even after many overwrites. Time travel, materialized views, and streaming audit tables all reflect or lose the historical state.

Exam trap

The trap here is assuming BigQuery time travel retains history indefinitely, when its window is limited and expires long before many reproducibility requirements.

663
MCQeasy

You are fine-tuning a BERT model from Hugging Face Transformers on Vertex AI. You want to minimise cost for a short experiment. Which compute configuration should you use?

A.A custom training job with a single NVIDIA T4 GPU using spot VMs
B.A custom training job with a TPU v3-8 pod
C.A custom training job with 8 NVIDIA V100 GPUs using regular VMs
D.A standard n1-highmem-8 machine with no accelerator
AnswerA

A single T4 GPU on spot VMs gives the lowest cost for a short fine-tuning experiment. Spot capacity suits interruptible, brief jobs, and T4 provides sufficient memory and compute for BERT-scale training without paying for premium accelerators.

Why this answer

A single NVIDIA T4 GPU with spot VMs is the most cost-effective choice for a short BERT fine-tuning experiment on Vertex AI. T4 GPUs are inexpensive and well-suited for moderate training workloads, and spot VMs offer up to 60-70% discount over regular VMs. Since the experiment is short, the risk of preemption is acceptable, and the cost savings are significant.

This configuration balances performance and cost effectively.

Exam trap

PMLE often tests the misconception that more powerful hardware (TPUs or multiple high-end GPUs) is always better, but the key is matching compute to workload scale and cost constraints; candidates may overlook spot VMs as a cost-saving option for short, fault-tolerant jobs.

How to eliminate wrong answers

Option B is wrong because a TPU v3-8 pod is significantly more expensive and overpowered for a short BERT fine-tuning experiment; TPUs are optimized for large-scale training and require specific code adjustments, adding complexity and cost. Option C is wrong because using 8 NVIDIA V100 GPUs with regular VMs incurs very high costs and is unnecessary for a short experiment; V100s are powerful but overkill, and regular VMs lack the cost savings of spot instances. Option D is wrong because a standard n1-highmem-8 machine without any accelerator would be extremely slow for BERT fine-tuning, as BERT training benefits greatly from GPU acceleration; using only CPU would prolong training time and increase overall cost despite lower hourly rates.

664
MCQhard

You are training a TensorFlow model on Vertex AI using a custom container. The training job uses a single node with 4 GPUs and a global batch size of 1024. You notice that the training is slower than expected and GPU utilization is low. You suspect the input pipeline is the bottleneck. Which of the following should you do to improve training throughput?

A.Reduce the global batch size to 512 to decrease memory pressure on the GPUs.
B.Use tf.data to prefetch data and parallelize data extraction and transformation with num_parallel_calls and prefetch.
C.Increase the number of GPUs to 8 and double the global batch size to 2048.
D.Switch to using TPUs instead of GPUs for the training job.
AnswerB

Optimizing the input pipeline with tf.data by using prefetch and parallel calls allows data to be prepared while the GPU is training. This overlaps I/O and preprocessing with computation, increasing GPU utilization and throughput. It directly addresses the bottleneck by ensuring the GPU is not waiting for data, which is a common cause of low utilization in multi-GPU training.

Why this answer

Low GPU utilization in multi-GPU training often indicates the input pipeline cannot supply data fast enough. Using tf.data with prefetch and parallel calls overlaps data preprocessing with model training, keeping the GPUs busy. This is a standard optimization for TensorFlow input pipelines and directly targets the bottleneck without changing hardware or batch size.

Exam trap

The trap here is assuming that adding more hardware or changing batch size will solve low GPU utilization, when the real issue is data starvation from an inefficient input pipeline.

665
MCQmedium

Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?

A.The 99th percentile of the `prediction/online/response_latencies` metric.
B.The number of prediction requests that timeout.
C.Average prediction latency from the endpoint's logs.
D.The maximum prediction latency from the endpoint's monitoring dashboard.
AnswerA

The prediction/online/response_latencies metric records server-side latency for online predictions, and its 99th percentile aligns exactly with the stem's 500ms SLO threshold. Alerting on that percentile detects tail latency affecting the slowest 1% of requests, which averages or medians would obscure.

Why this answer

The `prediction/online/response_latencies` metric in Vertex AI provides a distribution of latency values, allowing you to query the 99th percentile directly. This aligns with the SLO requirement to monitor the tail latency, not the average or maximum, ensuring that the worst-case performance for 1% of requests stays under 500ms.

Exam trap

Google Cloud often tests the distinction between tail latency (percentiles) and central tendency (average) or extreme values (maximum), trapping candidates who confuse SLO monitoring with simple failure counts or averages.

How to eliminate wrong answers

Option B is wrong because the number of prediction requests that timeout is a count of failures, not a latency measurement; it does not capture the 99th percentile latency and would miss requests that complete but exceed 500ms. Option C is wrong because average prediction latency can mask high tail latencies; a low average could hide a significant number of requests exceeding 500ms, violating the SLO. Option D is wrong because the maximum prediction latency is a single extreme value, often an outlier due to cold starts or transient spikes, and does not represent the 99th percentile behavior required for the SLO.

666
MCQhard

A team is building a CI/CD pipeline for ML using Cloud Build. The pipeline trains a model and deploys it to Vertex AI. Recently, a change in the data processing step caused the model to be trained with a different data version, leading to a failed deployment because the model was invalid. How should the team prevent this in the future?

A.Add a manual review step before training
B.Pin all library versions in the Docker image
C.Use a data versioning tool (e.g., DVC) to track datasets and ensure the pipeline always uses the correct version
D.Schedule a cron job to check for data changes
AnswerC

DVC pins each dataset to an immutable version, so the pipeline references an explicit data revision rather than whatever the processing step last produced. This guarantees training and deployment use the same version, preventing the invalid-model failure.

Why this answer

The root cause is a data version mismatch, not a code or environment issue. A data versioning tool like DVC (Data Version Control) tracks dataset versions via hash-based pointers in Git, ensuring the pipeline retrieves the exact dataset version used during training. This prevents silent failures when data processing steps change the data schema or content, which library pinning or manual reviews cannot guarantee.

Exam trap

The trap here is that candidates confuse environment reproducibility (pinning libraries) with data reproducibility, assuming that locking code dependencies is sufficient to prevent model failures caused by data drift or version changes.

How to eliminate wrong answers

Option A is wrong because a manual review step before training introduces human latency and does not enforce data version consistency; it relies on a person to catch a version mismatch that may not be visually obvious. Option B is wrong because pinning library versions in the Docker image addresses dependency drift in code, not data versioning; the model failed due to a different data version, not a library incompatibility. Option D is wrong because scheduling a cron job to check for data changes is reactive and does not prevent the pipeline from using the wrong data version; it only alerts after the fact, and the pipeline would still train on incorrect data.

667
MCQhard

A team is monitoring a production ML system that includes multiple models and data processing pipelines. They want to set up a comprehensive alerting strategy that minimizes false positives while ensuring critical issues are promptly addressed. Which approach is the most effective?

A.Set up alerts for all possible error conditions
B.Use static thresholds based on historical data
C.Rely on manual monitoring during business hours
D.Use AIOps with anomaly detection to dynamically adjust thresholds
AnswerD

AIOps anomaly detection learns normal metric behaviour per model and pipeline, dynamically adjusting thresholds rather than relying on static values, which reduces false positives while still surfacing genuine deviations promptly across the many monitored signals.

Why this answer

AIOps with anomaly detection uses machine learning to dynamically adjust alert thresholds based on real-time system behavior, reducing false positives while ensuring critical issues are detected promptly. This approach adapts to changing data distributions and traffic patterns, unlike static thresholds that require manual tuning and often miss subtle anomalies. It is the most effective strategy for complex ML production systems where multiple models and pipelines interact, as it can correlate signals across components to identify genuine incidents.

Exam trap

The trap here is that candidates often choose static thresholds (Option B) because they seem simpler and more predictable, but they fail to recognize that production ML systems require adaptive thresholds to handle dynamic data distributions and avoid alert fatigue.

How to eliminate wrong answers

Option A is wrong because setting alerts for all possible error conditions leads to alert fatigue, overwhelming the team with noise and causing critical issues to be missed; it lacks prioritization and ignores the need for intelligent filtering. Option B is wrong because static thresholds based on historical data fail to adapt to concept drift, seasonal patterns, or sudden traffic spikes, resulting in either too many false positives or missed anomalies when the system behavior changes. Option C is wrong because relying on manual monitoring during business hours introduces unacceptable latency for critical issues that occur outside those hours, and human error or fatigue can cause delays in detection; it is not scalable for 24/7 production ML systems.

668
MCQhard

Your team uses Vertex AI Pipelines to automate the training and deployment of a recommendation model. The pipeline includes a step that evaluates the model and only deploys it if the evaluation metric exceeds a threshold. You need to ensure that the pipeline's artifacts, including the evaluation metrics and the deployed model, are tracked and can be traced back to the pipeline run for auditing. What should you do?

A.Configure the pipeline to send an email with the evaluation metrics and model details to a distribution list for record-keeping.
B.Enable Cloud Logging for the pipeline and rely on the logs to capture the evaluation metrics and model deployment events.
C.Store the evaluation metrics in a BigQuery table and the model in Cloud Storage, and record the pipeline run ID in both locations.
D.Use Vertex AI ML Metadata to log the evaluation metrics and model artifacts, and associate them with the pipeline run.
AnswerD

Vertex AI ML Metadata automatically tracks artifacts, executions, and contexts for Vertex AI Pipelines runs. By logging metrics and model artifacts, you create a lineage that links them to the pipeline run. This enables auditing and reproducibility, as you can trace which run produced which model and its metrics.

Why this answer

Vertex AI ML Metadata provides automatic lineage tracking for pipeline artifacts, including metrics and models. It records relationships between executions and artifacts, enabling auditing and reproducibility. Other methods lack the structured, integrated metadata store that ML Metadata offers, making them less reliable for tracing and compliance.

Exam trap

The trap here is assuming that manual logging or Cloud Logging can substitute for ML Metadata's automatic lineage tracking, when only ML Metadata provides a queryable artifact graph.

669
MCQeasy

A machine learning engineer is exporting a trained model from Vertex AI Training to the Model Registry. Which artifact should they upload as the model artifact?

A.The saved model directory containing the model file(s) and any custom dependencies.
B.Only the model checkpoint file (.ckpt or .h5).
C.The entire training directory including training code and logs.
D.A zip file of the training source code.
AnswerA

Vertex AI Model Registry expects a SavedModel directory holding the serialised graph, variables and assets, so custom dependencies travel with it. Uploading only weights or a checkpoint omits the serving signature, preventing deployment to an Endpoint.

Why this answer

When exporting a trained model from Vertex AI Training to the Model Registry, the correct artifact is the saved model directory that contains the model file(s) (e.g., SavedModel format for TensorFlow, model.pkl for scikit-learn) along with any custom dependencies required for serving. This ensures the model can be deployed consistently to endpoints or batch predictions, as the Model Registry expects a self-contained artifact that includes both the model binary and its runtime dependencies.

Exam trap

Google Cloud often tests the distinction between training artifacts (checkpoints, code) and deployable model artifacts, trapping candidates who confuse a checkpoint (used for resuming training) with a final, serving-ready model.

How to eliminate wrong answers

Option B is wrong because a model checkpoint file (.ckpt or .h5) is an intermediate training state, not a final deployable artifact; it lacks the serialized graph and serving signatures needed for inference. Option C is wrong because uploading the entire training directory, including training code and logs, introduces unnecessary files and violates the Model Registry's expectation of a minimal, serving-ready artifact. Option D is wrong because a zip file of the training source code contains no model weights or architecture, making it useless for deployment.

670
MCQmedium

A team wants to collect ground truth labels for their model deployed on Vertex AI Endpoint to perform model quality monitoring. They have a process that generates actual outcomes within 24 hours of prediction. What is the recommended approach for storing these labels?

A.Upload the ground truth labels to a BigQuery table with a schema that includes prediction timestamp and model version.
B.Use Vertex AI Experiments to log ground truth alongside training runs.
C.Store the ground truth labels in Cloud Storage as CSV files and reference them in the monitoring config.
D.Insert ground truth labels directly into the Vertex AI Endpoint's log sink.
AnswerA

BigQuery stores ground truth with prediction timestamp and model version columns, letting Vertex AI join labels to logged predictions for skew and drift analysis. The 24-hour outcome delay fits batch upload, and the schema keys each label to the exact prediction it evaluates.

Why this answer

Vertex AI Model Monitoring expects ground truth labels to be stored in a BigQuery table whose schema includes the prediction timestamp and model version (plus the prediction output and the actual label). This lets the service join ground truth back to logged predictions within the monitoring window and compute quality metrics like accuracy, precision, and recall. BigQuery is the only supported sink for ground truth in Vertex AI model quality monitoring.

Exam trap

PMLE often tests the assumption that any storage (GCS, Experiments) can hold ground truth — the trap is forgetting that Vertex AI model quality monitoring only reads ground truth from BigQuery with a specific schema.

How to eliminate wrong answers

Option B is wrong because Vertex AI Experiments tracks training-run metadata (parameters, metrics, artifacts) and has no join path to live endpoint predictions for monitoring. Option C is wrong because Cloud Storage CSV files are not a supported ground truth source for Vertex AI model monitoring — the service reads from BigQuery. Option D is wrong because there is no mechanism to insert labels into an endpoint's log sink; the endpoint logs predictions to BigQuery, and ground truth is a separate table joined on prediction ID/timestamp.

671
MCQmedium

A retail company wants to predict customer churn using historical purchase data stored in BigQuery. The data includes customer demographics, transaction history, and support interactions. The team is comfortable writing SQL and wants to avoid moving data to a separate environment. Which approach should they take?

A.Use the Cloud Natural Language API to analyze customer support interactions and combine results with purchase data in BigQuery.
B.Export the data to a CSV file and use Vertex AI AutoML Tables to train a classification model.
C.Use BigQuery ML to create a logistic regression model (LOGISTIC_REG) on the data directly in BigQuery.
D.Create a Dataflow pipeline to stream data to Cloud SQL and use Cloud SQL's built-in ML functions.
AnswerC

BigQuery ML trains LOGISTIC_REG models using SQL directly against BigQuery-resident data, so no extraction or separate environment is needed. This satisfies the team's SQL comfort and the constraint of avoiding data movement for churn prediction.

Why this answer

BigQuery ML allows the team to build and train a logistic regression model directly on data stored in BigQuery using SQL syntax, without moving data to a separate environment. The LOGISTIC_REG model type is specifically designed for binary classification tasks like churn prediction, and it runs entirely within BigQuery's serverless infrastructure, satisfying the team's requirement to avoid data movement.

Exam trap

This question tests the misconception that ML requires moving data to a separate platform (like Vertex AI or Cloud SQL), when in fact BigQuery ML provides a low-code, SQL-based solution that keeps data in place and meets the stated constraints.

How to eliminate wrong answers

Option A is wrong because the Cloud Natural Language API is used for text analysis (e.g., sentiment extraction), not for training a predictive churn model; it would require additional steps to combine results and does not provide a built-in classification model. Option B is wrong because exporting data to a CSV file and using Vertex AI AutoML Tables violates the requirement to avoid moving data to a separate environment, and it introduces unnecessary data egress and manual steps. Option D is wrong because Cloud SQL does not have built-in ML functions for training classification models; it is a relational database service, and streaming data through Dataflow to Cloud SQL adds complexity and does not leverage BigQuery's native ML capabilities.

672
Drag & Dropmedium

Drag and drop the steps to set up data lineage tracking for ML pipelines using Vertex AI Experiments in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Start with SDK setup, then create an experiment, log metrics, record artifacts, and review lineage.

673
Multi-Selectmedium

Which THREE are key capabilities of Vertex AI Feature Store?

Select 3 answers
A.Automatic generation of feature embeddings
B.Feature monitoring and validation to detect skew
C.Online serving for low-latency feature retrieval
D.Real-time streaming ingestion from Apache Kafka
E.Offline batch serving for training
AnswersB, C, E

Feature monitoring and validation directly satisfies the requirement to detect skew between training and serving data. Vertex AI Feature Store continuously tracks feature distributions, alerting when drift or training-serving skew emerges, ensuring models consume consistent inputs. This capability is intrinsic to the managed feature store, distinguishing it from generic storage.

Why this answer

Option B is correct because Vertex AI Feature Store provides feature monitoring and validation capabilities that detect training-serving skew and drift by comparing statistics of ingested feature values against baselines. Option C is correct because Feature Store offers online serving, exposing features through a low-latency endpoint so models can retrieve the latest feature values at prediction time. Option E is correct because Feature Store supports offline batch serving, allowing point-in-time-correct feature retrieval for training datasets via BigQuery or batch export.

Option A is not a core capability of Feature Store; embeddings are generated by models or Vertex AI Embeddings APIs, not by the feature store itself. Option D is not a built-in capability; Feature Store ingests data through its API or BigQuery, and Kafka would require a custom pipeline rather than native streaming ingestion.

Exam trap

Google Cloud often tests the misconception that Vertex AI Feature Store includes automatic embedding generation or direct Kafka integration, when in fact these are separate services or require custom implementation.

674
MCQeasy

A data engineer wants to compute feature aggregates over a large dataset stored in BigQuery and write the results to Vertex AI Feature Store. The pipeline must handle both batch and streaming data. Which Google Cloud service should they use?

A.BigQuery scheduled queries
B.Cloud Functions triggered by Pub/Sub
C.Cloud Dataproc with Spark
D.Cloud Dataflow with Apache Beam
AnswerD

Cloud Dataflow with Apache Beam provides unified batch and streaming pipelines, reading from BigQuery and writing aggregates to Vertex AI Feature Store. This satisfies the stem's requirement to handle both data modes within one managed service.

Why this answer

Cloud Dataflow with Apache Beam is a unified stream and batch data processing service. It can read from BigQuery, compute aggregates, and write to Vertex AI Feature Store, handling both batch and streaming data with the same pipeline code. This makes it the ideal choice for the requirement.

Exam trap

PMLE often tests the choice between batch-only and unified processing services, and candidates may pick BigQuery scheduled queries or Dataproc for streaming, missing Dataflow's unified capability.

How to eliminate wrong answers

Option A is wrong because BigQuery scheduled queries only handle batch processing and cannot process streaming data. Option B is wrong because Cloud Functions triggered by Pub/Sub is for lightweight event-driven processing, not large-scale feature aggregation. Option C is wrong because Cloud Dataproc with Spark is primarily for batch processing and does not natively handle streaming as seamlessly as Dataflow.

675
Multi-Selecthard

A team uses Vertex AI Pipelines for continuous training triggered by model drift. They want to monitor the pipeline execution cost and optimize resource usage. Which THREE metrics should they track? (Choose 3)

Select 3 answers
A.Pipeline execution duration
B.Number of failed pipeline runs
C.Model accuracy on validation set
D.Total GPU hours consumed per pipeline run
E.Cost per pipeline run in Cloud Billing
AnswersA, D, E

Pipeline execution duration exposes wall-clock time each run consumes, revealing whether drift-triggered retraining is becoming slower and therefore costlier. Tracking it against a baseline highlights inefficient steps or resource contention, directly supporting the optimisation goal alongside billing and GPU-hour metrics.

Why this answer

Option A (Pipeline execution duration) is correct because the wall-clock time a Vertex AI Pipeline run takes directly drives the compute resources billed by the underlying services (Vertex AI Training, Dataflow, etc.), so tracking duration is essential for spotting inefficiencies and optimizing resource usage. Option D (Total GPU hours consumed per pipeline run) is correct because GPUs are the most expensive accelerator resource in Vertex AI; measuring GPU hours per run quantifies accelerator consumption and reveals over-provisioning or idle GPU time that can be right-sized. Option E (Cost per pipeline run in Cloud Billing) is correct because Cloud Billing cost data (exported to BigQuery and labeled per pipeline run) gives the actual monetary cost, which is the ground truth for monitoring execution cost and validating optimization efforts.

Option B (Number of failed pipeline runs) is not a cost or resource-usage metric; it measures reliability, so it does not directly address cost monitoring or resource optimization. Option C (Model accuracy on validation set) is a model-quality metric, not an execution-cost or resource-utilization metric, so it is out of scope for this question.

Exam trap

The trap is selecting model-quality metrics (accuracy, failure count) as cost metrics — the exam tests whether you distinguish operational cost/resource metrics from model performance metrics.

Page 8

Page 9 of 11

Page 10

All pages