Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 76–150

775 questions total · 11pages · All types, answers revealed

Page 1

Page 2 of 11

Page 3
76
MCQeasy

Refer to the exhibit. A team runs the command above and sees only two models. They know there is a model 'model-v3' created three days ago. What is the most likely reason it is not listed?

A.The model was created in a different region.
B.The model is in a different project.
C.The model is not deployed to an endpoint.
D.The model's display name contains a hyphen.
E.The model was created by a different user.
AnswerA

A model registered in another region is stored in that region's registry, so listing commands scoped to the current region will not return it. This directly satisfies the stem's constraint: the team queries one region while 'model-v3' resides elsewhere, hence only two models appear.

Why this answer

The `gcloud ai models list` command lists models within a specific region, as Vertex AI models are regional resources. If 'model-v3' was created in a different region, it would not appear in the output unless the `--region` flag is set to that region. This is the most likely reason the model is missing from the list.

Exam trap

Google Cloud often tests the regional scope of Vertex AI resources, trapping candidates who assume model listing is global or project-wide, when in fact it is region-specific and requires the correct `--region` flag.

How to eliminate wrong answers

Option B is wrong because the `gcloud ai models list` command operates within a single project (the current configured project or one specified with `--project`), but the question states the team sees only two models, implying they are in the correct project; a different project would require explicit project specification. Option C is wrong because model listing does not require deployment to an endpoint; Vertex AI lists all models in the project/region regardless of deployment status. Option D is wrong because hyphens in display names are allowed and do not affect listing; the command lists models by their resource name or display name without filtering on special characters.

Option E is wrong because model listing is not user-scoped; all models in the project/region are visible to any user with appropriate permissions, regardless of who created them.

77
Multi-Selectmedium

You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)

Select 2 answers
A.Deploy the model to multiple endpoints and use a load balancer to distribute traffic.
B.Configure autoscaling based on CPU utilization with a target utilization that triggers scale-out early.
C.Use a larger machine type with more CPU and memory for each replica to increase per-replica throughput.
D.Enable request-response logging to capture detailed latency metrics for each prediction.
E.Set a minimum replica count greater than 1 to ensure that the endpoint has warm replicas ready to serve traffic.
AnswersB, E

Autoscaling based on CPU utilization with a low target utilization triggers scale-out earlier, adding replicas before the existing ones become saturated. This proactive scaling helps handle traffic spikes by provisioning additional capacity in advance. It balances cost and performance by scaling out only when needed but doing so early enough to maintain low latency. This is a standard practice for latency-sensitive endpoints.

Why this answer

To handle traffic spikes and ensure low latency, you should maintain a minimum replica count greater than 1 to have warm replicas ready, and configure autoscaling with a low target utilization to trigger scale-out early. These two settings together provide both baseline capacity and proactive scaling, reducing the risk of cold starts and overload during spikes. Other options like logging or larger machines do not directly address dynamic traffic handling.

Exam trap

The trap here is assuming that logging or larger machines are sufficient for handling spikes, when the key is to have warm replicas and proactive autoscaling.

78
MCQmedium

A credit-risk team runs a tabular model on a Vertex AI Endpoint. They configured Vertex AI Model Monitoring with a training dataset and skew detection using the default threshold. After a week, they receive alerts that many features have high training-serving skew, but the model's business metrics (approval rate, default rate) are unchanged. They suspect the alerts are false positives due to a recent change in an upstream data pipeline that shifted feature distributions. What should they do to reduce these false alerts while still monitoring for real skew?

A.Recreate the monitoring job with a new training dataset that reflects the recent pipeline change, and adjust the skew threshold based on observed variance.
B.Disable skew detection and rely only on prediction drift monitoring.
C.Switch from skew detection to outlier detection and set the outlier threshold to a very low value.
D.Increase the monitoring frequency and lower the skew threshold to capture more data points.
AnswerA

Training-serving skew compares live serving inputs to the statistics of the training dataset. If the upstream pipeline legitimately changed feature distributions, the original training baseline is stale and will flag normal data as skewed. Updating the baseline to include recent representative data and setting a threshold that accounts for observed variance restores accurate detection. This aligns monitoring with the current production reality without disabling protection.

Why this answer

Training-serving skew compares live data to the training baseline. When an upstream pipeline changes distributions legitimately, the old baseline becomes invalid and triggers false skew alerts. Updating the monitoring configuration with a representative training dataset and a threshold based on observed variance realigns detection with current production behavior.

Disabling monitoring or making it more sensitive does not solve the underlying baseline mismatch.

Exam trap

The trap here is assuming that any skew alert must be suppressed by disabling monitoring or tightening thresholds, rather than recognizing that the training baseline itself may need to be refreshed after a legitimate pipeline change.

79
MCQeasy

A marketing team wants to analyze customer reviews for sentiment without writing code. Which Google Cloud service should they use?

A.Cloud Dataflow
B.Vertex AI Workbench
C.BigQuery ML
D.Cloud Natural Language API
AnswerD

Cloud Natural Language API provides pre-trained sentiment analysis via REST calls, returning sentiment scores and magnitudes without model training or code. This satisfies the no-code requirement, unlike Vertex AI, which demands custom model development and programming effort.

Why this answer

The Cloud Natural Language API (option D) is the correct choice because it provides pre-trained models for sentiment analysis, entity recognition, and syntax analysis via a simple REST API, requiring no code beyond sending HTTP requests. This aligns perfectly with the requirement to analyze customer reviews for sentiment without writing code, as the API abstracts all ML complexity.

Exam trap

Google Cloud often tests the distinction between services that require coding (like Dataflow or Workbench) versus those that offer pre-built, no-code APIs (like Cloud Natural Language API), leading candidates to mistakenly choose BigQuery ML because it uses SQL, which they perceive as 'low-code' but still requires explicit query writing and model management.

How to eliminate wrong answers

Option A is wrong because Cloud Dataflow is a fully managed stream and batch data processing service based on Apache Beam, requiring users to write code (e.g., Java or Python) to define data pipelines, making it unsuitable for a no-code sentiment analysis task. Option B is wrong because Vertex AI Workbench is a Jupyter-based notebook environment for building and deploying custom ML models, requiring users to write code (e.g., Python) to train or use models, not a no-code solution. Option C is wrong because BigQuery ML allows users to create and execute ML models using SQL queries, but it still requires writing SQL statements and managing model creation, which is not a no-code API for direct sentiment analysis of text.

80
MCQmedium

Two teams are collaborating on a project and want to use a shared Feature Store in Vertex AI. They need to ensure that features are discoverable and that access is controlled. What is the best practice?

A.Export features to CSV files in Cloud Storage and share the bucket
B.Build a custom feature pipeline using Dataflow and store in Cloud SQL
C.Each team stores features in their own BigQuery table and shares the table
D.Use Vertex AI Feature Store and grant appropriate IAM roles to each team
AnswerD

Vertex AI Feature Store provides a centralised registry where features are published and searchable across projects, satisfying the discoverability requirement. Granting granular IAM roles per team enforces least-privilege access to specific feature groups and resources, directly meeting the controlled-access constraint without duplicating feature data.

Why this answer

Vertex AI Feature Store provides a managed service for sharing features with access controls via IAM roles and enables feature discovery through the UI and API. Option A is wrong because CSV files in Cloud Storage lack feature store metadata, versioning, and online serving capabilities. Option B is wrong because building a custom pipeline with Dataflow and storing in Cloud SQL is not a managed feature store solution and does not provide the same discovery or access control features.

Option C is wrong because each team maintaining their own BigQuery table does not offer centralized feature discovery or unified access control; a Feature Store centralizes metadata and permissions.

81
Drag & Dropmedium

Drag and drop the steps to deploy a trained TensorFlow model to Vertex AI Prediction in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Export the model, upload to GCS, register as a model, deploy to endpoint, then test.

82
MCQmedium

An organization uses Vertex AI Pipelines and wants to track the lineage of datasets, models, and metrics across pipeline runs. They need to query upstream and downstream dependencies of an artifact. Which service should they use?

A.Vertex AI Feature Store
B.Vertex AI Experiments
C.Vertex AI Model Registry
D.Vertex AI Metadata
AnswerD

Vertex AI Metadata stores pipeline resources as a lineage graph of executions, artifacts and contexts, so you can query upstream and downstream dependencies of any artifact. Cloud Logging records events but holds no typed lineage relationships between datasets, models and metrics.

Why this answer

Vertex AI Metadata stores the ML metadata graph produced by Vertex AI Pipelines, including artifacts (datasets, models, metrics), executions, and their input/output relationships. It exposes APIs to traverse this graph in both directions, so you can query upstream sources and downstream dependents of any artifact. This is the correct service for cross-run lineage queries.

Exam trap

PMLE often tests the confusion between Experiments (which run produced these metrics) and Metadata (which artifacts depend on which), so candidates choose Experiments when the question asks for dependency traversal.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store serves and monitors feature values, not pipeline artifact lineage. Option B is wrong because Vertex AI Experiments compares training runs and their metrics but does not provide a queryable artifact dependency graph. Option C is wrong because Vertex AI Model Registry manages model versions and deployment stages, not the full dataset-to-metric lineage across pipeline runs.

83
MCQmedium

An ML team uses Vertex AI Pipelines to train and evaluate models. They want to ensure that only models meeting a minimum accuracy threshold are registered in Vertex AI Model Registry. Which approach should they take?

A.Register all models in Model Registry and use a separate Cloud Function triggered by a Model Registry event to delete models that do not meet the threshold.
B.Add a condition in the pipeline that checks the evaluation metric and only executes the ModelUploadOp if the metric meets the threshold.
C.Configure the ModelUploadOp to include the evaluation metric as a label, and rely on downstream consumers to filter models by label.
D.Use a Vertex AI Model Evaluation component to compute metrics, then manually review the metrics and register the model only if it passes.
AnswerB

Vertex AI Pipelines supports conditional execution using the condition parameter on tasks. By adding a condition that evaluates the model's accuracy metric, you can gate the ModelUploadOp so that only models meeting the threshold are uploaded to Model Registry. This enforces the quality gate directly in the pipeline without manual intervention.

Why this answer

The most reliable way to enforce a quality gate is to use a conditional execution in Vertex AI Pipelines that checks the evaluation metric and only runs the ModelUploadOp when the threshold is met. This prevents unqualified models from being registered and automates governance.

Exam trap

The trap here is assuming that post-registration filtering or manual review can enforce quality, when the pipeline itself should prevent registration of subpar models.

84
Multi-Selecthard

A retail company wants to implement a recommendation system using Recommendations AI. They need to generate personalized recommendations for users based on their browsing history and purchase behavior. Which THREE recommendation types are available in Recommendations AI?

Select 3 answers
A.trending-now
B.recommended-for-you
C.others-you-may-like
D.frequently-bought-together
E.most-popular
AnswersB, C, D

Recommended-for-you is a genuine Recommendations AI type that personalises suggestions per user from their browsing history and past behaviour. This directly matches the stem's requirement to generate personalised recommendations based on browsing and purchase activity.

Why this answer

Recommendations AI offers several built-in recommendation types, and recommended-for-you (B) is correct because it produces personalized product suggestions for a user based on that user's own browsing history, purchase behavior, and other interaction events. others-you-may-like (C) is also correct because it generates personalized recommendations related to a specific item the user is currently viewing, which fits the retail scenario of tailoring suggestions from user behavior. frequently-bought-together (D) is correct because it recommends complementary products commonly purchased alongside a given item, a standard Recommendations AI type for cross-sell use cases. The unmarked options do not belong: trending-now (A) and most-popular (E) are not valid Recommendations AI recommendation type identifiers, as the platform's types are named like recommended-for-you, others-you-may-like, frequently-bought-together, and similar-item, rather than those generic popularity labels.

Exam trap

PMLE often tests the distinction between valid recommendation types in Recommendations AI and generic recommendation strategies like 'trending-now' or 'most-popular', which are not available as predefined types.

85
MCQhard

A company has a large-scale ML system that uses Vertex AI Pipelines to retrain models weekly. The pipeline includes a custom training job and a batch prediction step. After moving to production, they observe that batch prediction jobs often fail with 'Quota exceeded' errors. The project has sufficient CPU quota. What is the most likely cause?

A.The pipeline is exceeding the maximum number of concurrent pipeline runs.
B.The batch prediction job is requesting a specific accelerator type that has a separate quota limit.
C.The batch prediction job is using a machine type that is not available in the region.
D.The custom training job is consuming all available quota before the batch prediction job starts.
AnswerB

Accelerator quotas are tracked separately from CPU quota in Google Cloud. A batch prediction job requesting GPUs or TPUs can therefore hit 'Quota exceeded' despite ample CPU quota, making the accelerator request the most likely cause.

Why this answer

The most likely cause is that the batch prediction job is requesting a specific accelerator type (e.g., GPU or TPU) that has a separate quota limit from CPU quota. In Vertex AI, accelerator quotas are distinct from general compute (CPU) quotas, and even if the project has sufficient CPU quota, the accelerator quota may be exhausted, causing 'Quota exceeded' errors.

Exam trap

Google Cloud often tests the misconception that all quota errors are related to CPU or memory, but the trap here is that accelerator types (GPUs/TPUs) have their own independent quota limits that are easily overlooked when CPU quota appears sufficient.

How to eliminate wrong answers

Option A is wrong because exceeding the maximum number of concurrent pipeline runs would result in pipeline submission failures or throttling, not batch prediction job failures with 'Quota exceeded' errors; Vertex AI Pipelines enforces concurrency limits separately. Option C is wrong because if a machine type is not available in the region, the error would be a resource availability error (e.g., 'Machine type not found'), not a quota exceeded error. Option D is wrong because the custom training job and batch prediction job run sequentially within the same pipeline; the training job completes before the batch prediction job starts, so it cannot consume quota during the batch prediction step.

86
Multi-Selectmedium

A company wants to implement a centralized model registry for governance. Which two features should they use? (Choose two.)

Select 2 answers
A.Vertex AI Feature Store
B.Vertex AI Model Registry
C.Vertex AI Experiments
D.Model versioning and aliases
E.Vertex AI Metadata
AnswersB, D

Vertex AI Model Registry is the centralised catalogue that tracks model artefacts, versions and lineage across projects. It satisfies the governance requirement by giving one authoritative location to register, discover and manage models throughout their lifecycle.

Why this answer

Vertex AI Model Registry (B) is the centralized repository for managing the lifecycle of ML models, providing a single source of truth for governance across an organization, which directly matches the requirement for a centralized model registry. Model versioning and aliases (D) are core capabilities of the Model Registry that let teams track multiple model versions, assign meaningful aliases (e.g., 'production', 'staging'), and control which version is deployed, which is essential for governance and reproducibility. Together, B and D provide the registry plus the version/alias controls needed for centralized model governance.

Vertex AI Feature Store (A) manages feature storage and serving, not model registration, so it does not fulfill the model registry requirement. Vertex AI Experiments (C) tracks experiment runs, parameters, and metrics for training, not model governance or registration. Vertex AI Metadata (E) stores metadata about artifacts and executions in a lineage graph, but it is not the centralized model registry itself.

Exam trap

The trap here is confusing Vertex AI's data/experiment tracking services (Feature Store, Experiments, Metadata) with the model governance service (Model Registry) — candidates often pick Metadata because it sounds like the 'registry' layer.

87
MCQhard

You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?

A.Deploy v2 to a new endpoint and update your clients to use the new endpoint.
B.Use Vertex AI Experiments to compare v1 and v2, then redeploy v2 with 100% traffic.
C.Update the traffic split configuration on the endpoint multiple times over the 3 days to gradually increase v2's percentage.
D.Delete v1 from the endpoint so that all traffic automatically goes to v2.
AnswerC

Updating the endpoint's traffic split repeatedly lets you raise v2's percentage incrementally, satisfying the gradual three-day shift while keeping v1 serving the remainder. Vertex AI supports modifying the deployed model traffic split on a live endpoint without redeployment, so risk is bounded at each step.

Why this answer

Vertex AI endpoints support live traffic splitting between deployed models, allowing you to gradually shift traffic from v1 to v2 by updating the traffic split configuration multiple times over the 3-day period. This approach minimizes risk by enabling incremental rollouts and immediate rollback if issues arise, without requiring client-side changes or downtime.

Exam trap

The trap here is that candidates may assume deleting the old model or redeploying with 100% traffic is acceptable, but the question explicitly requires a gradual shift over 3 days to avoid risk, which only incremental traffic split updates can achieve.

How to eliminate wrong answers

Option A is wrong because deploying v2 to a new endpoint and updating clients introduces unnecessary complexity, potential downtime, and defeats the purpose of gradual traffic shifting; it also requires client-side changes, which is riskier and not aligned with the goal of avoiding risk. Option B is wrong because Vertex AI Experiments are used for offline model evaluation and comparison, not for live traffic management; redeploying v2 with 100% traffic would be an abrupt switch, not a gradual shift over 3 days. Option D is wrong because deleting v1 from the endpoint would immediately route 100% of traffic to v2, which is an abrupt change, not gradual, and violates the requirement to shift traffic gradually over 3 days to avoid risk.

88
MCQeasy

A data scientist wants to automate the retraining of a model when new data arrives in Cloud Storage. Which Google Cloud service is most appropriate for orchestrating this workflow?

A.Cloud Run
B.Vertex AI Predictions
C.Cloud Scheduler
D.Cloud Composer
E.Cloud Functions
AnswerD

Cloud Composer is a managed Apache Airflow service that orchestrates multi-step workflows, including Cloud Storage event triggers and Vertex AI training jobs. It satisfies the retraining-on-new-data constraint by scheduling and coordinating the pipeline end to end.

Why this answer

Cloud Composer (D) is the most appropriate service for orchestrating a retraining workflow because it is a fully managed workflow orchestration service built on Apache Airflow. It allows you to define a Directed Acyclic Graph (DAG) that triggers model retraining when new data arrives in Cloud Storage, handling dependencies, scheduling, and monitoring across multiple steps such as data validation, training, and deployment.

Exam trap

The trap here is that candidates often confuse event-triggered compute services (like Cloud Functions) with full workflow orchestration, failing to recognize that retraining pipelines require multi-step dependency management, retries, and monitoring that only a dedicated orchestrator like Cloud Composer provides.

How to eliminate wrong answers

Option A (Cloud Run) is wrong because it is a serverless compute platform for running stateless containers, not a workflow orchestrator; it lacks native scheduling and dependency management for multi-step pipelines. Option B (Vertex AI Predictions) is wrong because it is a service for deploying models to serve predictions, not for orchestrating the retraining workflow triggered by new data. Option C (Cloud Scheduler) is wrong because it is a cron job service that triggers single actions at fixed times, not a workflow orchestrator that can handle event-driven triggers, conditional logic, and multi-step dependencies.

Option E (Cloud Functions) is wrong because it is a lightweight, event-driven compute service for single-purpose functions; while it can be triggered by Cloud Storage events, it cannot orchestrate complex multi-step pipelines with retries, branching, or monitoring.

89
MCQhard

A company uses Vertex AI Feature Store with an online store for low-latency serving. They observe high latency during peak hours. The feature values are small (< 1 KB each) and the workload is read-heavy. Which change would most effectively reduce latency?

A.Enable caching on the client side
B.Switch from Bigtable online store to Optimized online store
C.Use a larger machine type for Bigtable
D.Increase the number of Bigtable nodes
AnswerB

The optimized online store is purpose-built for low-latency, high-throughput serving of small feature values, so migrating from the Bigtable online store addresses the peak-hour latency directly. Increasing node count or reducing feature size does not change the underlying serving architecture.

Why this answer

Switching from Bigtable online store to Optimized online store is recommended for read-heavy workloads with small feature values, offering lower latency at high QPS.

90
Multi-Selecthard

A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?

Select 3 answers
A.Using Vertex AI Vizier to optimize the model parallelism strategy
B.Enabling Vertex AI AutoML to automatically distribute the model
C.Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe
D.A custom container with the distributed framework (e.g., PyTorch DDP) installed
E.Setting the --worker-machine-count flag when submitting the job
AnswersC, D, E

Manual implementation of pipeline parallelism is required as Vertex AI does not provide built-in model parallelism.

Why this answer

Pipeline parallelism requires explicit implementation in the training script, such as using `torch.distributed.pipeline.sync.Pipe` in PyTorch, to split the model layers across multiple GPUs. This is necessary when the model does not fit on a single GPU, and Vertex AI does not automatically handle model parallelism—it must be coded by the user.

Exam trap

This question tests the misconception that Vertex AI automatically handles model parallelism (e.g., via AutoML or Vizier), when in reality the user must manually implement it in the training script using frameworks like PyTorch or TensorFlow.

91
MCQhard

A company has an existing TensorFlow model for fraud detection that they want to use for predictions in BigQuery. They want to call the model from SQL queries without moving data out of BigQuery. How should they deploy the model?

A.Import the TensorFlow model directly into BigQuery ML
B.Deploy the model to Vertex AI Prediction and use a remote model in BigQuery ML
C.Export BigQuery data to Cloud Storage and use AI Platform Prediction
D.Use AutoML Tables to retrain the model in BigQuery ML
AnswerA

Why this answer

BigQuery ML (BQML) natively supports importing TensorFlow models directly, allowing you to use them for predictions via SQL without moving data out of BigQuery. This is the simplest and most efficient approach because it eliminates the need for external services or data export, leveraging BQML's built-in `CREATE MODEL` statement with the `OPTIONS(model_type = 'TENSORFLOW')` clause.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing Vertex AI or AI Platform, not realizing that BigQuery ML has native TensorFlow support, which is the most direct and low-code way to meet the requirement of keeping data in BigQuery.

How to eliminate wrong answers

Option B is wrong because deploying to Vertex AI Prediction and using a remote model adds unnecessary complexity and latency; while it works, it requires setting up a remote model and a connection, which is not the simplest or most direct method when BQML directly supports TensorFlow imports. Option C is wrong because exporting BigQuery data to Cloud Storage and using AI Platform Prediction moves data out of BigQuery, violating the requirement to keep data in BigQuery and adding extra steps and cost. Option D is wrong because AutoML Tables retrains the model from scratch, which does not reuse the existing TensorFlow model and may produce different results, whereas the requirement is to use the existing model as-is.

92
MCQeasy

A junior engineer on your team has a trained scikit-learn model saved as a local joblib file and wants other teams to be able to discover it, view its evaluation metrics, and deploy it to a Vertex AI Endpoint. Which action should they take first?

A.Create a Vertex AI Endpoint and manually copy the joblib file onto the underlying prediction nodes.
B.Upload the model artifact to Vertex AI Model Registry with its serving container and metadata.
C.Copy the joblib file into a shared Cloud Storage bucket and email the object path.
D.Commit the joblib file to the team's Git repository and share the repository link.
AnswerB

Vertex AI Model Registry is the central catalog where models become discoverable, versioned, and deployable to endpoints. Importing the joblib artifact with the appropriate pre-built scikit-learn serving container and attaching evaluation metrics gives other teams visibility and a one-click path to deployment. This is the correct first step for sharing a model across teams.

Why this answer

To let other teams discover, evaluate, and deploy a model, it must be registered in a shared catalog. Vertex AI Model Registry stores the artifact, its serving container, and metadata such as evaluation metrics, and it integrates directly with endpoints for deployment. Git, raw Cloud Storage objects, and manual node manipulation all lack the cataloging, versioning, and deployment integration the scenario requires.

Exam trap

The trap here is treating a shared file location as sufficient for model sharing, when cross-team discovery and deployment require registration in a model catalog with serving metadata.

93
MCQmedium

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a data preprocessing step, a training step, and an evaluation step. The evaluation step must run only if the training step succeeds and the evaluation metric meets a threshold. The engineer wants to define this logic natively in the pipeline without writing a custom component that exits with a specific code. Which Vertex AI Pipelines feature should they use?

A.Use a pipeline parameter of type 'bool' to control whether the evaluation step runs, and set it based on the metric before submitting the pipeline.
B.Configure the pipeline's 'exit handler' to check the metric and either continue or fail the pipeline.
C.Use the kfp.dsl.Condition context manager to conditionally execute the evaluation step based on the metric value.
D.Set the evaluation step's 'trigger' policy to 'on_success' and pass the metric threshold as a pipeline parameter.
AnswerC

The kfp.dsl.Condition context manager allows you to define conditional execution in a pipeline based on the value of a pipeline parameter or a task output. You can compare the evaluation metric to a threshold and conditionally run subsequent steps. This is a native KFP feature and does not require custom exit-code logic, making it the correct approach for conditional execution based on a metric.

Why this answer

Vertex AI Pipelines supports conditional execution through the kfp.dsl.Condition context manager. This allows you to branch based on the value of a task output, such as an evaluation metric. You can wrap the evaluation step in a Condition that checks if the metric exceeds a threshold.

The other options either misuse trigger policies, misunderstand exit handlers, or rely on static pipeline parameters, none of which provide runtime conditional execution based on a metric.

Exam trap

The trap here is confusing task trigger policies (which are about upstream success/failure) with conditional execution based on data or metrics.

94
MCQmedium

A logistics company wants to classify shipping documents into categories such as invoice, packing slip, and bill of lading. They have a small set of labeled documents (about 50 per category) and want to use a low-code approach. They need a model that can be trained quickly and deployed for online predictions. Which Google Cloud service should they use?

A.AutoML Natural Language
B.Document AI
C.Natural Language API
D.BigQuery ML
AnswerA

AutoML Natural Language allows training custom text classification models with as few as 50 labeled documents per category. It provides a low-code UI and API, and can deploy models for online predictions. It is designed for exactly this scenario: small labeled datasets and quick training.

Why this answer

AutoML Natural Language is designed for custom text classification with small labeled datasets. It provides a low-code environment, automatically handles preprocessing, and allows easy deployment for online predictions, making it ideal for classifying shipping documents.

Exam trap

The trap here is assuming that the pre-trained Natural Language API can be customized, but it only offers general categories, not user-defined ones.

95
MCQeasy

You are deploying a scikit-learn model to Vertex AI for online prediction. The model was trained on a dataset with numerical features and expects input data in a specific JSON format. You have created a custom container that serves the model using a Flask app. After deploying the model to a Vertex AI Endpoint, you send a prediction request with a JSON payload, but the response is an error indicating that the input format is invalid. What is the most likely reason for this error?

A.The custom container's Flask app does not parse the incoming JSON correctly because it expects a different key or structure than what was sent.
B.The Vertex AI Endpoint requires the input data to be base64-encoded.
C.The Vertex AI Endpoint automatically converts the input to a TensorFlow tensor, which is incompatible with scikit-learn models.
D.The model artifact was not uploaded to Cloud Storage, so the container cannot load the model and returns an input format error.
AnswerA

When you deploy a custom container, the prediction request payload is passed to your container as-is. If your Flask app expects a specific JSON schema (e.g., a key named 'instances' or 'data'), but you send a different structure, it will fail to parse the input. This is a common mistake when the container's preprocessing logic does not match the request format.

Why this answer

The correct answer is that the custom container's Flask app does not parse the incoming JSON correctly. Vertex AI passes the raw request to the container, so the container must implement logic to extract and format the input as expected by the model. If the JSON structure sent by the client does not match what the container expects, an input format error occurs.

Ensuring alignment between the client request and container preprocessing is essential.

Exam trap

The trap here is assuming that Vertex AI automatically transforms the input for scikit-learn, when in fact the custom container is fully responsible for parsing and preprocessing the request.

96
MCQmedium

You need to run batch predictions on 10 TB of text data stored in BigQuery using a custom container model hosted in Vertex AI. What is the most cost-effective and simple approach?

A.Use Vertex AI batch prediction with BigQuery source and sink.
B.Use Cloud Run jobs to read from BigQuery and write results back.
C.Export BigQuery data to GCS, then run a Dataflow pipeline to call the model's online prediction endpoint for each row.
D.Use Cloud Dataproc to spin up a Spark cluster and run the model inference in parallel.
AnswerA

Vertex AI batch prediction natively reads from and writes to BigQuery, avoiding data export and re-import. This directly satisfies the cost-effective and simple constraint for 10 TB of text, since no intermediate Cloud Storage staging or custom extraction code is needed.

Why this answer

Vertex AI batch prediction natively supports BigQuery as both source and sink, allowing you to run predictions on 10 TB of text data without any data movement or intermediate storage. This is the most cost-effective and simple approach because it eliminates the need for exporting data, managing infrastructure, or calling online endpoints, and it leverages Vertex AI's optimized batch inference infrastructure that scales automatically.

Exam trap

The exam often tests the misconception that you must export data from BigQuery to GCS before running batch predictions, when in fact Vertex AI batch prediction can directly read from and write to BigQuery, making the export step unnecessary and cost-inefficient.

How to eliminate wrong answers

Option B is wrong because Cloud Run jobs have a maximum request timeout of 60 minutes and are not designed for processing 10 TB of data efficiently; they would require complex batching and retry logic, and would incur higher costs due to per-request pricing and lack of native BigQuery integration. Option C is wrong because exporting data to GCS and then using Dataflow to call the online prediction endpoint for each row introduces unnecessary data movement, storage costs, and network latency; online endpoints are designed for low-latency single requests, not high-throughput batch processing, and this approach would be both slower and more expensive. Option D is wrong because Cloud Dataproc requires you to manage a Spark cluster, handle autoscaling, and write custom inference code, which adds operational complexity and cost for a task that Vertex AI batch prediction can handle natively with no infrastructure management.

97
MCQmedium

You manage a Vertex AI Model Monitoring job on an Endpoint that serves an image classification model. The monitoring job reports feature skew for the input feature 'brightness' but no prediction drift. You want to determine whether the skew is caused by a change in the distribution of incoming images compared to the training data. Which monitoring configuration should you inspect first?

A.The prediction drift configuration's default threshold for the model's output.
B.The endpoint's traffic split configuration.
C.The training dataset specified in the monitoring job's skew configuration.
D.The sampling rate set in the monitoring job.
AnswerC

Feature skew compares the live prediction input distribution to the training data distribution. Inspecting the training dataset baseline reveals whether the baseline itself has shifted or is unrepresentative, which directly explains the reported skew. Without verifying the baseline, you cannot determine if the skew is genuine or an artifact of a stale or incorrect training set.

Why this answer

Feature skew in Vertex AI Model Monitoring is calculated by comparing the distribution of live prediction inputs to the training data distribution. When skew is flagged, the first step is to verify that the training dataset used as the baseline is correct and representative. If the baseline is outdated or mismatched, the skew alert may be misleading.

Inspecting the training dataset configuration is therefore the correct initial diagnostic step.

Exam trap

The trap here is confusing feature skew with prediction drift and looking at output-related thresholds or sampling settings instead of the training baseline.

98
MCQmedium

You need to perform batch predictions on 10 TB of data stored in BigQuery using Vertex AI. The model requires some preprocessing that cannot be expressed in SQL. What is the most scalable approach?

A.Use a Cloud Function to preprocess each row and write to a new BigQuery table, then run batch prediction.
B.Use Dataflow to read from BigQuery, perform preprocessing, write results to GCS, then run Vertex AI batch prediction job with GCS source.
C.Use Vertex AI batch prediction with BigQuery source and include preprocessing logic in the model container.
D.Export BigQuery data to CSV, run a local Python script for preprocessing, then upload to GCS and start a batch prediction job.
AnswerB

Dataflow handles the arbitrary preprocessing that SQL cannot express, reading from BigQuery and writing prepared data to Cloud Storage. Vertex AI batch prediction then consumes that GCS source at scale, decoupling preprocessing from the prediction job.

Why this answer

Dataflow (Apache Beam) provides a fully managed, auto-scaling, serverless execution environment that can read from BigQuery, apply arbitrary Python/Java preprocessing logic (e.g., feature engineering, normalization) that cannot be expressed in SQL, and write the preprocessed results to Cloud Storage (GCS). Vertex AI batch prediction can then read from GCS as input, making this the most scalable approach for 10 TB of data without requiring custom model container changes or manual data movement.

Exam trap

A common misconception is that Cloud Functions can handle large-scale batch processing, but the trap here is that Cloud Functions are designed for event-driven, short-lived tasks, not for processing terabytes of data in a batch pipeline. Dataflow is the appropriate Google Cloud service for this scenario.

How to eliminate wrong answers

Option A is wrong because Cloud Functions have a 9-minute timeout and limited memory (up to 8 GB), making them unsuitable for processing 10 TB of data row-by-row; they would require an impractical number of invocations and lack built-in parallelization for large-scale batch workloads. Option C is wrong because Vertex AI batch prediction with a BigQuery source does not support preprocessing logic inside the model container — the container receives raw data and must handle all transformations itself, which couples preprocessing to the model and violates separation of concerns; also, BigQuery source does not allow custom preprocessing steps before inference. Option D is wrong because exporting 10 TB of data to CSV, running a local Python script (single machine, no distributed processing), then uploading to GCS is not scalable — it creates a bottleneck at the local script, requires significant network transfer, and does not leverage managed services for parallel processing.

99
MCQhard

An ML engineer is training a very large PyTorch model on Vertex AI using a TPU v3 pod. The training is slower than expected, and the TPU utilization is low. What is the most likely cause?

A.The data pipeline is a bottleneck; the TPU is waiting for data.
B.The learning rate schedule is too aggressive.
C.The model is using a single TensorFlow operation not supported by TPU.
D.The batch size is too large for the TPU memory.
AnswerA

TPUs stall when input tensors arrive slower than the accelerator can consume them, so low utilisation with slow training points to host-side starvation. Feeding data via tf.data with parallel extraction and prefetching keeps the TPU v3 pod saturated rather than idle waiting on I/O.

Why this answer

The most likely cause of low TPU utilization is a data pipeline bottleneck, where the TPU spends a significant amount of time idle waiting for the next batch of data to be loaded and preprocessed. TPU v3 pods are designed for high-throughput matrix operations and can process data far faster than a typical CPU-based data loader can supply it, especially if the data pipeline uses inefficient I/O, lacks prefetching, or has insufficient workers. This mismatch starves the TPU, leading to low utilization and slower training.

Exam trap

Google often tests the misconception that low utilization is caused by model architecture or hyperparameter issues, when in reality the most common bottleneck in distributed TPU training is the data pipeline, not the compute or memory limits.

How to eliminate wrong answers

Option B is wrong because an aggressive learning rate schedule may cause training instability or divergence, but it does not directly cause low TPU utilization; utilization is a measure of hardware activity, not training convergence. Option C is wrong because the question explicitly states the model is a PyTorch model, and while PyTorch has limited TPU support compared to TensorFlow, the issue is not a single unsupported operation—such an operation would typically raise an error or fall back to CPU, not cause low utilization across the entire pod. Option D is wrong because a batch size that is too large for TPU memory would cause an out-of-memory (OOM) error, not low utilization; the TPU would fail to allocate the batch, not run slowly.

100
MCQeasy

A team wants to use Vertex AI Workbench for collaborative notebook development. They need a persistent environment that can be stopped and restarted without losing installed packages and data. Which instance type should they choose?

A.User-managed notebooks
B.Managed notebooks
C.Colab Enterprise notebooks
D.Vertex AI Pipelines
AnswerA

User-managed notebooks give a persistent Compute Engine instance with a customisable environment, so installed packages and data survive stopping and restarting. This satisfies the stem's requirement for persistence, unlike managed notebooks with ephemeral or containerised runtimes.

Why this answer

User-managed notebooks are Compute Engine instances where the user controls the environment, including installed packages, kernel state, and persistent disk contents. Because the underlying VM and its boot/data disks persist across stop and start operations, installed packages and data remain intact. This makes them the right choice for a persistent, customizable environment.

Exam trap

The trap is assuming 'managed' means 'better persistence' — in fact, user-managed notebooks give the user direct control over the persistent VM and disk, which is what preserves custom packages and data.

How to eliminate wrong answers

Option B is wrong because managed notebooks are designed for ephemeral, auto-managed environments with automatic idle shutdown and dependency management; while they persist some state, they are optimized for reproducibility rather than user-controlled persistence of arbitrary installed packages. Option C is wrong because Colab Enterprise notebooks are a managed, browser-based collaborative environment where the runtime is typically ephemeral and not intended for long-lived custom package persistence. Option D is wrong because Vertex AI Pipelines is an orchestration service for ML workflows, not a notebook environment at all.

101
MCQeasy

An ML engineer needs to monitor the online prediction latency of a Vertex AI Endpoint. Which metrics should they look at in Cloud Monitoring?

A.p50, p95, p99 latency
B.Request count and error rate
C.Skew and drift scores
D.CPU/GPU utilization
AnswerA

Percentile latencies (p50, p95, p99) expose the tail behaviour that averages hide, which matters for online prediction where a minority of slow requests breach user-facing SLAs. Vertex AI publishes these under `prediction/latencies` in Cloud Monitoring, directly satisfying the requirement to monitor endpoint latency distribution rather than throughput or resource saturation.

Why this answer

To monitor online prediction latency on a Vertex AI Endpoint, the engineer should look at percentile latency metrics such as p50, p95, and p99. These metrics provide a distribution of latency, showing typical and tail latencies, which are crucial for understanding user experience and identifying outliers.

Exam trap

PMLE often tests the difference between latency metrics and other monitoring metrics, and candidates may confuse latency with throughput or resource utilization.

How to eliminate wrong answers

Option B is wrong because request count and error rate are important but do not directly measure latency; they measure traffic and errors. Option C is wrong because skew and drift scores are used for model monitoring of data distributions, not latency. Option D is wrong because CPU/GPU utilization measures resource usage, not the latency experienced by the client.

102
MCQhard

A company uses Vertex AI Model Monitoring on an Endpoint that serves a regression model. They configure monitoring for both feature skew and prediction drift with a 10% threshold. After a week, they receive an alert that prediction drift exceeds the threshold, but feature skew remains below threshold. They want to understand what this indicates about the model's performance. What should they conclude?

A.The model is performing well because prediction drift is expected as new data arrives.
B.The model's input feature distributions have shifted significantly compared to training.
C.The training data used for the skew baseline is no longer valid, causing the prediction drift alert.
D.The model's predictions have changed distribution over time, which may indicate degradation in model performance.
AnswerD

Prediction drift measures changes in the distribution of the model's outputs over time. An alert on prediction drift alone, with no feature skew, suggests that the model's predictions are shifting even though inputs appear stable. This can happen due to concept drift or model degradation, where the relationship between inputs and outputs changes. It signals a potential need to investigate model performance or retrain.

Why this answer

Prediction drift monitors changes in the distribution of model outputs over time. When prediction drift exceeds a threshold while feature skew remains low, it suggests that the model's predictions are shifting even though input features are stable. This can indicate concept drift, where the relationship between features and target changes, potentially degrading model performance.

The appropriate response is to investigate model performance and consider retraining.

Exam trap

The trap here is assuming that any drift alert means input data has changed, ignoring that prediction drift can occur independently of feature skew.

103
Multi-Selecthard

You are monitoring a production model that is experiencing gradual decay in AUC. Which THREE metrics should you set up alerts for to diagnose the root cause? (Choose three.)

Select 3 answers
A.Concept drift score measured by comparing predicted vs actual outcomes.
B.Training-serving skew for categorical features with high importance.
C.Average prediction latency over the past hour.
D.Feature drift score for key numerical features.
E.Model staleness (days since last retraining).
AnswersA, B, D

Concept drift compares predicted outcomes against actual labels, revealing whether the relationship the model learned has changed. Since AUC decay stems from that relationship shifting, alerting on concept drift pinpoints the root cause rather than merely the symptom.

Why this answer

Option A is correct because concept drift — measured by comparing predicted values against actual outcomes — directly captures the degradation in the input-to-target relationship that would cause AUC decay. Option B is correct because training-serving skew in high-importance categorical features means the model is receiving feature values at inference that differ from those seen during training, which degrades ranking quality and AUC. Option D is correct because feature drift in key numerical features signals that the input distribution has shifted, a common root cause of gradual AUC decay.

Option C is not correct because average prediction latency is an operational performance metric, not a model-quality diagnostic for AUC decay. Option E is not correct because model staleness is a coarse proxy for retraining cadence and does not by itself diagnose the root cause of AUC degradation.

Exam trap

Google Cloud often tests the distinction between metrics that indicate a symptom (e.g., latency, staleness) versus metrics that directly measure the cause of performance decay (drift scores), leading candidates to select operational metrics instead of diagnostic ones.

104
MCQeasy

A machine learning team wants to share features across multiple models to reduce training-serving skew and ensure consistency. Which Vertex AI service should they use?

A.Vertex AI Workbench
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Experiments
AnswerC

Vertex AI Feature Store provides a centralised repository where features are computed once and served identically to both training and prediction pipelines, directly eliminating training-serving skew. Sharing one feature definition across multiple models satisfies the consistency requirement, since online and batch serving draw from the same managed source rather than duplicated transformations.

Why this answer

Vertex AI Feature Store centralizes feature storage, ensuring the same features are used for training and serving, reducing training-serving skew.

105
Multi-Selectmedium

A company wants to train a custom machine learning model on Vertex AI using a pre-built container for scikit-learn. They want to use spot VMs to reduce costs. However, the training job fails intermittently due to preemption. Which TWO actions should they take to ensure the training job completes successfully?

Select 2 answers
A.Use a larger machine type to reduce training time
B.Increase the number of parallel trials in hyperparameter tuning
C.Set the worker_pool_specs to use spot VMs by setting spot=True
D.Set the max_retry_count in the worker pool spec to a value greater than 0
E.Implement checkpointing in the training code to save model state periodically to Cloud Storage
AnswersD, E

Vertex AI will retry the job if preempted up to max_retry_count times.

Why this answer

To handle spot VM preemptions, the training job must be restartable. Using checkpoints allows the job to resume from the last saved state. Vertex AI automatically retries on preemption if the job is restartable (managed by the service).

Setting max_retry_count in the worker pool spec allows Vertex AI to automatically restart the job after preemption. Also, reducing machine type or increasing parallel trials are not direct solutions.

106
MCQeasy

A team is using Cloud Composer to orchestrate ML workflows. They want to allow multiple data scientists to contribute DAGs without interfering with each other. What is the recommended approach?

A.Give each data scientist write access to the DAGs folder in Cloud Storage
B.Use a complex naming convention for DAG files to avoid overwriting
C.Store DAGs in a source control repository and use CI/CD to deploy to Cloud Composer
D.Create a separate Cloud Composer environment for each data scientist
AnswerC

Cloud Composer syncs DAGs from a single bucket, so direct edits cause collisions between data scientists. A source control repository with CI/CD deployment gives each contributor isolated branches and controlled merges into the shared environment.

Why this answer

Cloud Composer (based on Apache Airflow) recommends managing DAGs via source control and CI/CD pipelines to ensure version control, code review, and consistent deployment. This prevents conflicts when multiple data scientists contribute, as each change is tracked and tested before being synced to the DAGs folder in Cloud Storage, avoiding overwrites or broken workflows.

Exam trap

The trap here is that candidates may assume direct write access or naming conventions are sufficient for collaboration, but the Google Cloud recommended approach emphasizes source control and CI/CD to enforce code quality and prevent deployment conflicts.

How to eliminate wrong answers

Option A is wrong because giving each data scientist direct write access to the DAGs folder in Cloud Storage bypasses version control and can lead to accidental overwrites, conflicts, or deployment of untested code, breaking production workflows. Option B is wrong because a complex naming convention does not prevent race conditions or overwrites when multiple data scientists upload files simultaneously; it only reduces the probability of name collisions but does not address the core need for controlled, auditable deployments. Option D is wrong because creating a separate Cloud Composer environment for each data scientist is cost-prohibitive, inefficient, and defeats the purpose of shared orchestration; it also introduces overhead in managing multiple environments and does not solve the collaboration problem at scale.

107
Drag & Dropmedium

Drag and drop the steps to set up a distributed training job on Vertex AI using a custom container in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct sequence for setting up a distributed training job on Vertex AI using a custom container is to first prepare the training code and Dockerfile, then build and push the container image to a registry, then configure the job with the image URI and distributed settings, and finally submit the job. This order ensures that all dependencies are met at each step.

108
MCQhard

A company uses a custom container on Vertex AI Prediction. They want to send custom metrics from their prediction container to Cloud Monitoring. Which method should they use?

A.OpenCensus or OpenTelemetry SDK
B.Vertex AI built-in metrics
C.Stackdriver Monitoring agent installed in the container
D.Cloud Logging log-based metrics
AnswerA

Custom containers cannot rely on Vertex AI's automatic metric collection, so the OpenCensus or OpenTelemetry SDK must instrument the code and export custom metrics to Cloud Monitoring. This satisfies the requirement to emit metrics from inside the prediction container itself.

Why this answer

OpenCensus and OpenTelemetry are the recommended open-source frameworks for exporting custom metrics from custom containers on Vertex AI Prediction to Cloud Monitoring. They provide a standardized way to instrument your application code, collect metrics, and send them directly to Cloud Monitoring via the Cloud Monitoring API, without requiring additional agents or log-based workarounds.

Exam trap

The trap here is that candidates often confuse built-in Vertex AI metrics (which are automatic but limited) with the need for custom metrics, or they incorrectly assume that log-based metrics are the simplest path, when in fact OpenCensus/OpenTelemetry are the direct and recommended method for custom containers.

How to eliminate wrong answers

Option B is wrong because Vertex AI built-in metrics only cover default infrastructure metrics (e.g., CPU, memory, request latency) and cannot capture custom application-level metrics defined by the user. Option C is wrong because the Stackdriver Monitoring agent (now the Ops Agent) is designed for VM-based environments and is not intended to be installed inside a container; it would add unnecessary overhead and is not the recommended pattern for custom containers on Vertex AI. Option D is wrong because Cloud Logging log-based metrics require you to write metrics as structured log entries and then define metric filters, which is an indirect, higher-latency approach compared to directly exporting metrics via OpenCensus/OpenTelemetry, and it is not the standard method for custom containers in Vertex AI Prediction.

109
Drag & Dropmedium

Drag and drop the steps to set up a feature store for ML features using Vertex AI Feature Store in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct sequence to set up a feature store using Vertex AI Feature Store is: first define the entity type and its features (schema), then ingest feature data into the store, then serve those features to models for predictions, and finally monitor the store for data quality, freshness, and drift. This ensures a logical flow from definition to consumption to observability.

110
MCQhard

You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?

A.Compress the model file and decompress on startup.
B.Use a machine type with local SSD to speed up model loading.
C.Switch to batch prediction to avoid online cold start.
D.Set minReplicas to 1 to keep at least one instance always running.
AnswerD

Keeping minReplicas at 1 maintains a warm instance, so the 5 GB model file is already loaded and the container initialised. This directly eliminates the 10-minute cold start, since new requests hit a running replica rather than triggering a fresh container pull and model load.

Why this answer

Setting minReplicas to 1 keeps at least one prediction instance warm, eliminating the cold start caused by loading the 5 GB model from scratch on each new deployment. Vertex AI Prediction scales replicas based on traffic; with minReplicas=1, the model stays loaded in memory and responds immediately. Compression, local SSD, and batch prediction do not address the fundamental issue of keeping a warm online endpoint.

Exam trap

PMLE often tests cold-start mitigation — candidates confuse storage-level optimizations (SSD, compression) with the actual fix, which is keeping a warm replica via minReplicas.

How to eliminate wrong answers

Option A is wrong because compressing and decompressing on startup still requires loading the full model into memory on every cold start, adding CPU overhead and not eliminating the latency. Option B is wrong because local SSD may speed disk reads but the dominant cold-start cost is model loading and container initialization, which minReplicas avoids entirely. Option C is wrong because batch prediction is asynchronous and not suitable for online low-latency serving; it sidesteps rather than solves the cold-start problem.

111
Multi-Selectmedium

An ML engineer is configuring Vertex AI Model Monitoring for drift detection on a deployed endpoint. Which TWO settings directly affect the frequency and accuracy of drift detection? (Choose 2)

Select 2 answers
A.Model version
B.Explanation method
C.Sampling rate
D.Alerting threshold
E.Monitoring frequency
AnswersC, E

Sampling rate determines what proportion of prediction requests are logged and analysed, directly governing how much data feeds the drift baseline. A higher rate improves detection accuracy for low-traffic endpoints, while a lower rate reduces cost but risks missing subtle distribution shifts.

Why this answer

Sampling rate controls what fraction of predictions is analyzed; monitoring frequency controls how often the distribution comparison is performed. Both directly impact detection speed and accuracy.

112
MCQmedium

A financial services company uses a custom deep learning model on Vertex AI to automatically approve or reject credit card transactions. The model is explainable using Vertex Explainable AI, and the company monitors feature attribution drift with thresholds defined per feature. Last week, the monitoring system flagged that the mean absolute attribution score for the 'transaction_amount' feature increased from 0.35 to 0.55. The overall model accuracy, measured on a daily batch of labeled transactions, has remained around 97%. The operations team is concerned about potential compliance issues due to changing model behavior. What should the data scientist do?

A.Tune the alert threshold for 'transaction_amount' to 0.6 to avoid future false alarms.
B.Retrain the model by increasing regularization to reduce the importance of the 'transaction_amount' feature.
C.Investigate whether there has been a shift in the distribution of 'transaction_amount' values in the recent transaction data, which could explain the attribution change.
D.Disable the feature attribution drift monitoring for 'transaction_amount' since the model accuracy is stable.
AnswerC

Rising mean absolute attribution for transaction_amount with stable accuracy suggests the feature's values themselves have shifted, changing its influence. Checking the recent distribution confirms whether input drift explains the attribution change, satisfying the compliance concern about altered model behaviour.

Why this answer

A shift in the distribution of the 'transaction_amount' feature (e.g., due to seasonality or a new customer segment) can naturally cause its attribution score to change without indicating model degradation. Vertex Explainable AI computes feature attributions relative to the current data distribution; if the input values shift, the model's reliance on that feature may legitimately increase. Investigating the distribution shift is the first diagnostic step before adjusting thresholds or retraining, as stable accuracy does not rule out data drift that could lead to compliance issues.

Exam trap

The trap here is that candidates assume stable accuracy means the model is fine, but the PMLE exam tests that feature attribution drift can indicate a change in model behavior that accuracy alone cannot detect, especially for compliance-sensitive applications.

How to eliminate wrong answers

Option A is wrong because tuning the alert threshold to 0.6 without understanding the root cause ignores the possibility of a real distribution shift or model behavior change, and could mask a genuine compliance risk. Option B is wrong because increasing regularization to reduce the importance of 'transaction_amount' is a premature intervention that could harm model performance and does not address why the attribution changed; it assumes the change is harmful without evidence. Option D is wrong because disabling monitoring for a feature based solely on stable accuracy is dangerous—accuracy can remain high while feature attributions drift, leading to biased or non-compliant decisions that accuracy alone does not capture.

113
MCQeasy

A data scientist wants to perform feature engineering on a large dataset stored in BigQuery before training a model. Which feature engineering tool is most appropriate?

A.Use Vertex AI Feature Store to store engineered features
B.Export data to Cloud Dataproc for feature engineering
C.Create a Dataflow pipeline to compute features
D.Use BigQuery ML TRANSFORM clause
AnswerD

BigQuery ML's TRANSFORM clause applies feature engineering expressions inside BigQuery, so transformations run where the large dataset already resides. This satisfies the constraint of processing data in place without exporting it, and the transformations are automatically applied during training and prediction.

Why this answer

BigQuery ML's TRANSFORM clause allows feature engineering to be performed directly within BigQuery using SQL, eliminating the need to export data or build separate pipelines. It automatically applies transformations during model training and prediction, ensuring consistency and efficiency for large datasets. This is the most integrated and appropriate tool for feature engineering on BigQuery data.

Exam trap

PMLE often tests the misconception that external tools like Dataproc or Dataflow are required for feature engineering on BigQuery data, when in fact BigQuery ML's TRANSFORM clause is designed for this purpose.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store is designed for storing, serving, and sharing engineered features, not for performing the initial feature engineering on raw BigQuery data. Option B is wrong because exporting data to Cloud Dataproc adds unnecessary complexity and latency, and is not the most appropriate tool when BigQuery ML can handle it natively. Option C is wrong because creating a Dataflow pipeline for feature engineering is overkill and less integrated than using BigQuery ML's built-in TRANSFORM clause.

114
Multi-Selecthard

A machine learning team needs to ensure that the same features used for training are used for serving in production to avoid training-serving skew. They use Vertex AI Feature Store. Which THREE actions should they take?

Select 3 answers
A.Enable point-in-time correct retrieval when creating training datasets
B.Use different feature views for training and serving to compare performance
C.Use the same feature view for both training data export and online serving
D.Export training data from the online store directly
E.Set up feature monitoring to detect drift in feature distributions
AnswersA, C, E

Point-in-time correct retrieval reconstructs the feature values that existed at each training label's timestamp, preventing future data from leaking into training. This directly satisfies the stem's requirement that training features match those available at serving time, eliminating training-serving skew.

Why this answer

Option A is correct because point-in-time correct retrieval ensures training datasets are built from the exact feature values that were available at the time of each event, preventing label leakage and keeping training data consistent with what would have been served historically. Option C is correct because using the same feature view for both training data export and online serving guarantees that the identical feature transformation and source are used in both paths, which is the core defense against training-serving skew. Option E is correct because feature monitoring detects drift in feature distributions between training and serving, surfacing skew or data quality issues so the team can remediate them.

Option B is incorrect because deliberately using different feature views for training and serving introduces skew rather than preventing it. Option D is incorrect because exporting training data directly from the online store is not the recommended pattern; training datasets should come from the offline store with point-in-time correctness, while the online store serves low-latency predictions.

115
MCQmedium

A team wants to run a Vertex AI pipeline that deploys a model, runs a smoke test against the endpoint, and automatically rolls back if the smoke test fails. They need the deployment step to be reversible and the endpoint to remain available during the update. Which approach should they use?

A.Deploy the new model to a separate endpoint, run the smoke test there, and then update the application's endpoint URL after the test passes.
B.Use a Vertex AI pipeline with a custom training component that trains and deploys the model in a single step, and rely on pipeline retry policies for rollback.
C.Deploy the new model as a second deployment resource on the endpoint with 0% traffic, run the smoke test, then shift traffic and delete the old deployment if the test passes.
D.Undeploy the current model, deploy the new model on the same endpoint, and run the smoke test; if it fails, redeploy the previous model.
AnswerC

Creating a new deployment resource on the same endpoint keeps the existing model serving while the new one is validated. Traffic can be shifted only after the smoke test passes, and the old resource can be removed afterward. This pattern preserves availability and makes rollback trivial because the old deployment is untouched until the new one is proven.

Why this answer

Adding a new deployment resource to the existing endpoint lets the team validate the new model without disrupting current traffic. Traffic can be moved only after the smoke test succeeds, and removing the old resource completes the rollout; if the test fails, the old resource is still serving and no rollback action is needed.

Exam trap

The trap here is treating redeployment after failure as equivalent to a rollback, when a true rollback leaves the prior deployment serving continuously with no gap.

116
MCQmedium

An organisation wants to use Document AI to process contracts but requires human review for high-risk clauses. Which feature should they enable?

A.Human-in-the-Loop (HITL)
B.Batch Processing
C.Online Prediction
D.AutoML Training
AnswerA

Human-in-the-Loop routes low-confidence or high-risk clause extractions to human reviewers before finalising results, satisfying the stem's requirement for mandatory human review of high-risk clauses. Document AI's confidence thresholds trigger this escalation automatically, so contracts proceed without manual intervention except where risk demands it.

Why this answer

Human-in-the-Loop (HITL) is the correct feature because it allows Document AI to automatically process contracts while routing high-risk clauses to human reviewers for validation. This balances automation efficiency with the need for expert oversight on sensitive content, which is a core requirement for compliance-driven document processing.

Exam trap

The trap here is that candidates confuse HITL with AutoML Training, thinking that training a model with human-labeled data is the same as having a human review live predictions, but HITL is a runtime workflow, not a training process.

How to eliminate wrong answers

Option B (Batch Processing) is wrong because it handles large volumes of documents asynchronously but does not include any mechanism for human review or intervention on specific clauses. Option C (Online Prediction) is wrong because it provides real-time predictions on individual documents but lacks the built-in workflow to pause and escalate high-risk clauses to a human. Option D (AutoML Training) is wrong because it is used to train custom models on labeled data, not to manage human review workflows during inference.

117
MCQmedium

A data science team needs to share features across multiple ML models while ensuring consistency between training and serving. Which approach best achieves this?

A.Store features in a shared BigQuery dataset without versioning
B.Export features to CSV files shared via Cloud Storage
C.Use Vertex AI Feature Store to define and serve features for both training and online prediction
D.Each team maintains its own feature engineering code in separate pipelines
AnswerC

Vertex AI Feature Store provides a centralised repository where feature values are defined once and served consistently to both training jobs and online prediction, eliminating training-serving skew. This shared definition satisfies the consistency requirement while allowing multiple models to reuse the same features.

Why this answer

Vertex AI Feature Store provides a central repository where features are defined once and reused across models, reducing training-serving skew.

118
Multi-Selecthard

An ML engineer is designing a Vertex AI pipeline that includes a custom training component. The component must read training data from a Cloud Storage bucket and write the trained model to a Vertex AI Model Registry. The engineer wants to ensure the component can access these resources securely. Which two configurations should the engineer implement? (Choose two.)

Select 2 answers
A.Use the default Compute Engine service account for all pipeline runs.
B.Set the component's environment variable GOOGLE_APPLICATION_CREDENTIALS to point to a mounted secret.
C.Specify the pipeline's service account when submitting the pipeline run.
D.Embed a service account key file in the component's container image.
E.Grant the Vertex AI Pipelines service account the necessary IAM roles for Cloud Storage and Vertex AI.
AnswersC, E

When submitting a pipeline run, you can specify a service account that the pipeline will use. This service account's permissions determine what resources the components can access. By specifying a service account with the right roles for Cloud Storage and Vertex AI, the engineer ensures secure access. This is the correct way to manage authentication for pipeline components in Vertex AI.

Why this answer

To securely access Cloud Storage and Vertex AI Model Registry, the pipeline's service account must have the required IAM roles, and that service account must be specified when submitting the pipeline run. This ensures the component authenticates with the correct permissions without embedding credentials.

Exam trap

The trap here is thinking that embedding credentials or using environment variables is necessary, when in fact Vertex AI Pipelines uses the pipeline's service account for authentication.

119
MCQmedium

A company wants to use Vertex AI Vizier to tune hyperparameters for a PyTorch model. They have a limited budget of 50 training jobs. The objective metric is validation accuracy, and they want to find the best configuration efficiently. Which algorithm should they choose?

A.Bayesian optimization using Vertex AI Vizier.
B.Random search with 50 random configurations.
C.Use a custom algorithm implemented in the training code.
D.Grid search with 50 evenly spaced points.
AnswerA

Bayesian optimisation builds a probabilistic surrogate model of validation accuracy and selects each configuration to maximise expected improvement, converging in far fewer trials than grid or random search. This suits the 50-job budget while targeting the stated objective metric.

Why this answer

Vertex AI Vizier is a hyperparameter tuning service that uses Bayesian optimization to efficiently search the hyperparameter space. With a limited budget of 50 trials, Bayesian optimization is more sample-efficient than random or grid search, making it the best choice.

Exam trap

The trap is underestimating the efficiency of Bayesian optimization. Candidates might think random search is sufficient with 50 trials, but Bayesian optimization is designed for sample efficiency, which is critical when the budget is limited.

How to eliminate wrong answers

Option B is wrong because random search does not use information from previous trials and is less efficient with a small budget. Option C is wrong because implementing a custom algorithm in the training code would require significant effort and may not be as optimized as Vizier. Option D is wrong because grid search with 50 points is a brute-force approach that does not adapt to the results of previous trials and is inefficient for high-dimensional spaces.

120
Multi-Selectmedium

A machine learning team is collaborating on a project using Vertex AI Experiments to track model training runs. They want to ensure that all team members can reproduce any experiment by using the same code, data, and environment. Which THREE actions should the team take?

Select 3 answers
A.Store the training code in a Cloud Source Repository and tag commits with the experiment ID.
B.Build a custom container image for training and push it to Artifact Registry with a fixed tag.
C.Record the path and version of the training dataset in the experiment parameters.
D.Share a service account key with all team members so they can access the same resources.
E.Use Vertex AI's hyperparameter tuning job to automatically find the best parameters.
AnswersA, B, C

Tagging commits in Cloud Source Repositories binds each Vertex AI experiment to an immutable code revision. This satisfies reproducibility by letting any team member check out the exact commit that produced a given run, rather than relying on a mutable branch.

Why this answer

Option A is correct because storing the training code in Cloud Source Repositories and tagging commits with the experiment ID creates an immutable, versioned link between the exact code revision and the Vertex AI Experiment run, which is essential for reproducibility. Option B is correct because packaging the training environment in a custom container image pushed to Artifact Registry with a fixed (immutable) tag guarantees every team member runs the identical dependencies, libraries, and runtime, eliminating environment drift. Option C is correct because logging the dataset path and version as experiment parameters in Vertex AI Experiments captures the exact data snapshot used, so the same inputs can be retrieved and reused.

Option D is not appropriate because sharing a service account key is an insecure credential-management anti-pattern and does not itself contribute to code, data, or environment reproducibility. Option E is not appropriate because hyperparameter tuning optimizes model performance and does not by itself ensure that code, data, and environment are pinned and reproducible.

Exam trap

Google Cloud often tests the distinction between actions that enable reproducibility versus actions that improve model performance or access control, so candidates mistakenly select hyperparameter tuning or service account sharing as reproducibility measures.

121
MCQeasy

A company wants to implement a document processing solution that extracts key information from invoices and receipts. They have limited ML expertise and want to use a pre-trained solution as much as possible. Which Google Cloud service should they use?

A.Document AI with a pre-trained invoice processor.
B.AutoML Natural Language with custom entity extraction.
C.Vertex AI Workbench with custom Python scripts.
D.Cloud Vision API with OCR.
AnswerA

Document AI with a pre-trained invoice processor is correct because it provides a fully managed, pre-trained solution specifically designed for extracting structured data (e.g., vendor name, invoice number, line items) from invoices and receipts, requiring no custom model training or complex coding, which aligns with the company's limited ML expertise and desire for a pre-trained solution.

Why this answer

Document AI with a pre-trained invoice processor is the correct choice because it provides a fully managed, pre-trained solution specifically designed for extracting structured data (e.g., vendor name, invoice number, line items) from invoices and receipts. This aligns with the company's limited ML expertise and desire to use a pre-trained solution, requiring no custom model training or complex coding.

Exam trap

Google Cloud often tests the distinction between raw OCR (Cloud Vision API) and structured document understanding (Document AI), leading candidates to mistakenly choose Cloud Vision API for invoice processing when they only need text extraction, not structured data extraction.

How to eliminate wrong answers

Option B is wrong because AutoML Natural Language with custom entity extraction requires users to train a custom model with labeled data, which contradicts the requirement to use a pre-trained solution as much as possible. Option C is wrong because Vertex AI Workbench with custom Python scripts demands significant ML expertise to write and deploy custom code, which the company lacks. Option D is wrong because Cloud Vision API with OCR only extracts raw text from images, not the structured key-value pairs or specific fields needed for invoice processing.

122
Multi-Selectmedium

An ML engineer is using Vertex AI for distributed training of a PyTorch model across multiple nodes. The training job must use TPUs for high throughput. The engineer sets up the job configuration. Which THREE components are required for the training to work correctly? (Select 3)

Select 3 answers
A.A startup script to configure the TPU pod (e.g., `xla_lib.sh`)
B.A MultiWorkerMirroredStrategy configuration
C.A Docker image that includes PyTorch and the TPU library (torch-xla)
D.A TF_CONFIG environment variable set for each worker
E.A CustomJob with a TPU accelerator type (e.g., v3-32)
AnswersA, C, E

Startup scripts are often needed to initialize TPU devices.

Why this answer

A is correct because TPU pods require a startup script (e.g., `xla_lib.sh`) to initialize the XLA runtime, configure the TPU mesh, and set environment variables like `XRT_TPU_CONFIG`. Without this script, the TPU devices will not be discoverable by the PyTorch/XLA process, causing the training to fail with device-not-found errors.

Exam trap

Google Cloud often tests the distinction between TensorFlow and PyTorch distributed training configurations, and the trap here is assuming that `TF_CONFIG` or `MultiWorkerMirroredStrategy` are universal for all frameworks, when in fact PyTorch uses its own environment variables and the `torch-xla` library for TPU training.

123
MCQeasy

A machine learning engineer is using Vertex AI Pipelines and wants to run a custom Python function as a component. They need to pass a dataset artifact from a previous component and output a model artifact. Which decorator should they use to define the component in the Kubeflow Pipelines SDK v2?

A.@dsl.task
B.@dsl.pipeline
C.@dsl.component
D.@dsl.container
AnswerC

The @dsl.component decorator converts a plain Python function into a KFP v2 pipeline component, with typed parameters and artifacts. It supports declaring an input dataset artifact and an output model artifact, satisfying the stem's requirement for a custom Python component.

Why this answer

The correct decorator is @dsl.component because in Kubeflow Pipelines SDK v2, this decorator is used to define a custom Python function as a reusable pipeline component. It automatically handles input and output artifact serialization, such as passing a dataset artifact from a previous component and outputting a model artifact, by leveraging the component's type annotations and the KFP artifact system.

Exam trap

Candidates often confuse @dsl.component (for custom Python functions with artifact I/O) with @dsl.container (for pre-built container images) when the question emphasizes running a custom Python function.

How to eliminate wrong answers

Option A is wrong because @dsl.task is not a valid decorator in Kubeflow Pipelines SDK v2; it is a concept from Vertex AI custom jobs, not for defining pipeline components. Option B is wrong because @dsl.pipeline is used to define the entire pipeline graph, not an individual component function. Option D is wrong because @dsl.container is used to define a component that runs a container image directly, not a custom Python function with artifact handling.

124
Multi-Selecteasy

A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)

Select 2 answers
A.DVC uses a separate .dvc file to track data versions.
B.DVC can push data to remote storage like Google Cloud Storage.
C.DVC stores the actual data files in Git.
D.DVC only works with AWS S3 as remote storage.
E.DVC replaces Git for code versioning.
AnswersA, B

DVC stores a small .dvc metafile in Git containing the data's hash and path, while the actual dataset lives outside the repository. This lightweight pointer lets Git track dataset versions without bloating the repo, satisfying the requirement to version data alongside ML code.

Why this answer

Option A is correct because DVC creates small .dvc metafiles (e.g., data.dvc) that record the MD5 hash and path of the tracked data, allowing Git to version the pointer while the actual data lives in the DVC cache. Option B is correct because DVC supports many remote storage backends, including Google Cloud Storage, via `dvc remote add` and `dvc push`, so data can be stored outside the Git repository. Option C is wrong because DVC deliberately keeps large data files out of Git, storing them in its cache and remotes instead.

Option D is wrong because DVC supports S3, GCS, Azure Blob Storage, SSH, HDFS, and local remotes, not only S3. Option E is wrong because DVC complements Git for data versioning; it does not replace Git for source code version control.

Exam trap

The trap is assuming DVC stores data in Git or is AWS-only — candidates must remember DVC uses pointer files and supports multiple remotes, and that it augments rather than replaces Git.

125
Multi-Selecthard

A company wants to implement a CI/CD pipeline for their ML models using Vertex AI. They need to automatically retrain the model when new data arrives, but only if the model performance on a validation set has degraded by more than 5% compared to the current production model. Which three services or components should they incorporate into the automated pipeline? (Choose three.)

Select 3 answers
A.Dataflow pipeline to clean the new data before training
B.Vertex AI Evaluation component to compute model performance metrics on the validation set
C.Cloud Functions to trigger the pipeline when new data arrives in Cloud Storage
D.Vertex AI Model Registry alias update to promote the model if performance passes the threshold
E.Cloud Scheduler to run the pipeline on a fixed schedule
AnswersB, C, D

The Vertex AI Evaluation component computes metrics on the validation set, producing the performance figures needed to compare against the production model. This satisfies the stem's 5% degradation gate by supplying the quantitative basis for the retrain decision.

Why this answer

Option B is correct because the Vertex AI Evaluation component is the mechanism that computes model performance metrics (such as accuracy, AUC, or RMSE) on a validation set, which is exactly what is needed to compare the newly trained model against the current production model and detect a degradation greater than 5%. Option C is correct because Cloud Functions can be configured with a Cloud Storage trigger (via Eventarc) to fire when new data objects arrive in a bucket, thereby automatically kicking off the retraining pipeline without manual intervention. Option D is correct because the Vertex AI Model Registry uses aliases (for example, a 'production' alias) to point to a specific model version, and updating that alias is how a newly validated model gets promoted to production once it passes the performance threshold.

Option A is not required by the scenario, since the question focuses on triggering, evaluating, and promoting models rather than on data cleaning, and no data quality issue is stated. Option E is incorrect because Cloud Scheduler runs jobs on a fixed time-based schedule, whereas the requirement is event-driven retraining triggered by the arrival of new data.

Exam trap

In Google Cloud, the distinction between event-driven triggers (Cloud Functions/Eventarc) and scheduled triggers (Cloud Scheduler) is commonly tested. Candidates often mistakenly choose Cloud Scheduler when the requirement is for an event-driven retraining pipeline triggered by new data arrival.

126
Multi-Selectmedium

A regulated enterprise must prove to auditors that a specific production prediction can be traced to the exact model version, training data, and pipeline execution that produced it. Their ML workflows run on Vertex AI Pipelines. Which two practices should the team adopt? (Choose two.)

Select 2 answers
A.Log each prediction request and response to Cloud Logging and include the deployed model resource ID in the log entry.
B.Enable VPC Service Controls around the Vertex AI project to restrict data exfiltration.
C.Manually maintain a spreadsheet mapping model names to training dates and update it after each release.
D.Emit pipeline parameters, artifact URIs, and execution metadata so Vertex ML Metadata records the lineage graph automatically.
E.Store periodic screenshots of the Vertex AI console showing the deployed model and its metrics.
AnswersA, D

Capturing the model resource ID with each prediction creates the link from a specific prediction to a specific registered model version. Combined with lineage in Vertex ML Metadata, an auditor can follow that model version back to its training artifacts and pipeline execution, completing the end-to-end trace the regulation demands.

Why this answer

End-to-end traceability requires two linked records: automatic lineage of pipeline executions, artifacts, and models captured by Vertex ML Metadata, and a runtime record tying each prediction to the deployed model resource. Together they let an auditor walk from a prediction to a model version and onward to the training data and execution that created it. Screenshots, network controls, and manual spreadsheets provide no verifiable lineage.

Exam trap

The trap here is assuming that security controls or manual records constitute an audit trail, when traceability specifically requires machine-recorded lineage plus a prediction-to-model link.

127
MCQeasy

A data science team is using a shared Cloud Storage bucket to store training data. Multiple team members are simultaneously uploading new data files, and occasionally the wrong version of a file is used in training, leading to inconsistent results. Which best practice should the team implement to ensure data version consistency?

A.Use Cloud Composer to schedule a daily snapshot of the Cloud Storage bucket.
B.Migrate all training data to BigQuery and use time-travel queries to access historical versions.
C.Enable object versioning on the Cloud Storage bucket and use the version ID when referencing data files.
D.Restrict write access to the bucket to only one team member using IAM roles.
AnswerC

Object versioning retains every overwrite as a distinct, immutable generation, so concurrent uploads no longer silently replace earlier files. Referencing the explicit version ID pins each training run to one exact object, satisfying the requirement for consistent, reproducible data.

Why this answer

Enabling object versioning on a Cloud Storage bucket preserves each object's history, allowing the team to reference a specific version ID when reading data files. This ensures that every training run uses the exact same version of a file, eliminating inconsistency from concurrent uploads. The version ID acts as an immutable pointer, decoupling the training process from the bucket's live state.

Exam trap

Google Cloud often tests the distinction between data versioning (object-level immutability) and data backup (snapshots or time-travel), leading candidates to choose snapshot or database-centric solutions that do not provide per-file version consistency in a shared object store.

How to eliminate wrong answers

Option A is wrong because Cloud Composer schedules workflows (e.g., Airflow DAGs) but does not provide per-object version consistency; a daily snapshot captures a point-in-time state but does not prevent concurrent uploads from overwriting files between snapshots. Option B is wrong because BigQuery time-travel queries access table snapshots within a 7-day window, but the scenario involves files in Cloud Storage, not tables; migrating all training data to BigQuery is an unnecessary architectural change that does not address file-level versioning. Option D is wrong because restricting write access to one team member creates a bottleneck and single point of failure, violating the team's need for simultaneous uploads and not solving the core problem of identifying which version is used.

128
MCQhard

A machine learning engineer needs to share a trained model with the product team for integration. The model is stored in Cloud Storage, and the product team’s service account needs read access. The engineer wants to follow the principle of least privilege. Which IAM configuration should be used?

A.Generate a signed URL with read access and share it with the product team.
B.Grant the product team's service account the roles/storage.objectViewer role at the bucket level.
C.Grant the product team's service account the roles/storage.objectAdmin role at the bucket level.
D.Grant the product team's service account the roles/storage.objectViewer role at the project level.
AnswerB

Granting roles/storage.objectViewer at bucket level gives the service account read-only access to objects within that specific bucket, satisfying least privilege by avoiding project-wide permissions. This scopes access precisely to the shared model's location, unlike broader roles such as storage.objectAdmin or project-level grants that would exceed the required read capability.

Why this answer

Granting the product team's service account the roles/storage.objectViewer role at the bucket level provides read-only access to objects in that specific bucket, adhering to the principle of least privilege. This role allows the service account to list and read objects without granting broader permissions, such as modifying or deleting them, and scoping it to the bucket prevents unnecessary access to other buckets in the project.

Exam trap

The trap here is that candidates may confuse the principle of least privilege with convenience, choosing a signed URL (Option A) because it seems simple, or selecting a project-level role (Option D) without realizing it grants access to all buckets, both of which violate the core requirement of minimal necessary permissions.

How to eliminate wrong answers

Option A is wrong because generating a signed URL with read access creates a time-limited, publicly accessible URL that bypasses IAM authentication, which violates the principle of least privilege by not using the service account's identity and potentially exposing the model to unauthorized users if the URL is leaked. Option C is wrong because granting the roles/storage.objectAdmin role at the bucket level provides full control over objects, including delete and overwrite permissions, which exceeds the required read-only access and violates least privilege. Option D is wrong because granting the roles/storage.objectViewer role at the project level gives read access to all buckets in the project, not just the specific bucket containing the model, which violates least privilege by granting broader access than necessary.

129
MCQmedium

Your team has a production ML model on Vertex AI that shows a gradual decline in accuracy over the past week. The model is retrained weekly using the latest data. Which monitoring approach should you implement to detect the issue earlier?

A.Configure Vertex AI Model Monitoring to detect feature drift and alert when metrics exceed thresholds.
B.Create a Cloud Monitoring alert for prediction response count.
C.Use BigQuery ML to retrain the model more frequently.
D.Set up a Cloud Monitoring uptime check on the prediction endpoint.
AnswerA

Vertex AI Model Monitoring computes feature drift against a training baseline and raises alerts when thresholds are breached, catching distributional shifts before weekly retraining would. This detects the gradual accuracy decline earlier than waiting for the next scheduled retrain.

Why this answer

Vertex AI Model Monitoring is designed to detect data drift and concept drift by comparing the distribution of incoming prediction requests against a baseline (e.g., training data). A gradual decline in accuracy over a week strongly suggests feature drift—the input data distribution has shifted away from what the model was trained on. By configuring drift thresholds and alerts, you can catch the drift before it significantly degrades model performance, enabling earlier retraining or investigation.

Exam trap

PMLE often tests the distinction between infrastructure monitoring (uptime, latency, request count) and model-specific monitoring (data drift, skew, accuracy), so candidates must recognize that only Vertex AI Model Monitoring addresses data distribution changes that precede accuracy drops.

How to eliminate wrong answers

Option B is wrong because a Cloud Monitoring alert on prediction response count only tracks traffic volume, not model quality or data distribution—it would not detect accuracy degradation. Option C is wrong because retraining more frequently with BigQuery ML does not monitor the production model; it only addresses the symptom after accuracy has already dropped, and BigQuery ML is not the native monitoring tool for Vertex AI endpoints. Option D is wrong because an uptime check only verifies that the prediction endpoint is reachable and responding, not whether the predictions are correct or the input data has drifted.

130
MCQeasy

A team has a trained TensorFlow model running locally and wants to deploy it for low-latency online predictions on Google Cloud. Which service should they use?

A.Vertex AI Prediction
B.AI Platform Training
C.Cloud Run
D.Cloud Functions
AnswerA

Vertex AI Prediction provides a managed endpoint with low-latency online serving, letting the trained TensorFlow model be deployed as a container and queried in real time. It satisfies the low-latency online prediction constraint that batch or local serving cannot meet.

Why this answer

Vertex AI Prediction is the correct choice because it is a fully managed service designed specifically for deploying trained ML models for online (real-time) prediction with low latency. It supports importing TensorFlow SavedModel artifacts and automatically scales the serving infrastructure, including GPU/TPU support, to handle request traffic while providing built-in monitoring and explainability features.

Exam trap

Google Cloud often tests the distinction between training and prediction services, and the trap here is that candidates may confuse AI Platform Training (which is for model training) with AI Platform Prediction (now part of Vertex AI), or assume that any serverless compute like Cloud Run or Cloud Functions can handle ML inference without considering the need for GPU/TPU support and optimized serving infrastructure.

How to eliminate wrong answers

Option B (AI Platform Training) is wrong because it is a service for training ML models, not for serving predictions; using it for online predictions would require additional custom infrastructure and does not provide the low-latency serving endpoints needed. Option C (Cloud Run) is wrong because while it can host custom containers, it lacks native ML model serving optimizations such as automatic GPU/TPU acceleration, model versioning, and request batching, and would require you to manually build and manage a prediction server. Option D (Cloud Functions) is wrong because it is a serverless compute platform for event-driven, short-lived functions with a maximum timeout of 9 minutes and no support for GPU/TPU, making it unsuitable for low-latency online predictions that require persistent, stateful serving of large ML models.

131
MCQeasy

You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?

A.Delete the endpoint and recreate it with the new model.
B.Deploy the new model version to the same endpoint and then set traffic to 100% for the new version.
C.Use Cloud Load Balancing to switch traffic between two endpoints.
D.Create a new endpoint and update the client application to point to the new endpoint.
AnswerB

Deploying the new version alongside the existing one on the same endpoint, then shifting traffic to 100%, uses Vertex AI's traffic splitting to avoid downtime. Requests route to the old version until the switch, satisfying the zero-downtime requirement.

Why this answer

Vertex AI endpoints support canary deployments by allowing you to deploy a new model version to the same endpoint and then gradually shift traffic to it using the `traffic_split` parameter. Setting traffic to 100% for the new version after deployment ensures zero downtime, as the endpoint remains active and serves requests from the old version until the switch is complete.

Exam trap

The trap here is that candidates assume a new endpoint or load balancer is required for zero-downtime updates, but Vertex AI endpoints natively support traffic splitting between model versions on the same endpoint, making external components unnecessary.

How to eliminate wrong answers

Option A is wrong because deleting and recreating the endpoint causes downtime during the deletion and creation process, and the endpoint URL changes, requiring client updates. Option C is wrong because Cloud Load Balancing is an external traffic management layer that adds unnecessary complexity and latency; Vertex AI endpoints natively support traffic splitting without needing an external load balancer. Option D is wrong because creating a new endpoint changes the endpoint URL, which requires updating client applications, leading to potential downtime or misrouting during the transition.

132
MCQhard

A team is training a model using historical data and wants to avoid data leakage when joining feature values from a feature store. The features include time-varying data like user activity counts. Which retrieval method should they use when creating a training dataset?

A.Retrieve the latest feature values for each entity
B.Aggregate features over all historical data
C.Use random sampling of feature values
D.Use point-in-time correct retrieval with timestamp matching
AnswerD

Point-in-time correct retrieval with timestamp matching returns feature values as they existed at each training example's timestamp, preventing leakage from future user activity counts. Naive latest-value retrieval would expose post-event data, inflating offline metrics relative to production.

Why this answer

Point-in-time correct retrieval with timestamp matching ensures that for each training row, the feature values used are the ones that were actually available at the time of the label event, preventing future information from leaking into the training set. This is critical for time-varying features like user activity counts, where using the latest value would introduce look-ahead bias. By joining on entity ID and event timestamp, the feature store returns the feature value as of that timestamp, mimicking the production inference environment.

Exam trap

The trap here is confusing 'latest feature values' with 'correct feature values'—candidates often assume that using the most recent data is always best, but in training, it causes data leakage and inflated offline metrics that fail in production.

How to eliminate wrong answers

Option A is wrong because retrieving the latest feature values ignores the temporal relationship between features and labels, causing data leakage (future data used to predict past events). Option B is wrong because aggregating features over all historical data also incorporates future values relative to the label timestamp, and it destroys the point-in-time semantics needed for training. Option C is wrong because random sampling of feature values breaks the entity-timestamp association and introduces noise, not a valid retrieval strategy for feature stores.

133
MCQmedium

An ML engineer is building a pipeline component that takes a dataset URI and a model URI as inputs, and outputs a classification metrics artifact. Which KFP SDK v2 type should the output artifact be annotated with?

A.Dataset
B.Metrics
C.ClassificationMetrics
D.Model
AnswerC

ClassificationMetrics is the KFP v2 artifact type purpose-built for classification evaluation output, carrying fields such as confusion matrix, precision and recall. Annotating the output with it satisfies the stem's requirement for a classification metrics artifact.

Why this answer

In KFP SDK v2, the `ClassificationMetrics` type is specifically designed to output classification metrics such as confusion matrix, ROC curve, and AUC. The question asks for a component that outputs classification metrics, so `ClassificationMetrics` is the correct artifact type. Using `Metrics` would be too generic and not provide the structured schema needed for classification-specific visualizations in the KFP UI.

Exam trap

The trap here is that candidates confuse the generic `Metrics` type (which handles scalar values) with the specialized `ClassificationMetrics` type, not realizing that KFP SDK v2 requires the specific artifact type to enable proper UI rendering and schema validation for classification outputs.

How to eliminate wrong answers

Option A is wrong because `Dataset` is used for input or output of tabular data, not for metrics artifacts. Option B is wrong because `Metrics` is a generic artifact for scalar metrics (e.g., accuracy, loss) but lacks the structured fields (e.g., confusion matrix, ROC) required for classification metrics; it would not render classification-specific visualizations in the KFP UI. Option D is wrong because `Model` is used for serialized model artifacts, not for evaluation metrics.

134
MCQmedium

An organization runs a Vertex AI pipeline that includes a model evaluation step. Team members want to reuse previously computed evaluation metrics when re-running the pipeline with unchanged code and hyperparameters. Which feature should they enable?

A.Manually store outputs in Cloud Storage and check for existence
B.Enable pipeline caching (default behavior)
C.Use the importer component to fetch previous results
D.Disable caching for the evaluation component
AnswerB

Pipeline caching reuses outputs from previously executed components when the pipeline definition, code, and inputs are unchanged, so the evaluation step's metrics are not recomputed. Enabling it satisfies the stem's requirement to reuse prior evaluation metrics across identical re-runs.

Why this answer

Vertex AI Pipelines has caching enabled by default. When a pipeline is re-run with unchanged code, hyperparameters, and inputs, the evaluation step will reuse the cached output from the previous run, saving time and cost. Team members do not need to manually store outputs or disable caching.

Exam trap

The trap is overcomplicating the solution by suggesting manual storage or importer components; candidates often forget that Vertex AI Pipelines caching is enabled by default and automatically reuses outputs when inputs and code are unchanged.

How to eliminate wrong answers

Option A is wrong because manually storing outputs in Cloud Storage and checking for existence is a custom workaround that is unnecessary when pipeline caching is already available and enabled by default. Option C is wrong because the importer component is used to import external artifacts into a pipeline, not to reuse previously computed evaluation metrics from a prior run. Option D is wrong because disabling caching would force the evaluation step to re-run, which is the opposite of what they want.

135
Multi-Selecteasy

Which TWO are benefits of using Vertex AI Pipelines for ML workflow orchestration over deploying custom Airflow DAGs in Cloud Composer? (Choose TWO.)

Select 2 answers
A.Managed infrastructure without manual configuration
B.Built-in scheduling capabilities
C.Automatic artifact lineage tracking
D.Native integration with Vertex AI services
E.Support for arbitrary Python code in steps
AnswersC, D

Vertex AI Pipelines records each step's inputs and outputs as ML Metadata artifacts, automatically building lineage across the run. This satisfies the stem's orchestration benefit: unlike custom Airflow DAGs in Cloud Composer, where lineage must be coded manually, the managed service captures provenance without extra engineering effort.

Why this answer

Option C is correct because Vertex AI Pipelines automatically tracks artifact lineage through ML Metadata, recording inputs, outputs, and parameters of each pipeline step without requiring custom instrumentation. Option D is correct because Vertex AI Pipelines natively integrates with Vertex AI services such as Vertex AI Training, Vertex AI Endpoints, and the Model Registry, enabling seamless handoff between pipeline steps and managed ML resources. In contrast, option A is not the distinguishing benefit here since Cloud Composer is also a managed service that abstracts much of the underlying infrastructure.

Option B is incorrect because Cloud Composer (Airflow) also provides robust built-in scheduling capabilities. Option E is incorrect because both Airflow DAGs and Vertex AI Pipelines support arbitrary Python code in their steps, so this is not a unique advantage.

Exam trap

Google Cloud often tests the misconception that managed infrastructure and scheduling are unique to Vertex AI Pipelines, when in fact Cloud Composer also provides these features, so candidates must focus on the specific differentiators like native integration and automatic lineage tracking.

136
MCQhard

A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?

A.Deploy the model to a Vertex AI Batch Prediction job and use online predictions only for high-value transactions.
B.Set minReplicaCount to a value that covers baseline traffic and increase the autoscaling cool-down period to avoid rapid scale-down.
C.Configure the endpoint to use a smaller machine type so that replicas start faster and cold starts are shorter.
D.Increase maxReplicaCount to 20 and lower the autoscaling target CPU utilization to 50%.
AnswerB

Keeping a baseline number of warm replicas prevents cold starts during normal traffic fluctuations. Extending the cool-down period reduces thrashing, so replicas are not removed immediately after a spike. This combination maintains low p99 latency while avoiding unnecessary replica churn, which is the main cause of the 200 ms spikes.

Why this answer

Cold starts and rapid scale-down cause p99 latency spikes. Maintaining a baseline of warm replicas ensures that new requests do not hit a cold container, and extending the cool-down period prevents replicas from being removed too quickly. This stabilizes the replica count and keeps latency low without over-provisioning for peak traffic.

Exam trap

The trap here is focusing on increasing maximum capacity or lowering the scaling target, when the real issue is cold starts and replica thrashing.

137
MCQhard

You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?

A.Enable dynamic batching in Triton.
B.Use CPU-only instances to avoid GPU memory issues.
C.Increase the number of GPU replicas.
D.Apply model quantization using TensorRT.
AnswerD

TensorRT quantisation converts FP32 weights to lower precision such as FP16 or INT8, shrinking memory footprint and boosting throughput on NVIDIA GPUs. Triton serves the optimised engine, so GPU utilisation improves while latency drops, satisfying the memory and utilisation constraints.

Why this answer

Model quantization using TensorRT reduces the precision of model weights (e.g., from FP32 to FP16 or INT8), which directly decreases GPU memory usage and can improve throughput by enabling faster arithmetic operations on compatible NVIDIA GPUs. This technique is specifically designed to optimize GPU utilization and memory footprint for large transformer models deployed with Triton Inference Server.

Exam trap

Google often tests the distinction between throughput optimization techniques (like dynamic batching) and memory footprint reduction techniques (like quantization), leading candidates to mistakenly choose dynamic batching when the question specifically asks about reducing memory footprint.

How to eliminate wrong answers

Option A is wrong because dynamic batching improves throughput by grouping inference requests, but it does not reduce the memory footprint per model instance or optimize GPU utilization in terms of memory efficiency. Option B is wrong because CPU-only instances cannot run the large transformer model with acceptable latency or throughput, and the question explicitly requires GPU. Option C is wrong because increasing the number of GPU replicas scales horizontally, which increases total memory footprint and cost, rather than reducing memory footprint per replica or optimizing utilization of a single GPU.

138
Multi-Selecthard

A company uses Vertex AI Model Monitoring to detect training-serving skew. They have a categorical feature 'product_category' with high cardinality. The monitoring job alerts for skew, but the data scientists believe the model performance is still acceptable. Which THREE actions should the team take to investigate and resolve the alert?

Select 3 answers
A.Examine which categories have the largest distribution changes to understand the nature of the shift.
B.Adjust the alerting threshold based on historical drift patterns to reduce noise.
C.Compare model performance metrics (e.g., AUC) on the drifted segment vs. the non-drifted segment.
D.Remove the drifted categories from the feature set to eliminate the alert.
E.Ignore the alert because the model is performing well; monitoring alerts are often false positives.
AnswersA, B, C

Identifying specific categories helps assess whether the drift is due to seasonal effects or other benign causes.

Why this answer

Examining which categories have the largest distribution changes allows the team to pinpoint the root cause of the training-serving skew. In Vertex AI Model Monitoring, the skew alert is based on statistical distance metrics (e.g., Jensen-Shannon divergence) between training and serving distributions. By drilling down into the specific categories driving the divergence, the team can assess whether the shift is benign (e.g., seasonal) or problematic, rather than relying on aggregate model performance alone.

Exam trap

Google Cloud often tests the misconception that a model's aggregate performance metrics (e.g., AUC) are sufficient to dismiss drift alerts, but the trap is that drift can be localized to specific segments without affecting overall metrics, requiring per-segment evaluation.

139
MCQmedium

An ML team is using Vertex AI Pipelines to automate model training and deployment. They want to reuse components across multiple pipelines. What is the best practice for managing component code?

A.Define components inline in the pipeline definition
B.Embed component code in Cloud Composer DAGs
C.Copy the component definitions into each pipeline's YAML file
D.Use Cloud Functions to define components
E.Store components as container images in Artifact Registry and reference them from pipelines
AnswerE

Packaging components as versioned container images in Artifact Registry gives immutable, portable artefacts that any pipeline can reference by digest. This satisfies cross-pipeline reuse, since component code and dependencies travel together independently of the pipeline definition or local Python environments.

Why this answer

Vertex AI Pipelines natively supports reusable components by packaging them as container images stored in Artifact Registry. This allows teams to version, share, and reference components across multiple pipelines without duplicating code, ensuring consistency and reducing maintenance overhead. Container images encapsulate the component's runtime environment and logic, making them portable and independently deployable.

Exam trap

Google Cloud often tests the misconception that inline definitions or YAML duplication are acceptable for reuse, but the trap here is that candidates overlook the requirement for versioned, decoupled, and independently deployable components, which only container images in a registry can provide.

How to eliminate wrong answers

Option A is wrong because defining components inline in the pipeline definition tightly couples the component logic to a specific pipeline, preventing reuse across multiple pipelines and making versioning difficult. Option B is wrong because Cloud Composer DAGs are used for orchestrating Apache Airflow workflows, not for defining Vertex AI pipeline components; embedding component code in DAGs would violate separation of concerns and is not a supported pattern for Vertex AI Pipelines. Option C is wrong because copying component definitions into each pipeline's YAML file leads to code duplication, version drift, and increased maintenance burden, contradicting the goal of reusability.

Option D is wrong because Cloud Functions are event-driven serverless functions, not designed to define or host reusable pipeline components; they lack the containerized runtime and dependency management required by Vertex AI Pipelines.

140
MCQhard

A team has a pipeline that trains a model and then evaluates it. They want to conditionally deploy the model to a staging endpoint only if evaluation metrics exceed a threshold. Which KFP feature should they use?

A.Use dsl.Condition (deprecated) or dsl.If to check metrics and conditionally run deployment.
B.Use dsl.ParallelFor to evaluate and deploy in parallel.
C.Use an exit handler to deploy regardless of metrics.
D.Split the pipeline into two separate pipelines and run the second only if metrics are good.
AnswerA

dsl.If evaluates the evaluation task's output metric at runtime and gates the deployment task on the threshold, so deployment runs only when metrics pass. dsl.Condition is the deprecated predecessor; dsl.If is the current KFP SDK v2 construct satisfying this conditional-deployment requirement.

Why this answer

KFP provides `dsl.Condition` (deprecated) and `dsl.If` as first-class pipeline constructs to conditionally execute pipeline components based on runtime metrics or other pipeline outputs. By wrapping the deployment step inside a `dsl.If` block that checks whether evaluation metrics exceed a threshold, the pipeline can deploy the model to a staging endpoint only when the condition is met, avoiding unnecessary deployments for underperforming models.

Exam trap

Google often tests the distinction between conditional execution (`dsl.If`) and unconditional execution patterns (exit handlers, parallel loops), tempting candidates to choose a pattern that always runs the deployment step or runs it in parallel without any gate.

How to eliminate wrong answers

Option B is wrong because `dsl.ParallelFor` is designed for iterating over a collection of items to execute the same component in parallel, not for conditionally executing a component based on a runtime evaluation result. Option C is wrong because an exit handler (e.g., `dsl.ExitHandler`) always runs a specified component when the pipeline exits, regardless of success or failure, so it would deploy the model even if metrics are poor, which contradicts the requirement. Option D is wrong because splitting the pipeline into two separate pipelines loses the benefit of a single orchestrated workflow; it introduces manual coordination, external state management, and additional operational complexity, whereas KFP’s conditional constructs handle this natively within one pipeline.

141
MCQhard

The pipeline fails during the evaluate component with error "Model not found". What is the most likely cause?

A.The dataset_id is misspelled
B.The model_id parameter is referencing the wrong output
C.The training container did not produce a model artifact
D.The threshold value is invalid
AnswerB

The evaluate component resolves its input model via the model_id parameter, which must point to the trained model output produced by an earlier pipeline step. Referencing the wrong output leaves no artefact at that path, so evaluation fails with "Model not found".

Why this answer

The error 'Model not found' during the evaluate component indicates that the model_id parameter is referencing an output that does not exist or is incorrectly named. In Vertex AI Pipelines, the evaluate component takes the model artifact from a previous training step via an output parameter or artifact reference. If the model_id parameter points to a wrong output (e.g., a different step's output or a misspelled reference), the pipeline cannot locate the model.

This is the most likely cause because the error is specific to model resolution, not dataset or threshold issues.

Exam trap

Google Cloud often tests the distinction between resource resolution errors (like 'Model not found') and data/validation errors, tricking candidates into confusing dataset or threshold issues with pipeline step output references.

How to eliminate wrong answers

Option A is wrong because a misspelled dataset_id would cause a 'Dataset not found' or data loading error, not a 'Model not found' error during evaluation. Option C is wrong because if the training container did not produce a model artifact, the pipeline would fail earlier in the training step with an artifact missing error, not during the evaluate component. Option D is wrong because an invalid threshold value would cause a validation or scoring error within the evaluate step, not a 'Model not found' error, which is a resource resolution issue.

142
Multi-Selecteasy

Which THREE factors should be considered when choosing a compute option for serving a deep learning model in production on Google Cloud? (Choose three.)

Select 3 answers
A.Integration with Vertex AI for model monitoring
B.Autoscaling capabilities to handle variable traffic
C.GPU or TPU requirements for model inference
D.The programming language used for training
E.The color of the team's logo
AnswersA, B, C

Integration with Vertex AI satisfies the requirement for production serving, since Vertex AI provides managed endpoints with built-in model monitoring for drift and skew detection. This removes the need to build custom observability around the compute layer, directly addressing the operational constraint of maintaining model quality once deployed.

Why this answer

Option A is correct because integration with Vertex AI enables managed model monitoring, which is essential for detecting training-serving skew, drift, and performance degradation in production deep learning deployments. Option B is correct because autoscaling capabilities allow the serving infrastructure to dynamically adjust compute resources based on traffic patterns, ensuring cost efficiency and consistent latency under variable load. Option C is correct because deep learning inference often requires GPU or TPU acceleration to meet latency and throughput requirements, and the chosen compute option must support the necessary hardware accelerators.

Option D is not a relevant factor because the programming language used during training does not constrain the production serving compute option, as models are typically exported to framework-agnostic formats or served via standard runtimes. Option E is clearly irrelevant because the team's logo color has no bearing on compute selection for model serving.

Exam trap

The trap here is that candidates might think the training language (D) matters for serving, but Google Cloud serving infrastructure is language-agnostic as long as the model is exported in a supported format, making this a common distractor.

143
Multi-Selectmedium

A company wants to automatically retrain their model when data drift is detected. Which THREE components are needed to implement this pipeline?

Select 3 answers
A.Cloud Function to invoke Vertex AI Pipeline
B.Vertex AI Feature Store
C.Cloud Monitoring alert policy for drift metric
D.Pub/Sub topic
E.Cloud Storage bucket for storing training data
AnswersA, C, D

A Cloud Function provides the event-driven compute layer that reacts to a drift alert and triggers the Vertex AI Pipeline, satisfying the requirement for automatic retraining without manual intervention. It bridges drift detection to pipeline execution.

Why this answer

Option A (Cloud Function to invoke Vertex AI Pipeline) is correct because the Cloud Function acts as the automation trigger that programmatically starts the Vertex AI Pipeline run to retrain the model once drift is signaled. Option C (Cloud Monitoring alert policy for drift metric) is correct because Vertex AI Model Monitoring publishes drift metrics to Cloud Monitoring, and an alerting policy on that metric is what detects the drift condition and fires the event. Option D (Pub/Sub topic) is correct because the alert policy notification channel uses Pub/Sub to deliver the drift alert, which then triggers the Cloud Function, forming the event-driven chain (Monitoring alert → Pub/Sub → Cloud Function → Vertex AI Pipeline).

Option B (Vertex AI Feature Store) is not required here since it is for serving/online feature management, not for the drift-detection-and-retrain trigger mechanism. Option E (Cloud Storage bucket for storing training data) is not required as a distinct component of this pipeline, since training data storage is already assumed to exist and is not part of the drift-triggered automation chain.

Exam trap

PMLE often tests the confusion between components that are part of the ML workflow (like Feature Store or Cloud Storage) and those that are specifically needed for event-driven automation (Pub/Sub, Cloud Function, Monitoring alert). Candidates may incorrectly include storage or feature management components as necessary for the retraining trigger.

144
MCQhard

A large e-commerce company uses Vertex AI Pipelines to orchestrate its recommendation model training. The pipeline has several parallel components: feature engineering, model training, and model evaluation. Recently, they noticed that the pipeline often fails due to resource exhaustion in the Vertex AI custom training job for the model training component. The training job consumes significant memory and occasionally exceeds the allocated memory limit, causing the pod to be OOMKilled. The team has already increased the memory to the maximum allowed for the chosen machine type. They need to prevent the pipeline from failing while still using the same machine type. Which approach should they take?

A.Split the training component into multiple smaller steps that process data in chunks to reduce peak memory usage.
B.Use a larger machine type with more memory to accommodate the peaks.
C.Add a memory check step before training that estimates memory usage and skips training if it exceeds the limit.
D.Implement a retry policy with exponential backoff for the training component, so it automatically retries on failure.
AnswerA

Chunked processing bounds peak memory by loading and training on data subsets sequentially, so the custom training job stays within the machine type's memory ceiling. This satisfies the stem's constraint of preventing OOMKilled pod failures while retaining the same machine type, rather than exceeding its maximum.

Why this answer

Splitting the training component into smaller steps that process data in chunks directly addresses the root cause of OOMKilled failures—peak memory usage exceeding the allocated limit. By reducing the memory footprint per step, the pipeline can stay within the maximum memory of the existing machine type without requiring a larger instance. This approach aligns with best practices for Vertex AI custom training jobs, where resource limits are fixed per machine type and cannot be exceeded.

Exam trap

Google Cloud often tests the misconception that retry policies or pre-checks can solve resource exhaustion, but the correct approach is to redesign the component to reduce peak memory usage, as retries do not fix the underlying OOM condition.

How to eliminate wrong answers

Option B is wrong because it suggests using a larger machine type, which contradicts the requirement to keep the same machine type; it also may increase cost unnecessarily without solving the underlying memory inefficiency. Option C is wrong because adding a memory check step that skips training on high memory usage would cause the pipeline to fail or produce no model, which does not prevent failure—it merely avoids it by not running the component. Option D is wrong because implementing a retry policy with exponential backoff does not address the resource exhaustion; the training job will repeatedly fail with OOMKilled on each retry, wasting time and compute resources without resolving the memory limit issue.

145
Multi-Selecthard

Which THREE components should you include in a comprehensive model monitoring dashboard for a production ML system?

Select 3 answers
A.Team member roles and responsibilities
B.System resource utilization (CPU, memory, latency)
C.Input data quality metrics (missing values, outliers)
D.Training pipeline code version
E.Model performance metrics (accuracy, precision, recall) over time
AnswersB, C, E

System resource utilisation exposes infrastructure saturation that degrades inference before model metrics shift, satisfying the stem's production constraint. CPU, memory and latency reveal capacity exhaustion, memory leaks and queueing under live traffic, which data drift or accuracy checks cannot detect. Including these signals keeps the dashboard comprehensive across serving health and model behaviour.

Why this answer

Option B is correct because a production ML monitoring dashboard must track system resource utilization such as CPU, memory, and latency to detect infrastructure bottlenecks, scaling issues, and SLA violations that degrade inference serving. Option C is correct because input data quality metrics like missing values and outliers enable detection of data drift, schema violations, and upstream pipeline failures before they corrupt predictions. Option E is correct because tracking model performance metrics such as accuracy, precision, and recall over time is essential to detect model degradation, concept drift, and performance regressions in production.

Option A is not a monitoring dashboard component; team roles and responsibilities belong to governance and RACI documentation, not runtime observability. Option D is not appropriate for a monitoring dashboard either, since training pipeline code version is a lineage/reproducibility artifact tracked in model registries or CI/CD metadata, not a live production monitoring signal.

Exam trap

Google Cloud often tests the distinction between operational governance artifacts (like team roles) and actual monitoring metrics; the trap here is confusing project management documentation with the technical components of a live monitoring dashboard.

146
MCQhard

A model deployed on a Vertex AI Endpoint uses an image model with XRAI explainability. The team notices that the prediction distributions are shifting over time. They want to monitor prediction drift. However, the explainability feature is not enabled. What must the engineer do to enable monitoring prediction drift?

A.Re-deploy the model with a sampling rate of 100%
B.Configure Vertex AI Model Monitoring to monitor prediction drift
C.Enable Vertex AI Explainability with XRAI on the endpoint deployment
D.Enable request/response logging to BigQuery and build custom drift detection
AnswerB

Vertex AI Model Monitoring computes prediction drift independently of explainability; XRAI only affects feature attribution. Enabling the monitoring job with a training dataset baseline and drift thresholds satisfies the stem's requirement, since drift detection compares incoming prediction distributions against the reference distribution without needing explanations enabled.

Why this answer

Vertex AI Model Monitoring is the managed service that detects prediction drift (and training-serving skew) by comparing live prediction distributions against a baseline. It operates independently of explainability features — XRAI is only for feature attribution, not drift detection. Enabling Model Monitoring on the endpoint is the correct, purpose-built action to monitor prediction drift.

Exam trap

The trap here is conflating explainability with drift monitoring — candidates assume that because XRAI is mentioned in the stem, enabling it will somehow unlock drift detection, when in fact Model Monitoring is a completely separate Vertex AI feature.

How to eliminate wrong answers

Option A is wrong because a 100% sampling rate only controls how many requests are logged for analysis; it does not itself perform drift detection and is unnecessary (and costly) when Model Monitoring can sample intelligently. Option C is wrong because XRAI explainability provides feature attributions for individual predictions, not distributional drift metrics — enabling it does not satisfy the drift-monitoring requirement. Option D is wrong because building custom drift detection from BigQuery logs is a manual, non-managed workaround; Vertex AI Model Monitoring already provides this capability natively, so reinventing it is inefficient and error-prone.

147
MCQhard

Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?

A.Switch to CPU-only instances to reduce cost.
B.Increase maxReplicas to handle the load without caching.
C.Set up a Cloud CDN in front of the endpoint.
D.Use Cloud Memorystore to cache prediction results.
AnswerD

Cloud Memorystore caches repeated predictions for the same users within short windows, so identical requests bypass GPU inference entirely. This cuts both latency and GPU cost, directly addressing the repeated-request pattern while the endpoint stays GPU-backed.

Why this answer

Cloud Memorystore (Redis) in front of the Vertex AI endpoint lets you cache prediction results keyed by user/request signature, so repeated similar requests within short windows are served from cache instead of hitting the GPU-backed model. This reduces both latency (cache hit is sub-millisecond) and cost (fewer GPU inference calls).

Exam trap

PMLE often tests whether candidates reach for infrastructure scaling (more replicas, CDN) when the correct answer is application-level caching for repeated identical or near-identical requests.

How to eliminate wrong answers

Option A is wrong because switching to CPU-only instances would increase latency for a large recommendation model and likely reduce throughput, defeating the latency goal. Option B is wrong because increasing maxReplicas scales out GPU instances, which increases cost and does nothing to reduce per-request latency for repeated queries. Option C is wrong because Cloud CDN caches HTTP responses at edge locations but is not designed for dynamic, user-specific prediction payloads and does not integrate with Vertex AI prediction semantics.

148
Multi-Selectmedium

You are preparing to deploy a trained scikit-learn model to Vertex AI for online prediction. You need to create a custom container that serves the model. Which two of the following steps are required to ensure the container works with Vertex AI? (Choose two.)

Select 2 answers
A.Ensure the container runs as a non-root user for security.
B.Include the model artifacts in the container image or ensure they are accessible at runtime.
C.Implement a web server that listens on the port specified by the AIP_HTTP_PORT environment variable.
D.Expose a /health endpoint that returns a 200 status code.
E.Use the Vertex AI SDK to build and push the container image.
AnswersB, C

The container must have access to the trained model artifacts to perform predictions. You can either bake the model into the image or mount it from a Cloud Storage location at runtime. Without the model, the container cannot serve predictions. Vertex AI expects the model to be loaded when the container starts, so providing the artifacts is essential for a functional deployment.

Why this answer

For a custom container to work with Vertex AI, it must listen on the port specified by the AIP_HTTP_PORT environment variable and have access to the model artifacts. These are the two essential steps. Other aspects like health endpoints, SDK usage, or non-root user are optional or best practices but not mandatory for the container to serve predictions on Vertex AI.

Exam trap

The trap here is assuming that a /health endpoint or non-root user is required, when the actual requirements are listening on the correct port and providing model artifacts.

149
MCQhard

An e-commerce company uses a Vertex AI endpoint for product recommendations. Recently, the click-through rate (CTR) dropped significantly. Model monitoring shows no significant data drift or skew. Logs show increased latency but no errors. Which technique should the engineer use to diagnose the issue?

A.Increase the endpoint's request timeout value to accommodate the higher latency.
B.Enable autoscaling on the endpoint to reduce latency by adding more nodes.
C.Retrain the model with the most recent user interaction data.
D.Analyze the prediction output distribution using Vertex AI Model Monitoring for prediction drift and compare to a baseline.
AnswerD

With drift and skew ruled out, prediction drift isolates changes in the model's output distribution versus baseline. Comparing score distributions reveals whether the model itself shifted behaviour, explaining the CTR drop despite stable inputs and no errors.

Why this answer

The drop in CTR despite no data drift or skew suggests that the model's predictions have shifted in distribution (prediction drift), even if the input features remain stable. Vertex AI Model Monitoring can compare the current prediction output distribution against a baseline to detect such drift, which directly explains the CTR decline. The increased latency is a symptom, not the root cause, and fixing latency alone would not restore CTR.

Exam trap

Google Cloud often tests the distinction between data drift (input distribution changes) and prediction drift (output distribution changes), and candidates mistakenly assume that no data drift means the model is fine, overlooking that the model's predictions can still degrade due to concept drift.

How to eliminate wrong answers

Option A is wrong because increasing the request timeout does not address the root cause of the CTR drop; it only masks the latency issue and may lead to worse user experience if predictions are stale. Option B is wrong because enabling autoscaling reduces latency by adding nodes, but the CTR drop is not caused by latency; it is a prediction quality issue, and autoscaling does not fix prediction drift. Option C is wrong because retraining with recent data assumes the model is stale, but monitoring shows no data drift or skew, so the input distribution is fine; the problem is in the output distribution, and retraining without investigating prediction drift may not resolve the issue.

150
MCQmedium

You are using Vertex AI continuous evaluation (model monitoring) for your deployed model. You receive an alert that the prediction distribution is significantly different from the training distribution. What should you do first?

A.Roll back the model to the previous version immediately.
B.Increase the alerting threshold to reduce false positives.
C.Analyze the input data to understand if there is a skew or drift.
D.Retrain the model using the latest data and redeploy.
AnswerC

Analysing input data distinguishes training-serving skew from prediction drift, the two distinct axes Vertex AI monitoring reports. Skew arises from preprocessing inconsistencies between training and serving pipelines; drift reflects genuine changes in live data distribution. Inspecting inputs first identifies which mechanism triggered the alert, directing remediation correctly before any retraining or pipeline fix.

Why this answer

When a monitoring alert triggers, the first step is to investigate the root cause: check if input data has changed, retraining is needed, or there is a data pipeline issue. Simply rolling back or retraining without analysis might be premature.

Page 1

Page 2 of 11

Page 3

All pages