Courseiva

PMLE · domain

scenario questions

Practise Google Professional Machine Learning Engineer scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

775 questions209 easy344 medium222 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (775)

Click any question to see the full explanation, or start a practice session above.

1

A company needs to extract key fields from scanned invoices, such as invoice number and total amount, with high accuracy. They want a managed service and plan to use human review for low-confidence results. Which combination of services should they use?

Medium
2

You are training a scikit-learn model on Vertex AI using a custom training job. The training dataset is a 2 TB CSV file stored in Cloud Storage, and the job must run on a single CPU-only VM. Loading the entire file into memory fails because the machine has only 32 GB of RAM. You need to train the model without increasing the VM size and without rewriting the training code to use a distributed framework. What should you do?

Medium
3

A company wants to analyze videos to detect objects and track their movement over time. Which TWO Google Cloud services are suitable for this task?

Medium
4

You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?

Hard
5

You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?

Medium
6

A data scientist has deployed a model on Vertex AI Endpoints and wants to monitor the model's predictions for any drift over time. Which Vertex AI service should they use?

Easy
7

A data engineer wants to use BigQuery ML to train a model that predicts customer churn using a table with customer features and a label column. They want to use a deep neural network. Which model type should they specify?

Medium
8

You are responsible for monitoring a batch prediction pipeline that runs daily. Recently, the pipeline started failing intermittently with out-of-memory errors. The input data volume has not changed. What is the most likely cause?

Medium
9

A data science team has trained a TensorFlow model and wants to serve it online with minimal latency. Which Vertex AI deployment option should they use to ensure the model can handle traffic spikes without manual scaling?

Easy
10

A pipeline uses the Google Cloud Pipeline Components to perform AutoML training and batch prediction. Which two components from the GCPC library should they use? (Choose two.)

Medium
11

A company is deploying a model on Vertex AI for online predictions with strict latency SLOs. The model requires GPU acceleration. Which TWO configurations should they consider to meet the SLOs while optimizing cost?

Medium
12

A company uses Vertex AI Pipelines with prebuilt components for data processing, training, and deployment. They need to integrate a custom validation step written in Python. What is the correct way to include this as a component?

Hard
13

A team uses Cloud Composer to orchestrate a complex ML pipeline with many tasks. They notice that the DAG parsing time is very high, causing delays in task scheduling. Which action would most effectively reduce DAG parsing time?

Hard
14

A healthcare company wants to build a model to predict patient readmission risk using structured data in BigQuery. They have a dataset with 100,000 rows and 30 features, including numerical and categorical variables. They require a model that provides explainable predictions and can be trained quickly. They decide to use BigQuery ML. Which model type should they choose?

Medium
15

A team has deployed a model to a Vertex AI Endpoint and wants to monitor the model's performance in production. They need to track the number of prediction requests and the average latency. Which Google Cloud service should they use to collect and visualize these metrics?

Easy
16

An organization uses Cloud Dataflow to preprocess training data. Dataflow jobs are often failing because of insufficient quota for certain resources. The team has requested a quota increase, but the jobs still fail with 'quota exceeded' errors for a different resource. They want to proactively monitor and manage quotas to avoid failures. What is the best approach?

Hard
17

A data analyst wants to build a binary classification model to predict customer churn using SQL queries in BigQuery. Which BigQuery ML model type should they use?

Easy
18

Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?

Easy
19

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to automatically deploy the model to an endpoint only if the evaluation metric (e.g., accuracy) exceeds a threshold. The pipeline is defined using the Kubeflow Pipelines SDK. Which approach should the engineer use to implement this conditional deployment?

Medium
20

An ML engineer is designing a Vertex AI Pipeline that includes a custom training component. The component must read a dataset from a Cloud Storage bucket and write the trained model to another Cloud Storage location. The engineer wants the component to be reusable across pipelines and to ensure that the pipeline tracks the exact dataset and model artifacts. Which approach should the engineer take?

Medium
21

Your Vertex AI custom training job is failing with an out-of-memory error on a single GPU. You need to reduce memory usage without changing the model architecture. Which approach should you try first?

Medium
22

A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?

Medium
23

You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?

Medium
24

A company wants to analyze customer reviews for sentiment (positive, negative, neutral) using a pre-trained model with no training. They have text data stored in BigQuery. Which Google Cloud service should they use?

Medium
25

An MLOps team has deployed a model on Vertex AI Endpoints and wants to monitor for skew between training and serving data distributions. Which Vertex AI service should they use?

Easy
26

Which Vertex AI service is best suited for finding similar items in a large dataset based on embedding vectors, such as product recommendations or image similarity search?

Easy
27

You are preparing to train a large image classification model on Vertex AI using a custom training job. You want to optimize the training job for cost and performance. The dataset is stored in Cloud Storage as TFRecords and is about 2 TB. You plan to use a machine with 4 NVIDIA V100 GPUs. Which two actions should you take to improve training efficiency? (Choose two.)

Hard
28

A team wants to track the lineage of ML pipeline runs, including which datasets, parameters, and models were used in each execution. Which Vertex AI service should they use?

Easy
29

A team needs to quickly create a visual interface for data exploration and model building without writing code. They want to run AutoML jobs and visualize results. Which Google Cloud tool should they use?

Easy
30

A data scientist is creating a Vertex AI pipeline using the Kubeflow Pipelines SDK v2. Which TWO statements about pipeline parameters are correct? (Choose two.)

Easy
31

An ML team uses Delta Lake on Dataproc for data versioning. Which THREE benefits does Delta Lake provide?

Medium
32

You are a machine learning engineer working on a team that uses Vertex AI Feature Store. A colleague has created a new feature and wants to make it available to other teams for training and serving. You need to ensure that the feature can be discovered and reused across projects. What should you do?

Easy
33

A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?

Hard
34

A data scientist finishes training a model in a Vertex AI Workbench notebook and wants to save it so that the deployment team can later deploy it to an endpoint without re-running the notebook. The deployment team needs to see the model's version history and assign a production alias. Which action should the data scientist take?

Easy
35

A data science team is using AI Platform for training. They want to track hyperparameters and metrics across multiple experiments. What should they use?

Medium
36

A startup wants to build a product recommendation engine without writing custom training code. They have user-item interaction data stored in BigQuery. Which Google Cloud service should they use?

Easy
37

An ML engineer is using Cloud Build to trigger a Vertex AI Pipeline on every commit to a repository. The pipeline takes 2 hours. The engineer wants to only run the pipeline when changes are made to specific directories. How can this be achieved?

Medium
38

An ML engineer wants to monitor the performance of a Vertex AI Endpoint. Which TWO metrics are available in Cloud Monitoring for Vertex AI Endpoints? (Choose 2)

Easy
39

A team is using Vertex AI Pipelines to orchestrate a training workflow. They want to ensure that the pipeline can be reproduced exactly six months later for auditing purposes. They need to capture all necessary information to rerun the pipeline and obtain identical results. (Choose two.)

Hard
40

You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)

Medium
41

You have trained a scikit-learn model and want to deploy it to Vertex AI for online predictions. You need to minimize the effort to create a custom container and ensure the model is served with the default pre-built container. What should you do?

Easy
42

Your organization wants to automate the retraining of a model when new data is available and also on a weekly schedule. Which TWO services would you use together to achieve this? (Choose two.)

Medium
43

An ML engineer manages a Vertex AI Endpoint serving a recommendation model. The team wants to detect when the distribution of a specific numerical feature, average session duration, shifts significantly from its training distribution. They have configured Vertex AI Model Monitoring with a training dataset baseline and a monitoring frequency of one hour. After a week, no drift alerts have fired even though the feature's daily mean has visibly moved. What is the most likely cause?

Hard
44

Your company uses a custom container for model serving on Vertex AI. After a recent update, the model returns predictions but they are clearly wrong (e.g., negative probabilities for a classification model). The logs show no errors. What is the most likely cause?

Hard
45

A company wants to transcribe audio from customer service calls and then analyze the sentiment of the transcribed text. Which TWO Google Cloud services should they use?

Easy
46

A company wants to automatically retrain their model every night at 2 AM using Vertex AI Pipelines. Which approach should they use to trigger the pipeline on a schedule?

Easy
47

A data scientist trained a custom TensorFlow model using Vertex AI Training and wants to deploy it for online predictions with low latency (<100ms). Which deployment option on Google Cloud is best?

Medium
48

A data science team is building a real-time feature engineering pipeline for ML model training and serving. They need to compute features from streaming data, store them for low-latency serving, and ensure consistency between training and serving. Which TWO Google Cloud services should they use?

Medium
49

A company uses Vertex AI Model Registry to manage multiple model versions. They want to designate a model version as 'champion' for production deployment and another as 'challenger' for A/B testing. Which feature of the registry should they use?

Easy
50

A marketing team wants to build a model to predict which customers are likely to churn. They have a BigQuery table with customer demographics, usage metrics, and a binary churn label. They want to use BigQuery ML and need to evaluate the model's performance. Which two statements are true regarding model evaluation in BigQuery ML? (Choose two.)

Medium
51

A small marketing team has a CSV file of 2,000 labeled customer support tickets (each with a category such as 'billing' or 'technical'). They have no ML engineers and want a fully managed, low-code way to train a text classification model that they can later call from their internal web app. Which Google Cloud service should they use?

Easy
52

A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?

Easy
53

A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?

Medium
54

A data-processing pipeline using Dataflow needs to incorporate a custom ML prediction step. The team wants to maintain fast processing and minimize latency. What is the optimal approach?

Medium
55

A data scientist has deployed a model with Vertex AI Endpoints and enabled request/response logging to BigQuery. They want to compute a confusion matrix over time to monitor model quality. What should they do?

Medium
56

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline includes a step that trains a model and a subsequent step that evaluates the model. The engineer wants to ensure that the evaluation step runs only if the training step succeeds and that the pipeline fails if the model's accuracy is below a threshold. Which approach should the engineer use?

Hard
57

A machine learning model deployed on Vertex AI is returning erroneous predictions. The team needs to investigate the root cause by examining the prediction request and response details. Which Google Cloud tool is best suited for this?

Easy
58

A company is using Vertex AI Pipelines for ML workflows. They want to implement best practices for idempotent components and data passing. Which THREE practices should they adopt?

Hard
59

An ML engineer is designing a CI/CD pipeline for ML models using Cloud Build and Cloud Deploy. They want to automatically test model performance on a validation set before promoting to production. Which step should be included in the CI/CD pipeline?

Easy
60

A team wants to share feature definitions across multiple projects in their organization using Vertex AI Feature Store. What is the recommended approach?

Medium
61

A team is scaling a prototype ML model to production on Vertex AI. The model was developed using scikit-learn and requires custom preprocessing. They want to minimize operational overhead and ensure consistency between training and serving. Which approach should they use?

Medium
62

A non-technical user wants to build a binary classification model using Vertex AI. Which UI should they use?

Easy
63

A company wants to track the cost of their Vertex AI prediction endpoint. They use a custom machine type with 1 n1-standard-4 (4 vCPU, 15 GB memory) and 1 NVIDIA T4 GPU. The endpoint is configured for automatic scaling with min=1, max=5 replicas. Which cost monitoring approach should they use?

Medium
64

Which TWO strategies help ensure data consistency when multiple teams are contributing features to a shared Vertex AI Feature Store?

Hard
65

A financial institution uses a machine learning model to approve loans. They must monitor for fairness and bias. Which THREE Google Cloud tools or features can help them achieve this? (Choose 3.)

Hard
66

A large enterprise has multiple ML models deployed in production across different regions. They want to implement a centralized monitoring dashboard that tracks key performance indicators such as prediction accuracy, latency, and error rates for all models, with the ability to drill down into individual model versions. Which approach best meets these requirements?

Hard
67

Refer to the exhibit. A data scientist runs this Vertex AI training job code. What will be the outcome?

Easy
68

A fraud detection model is deployed to a Vertex AI Endpoint and configured with Vertex AI Model Monitoring for feature drift. The team wants the drift monitor to compare live production traffic against the exact statistics captured from the training dataset, so that alerts reflect deviation from the model's original data distribution rather than from recent traffic. Which configuration should they use?

Medium
69

A financial services company uses BigQuery ML to build a logistic regression model for fraud detection. The model is trained on the last 6 months of transaction data (about 50 million rows). After deployment, the fraud detection team notices a high false positive rate, causing customer dissatisfaction and extra manual review costs. The model is currently retrained monthly. The team wants to reduce false positives without sacrificing recall. They have access to real-time transaction streaming and can compute new features quickly. What is the most effective approach?

Medium
70

A financial company is building a fraud detection model. The dataset has 1% fraud cases and 99% legitimate transactions. Which technique should they use to handle the class imbalance?

Easy
71

A team is training a TensorFlow model on Vertex AI using a custom container. The training script writes checkpoints to a local directory inside the container. The job runs for 14 hours, and when it completes, the team cannot find the checkpoints in Cloud Storage. They need the checkpoints to be persisted so they can resume training and deploy the best model. What should they do?

Hard
72

An ML team has set up automated retraining triggered by Cloud Monitoring alerts. When a feature drift alert fires, a Cloud Function publishes to Pub/Sub, which triggers a Vertex AI Pipeline. However, the retraining pipeline is failing because the training data is not updated. What is the most likely cause?

Hard
73

A data analyst wants to use Vision API to detect custom objects in manufacturing images, but the pre-trained API does not recognize their specific components. They have 1000 labeled images. Which path offers the fastest time-to-value with minimal coding?

Medium
74

A machine learning pipeline in Vertex AI produces a dataset artifact, a trained model, and evaluation metrics. The team wants to query the lineage to find all downstream artifacts that depend on a particular dataset. Which Vertex AI service should they use?

Hard
75

An ML team wants to automatically track training runs, including hyperparameters and metrics, with minimal code changes. Which Vertex AI service should they use?

Easy
76

Refer to the exhibit. A team runs the command above and sees only two models. They know there is a model 'model-v3' created three days ago. What is the most likely reason it is not listed?

Easy
77

You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)

Medium
78

A credit-risk team runs a tabular model on a Vertex AI Endpoint. They configured Vertex AI Model Monitoring with a training dataset and skew detection using the default threshold. After a week, they receive alerts that many features have high training-serving skew, but the model's business metrics (approval rate, default rate) are unchanged. They suspect the alerts are false positives due to a recent change in an upstream data pipeline that shifted feature distributions. What should they do to reduce these false alerts while still monitoring for real skew?

Medium
79

A marketing team wants to analyze customer reviews for sentiment without writing code. Which Google Cloud service should they use?

Easy
80

Two teams are collaborating on a project and want to use a shared Feature Store in Vertex AI. They need to ensure that features are discoverable and that access is controlled. What is the best practice?

Medium
81

Drag and drop the steps to deploy a trained TensorFlow model to Vertex AI Prediction in the correct order.

Medium
82

An organization uses Vertex AI Pipelines and wants to track the lineage of datasets, models, and metrics across pipeline runs. They need to query upstream and downstream dependencies of an artifact. Which service should they use?

Medium
83

An ML team uses Vertex AI Pipelines to train and evaluate models. They want to ensure that only models meeting a minimum accuracy threshold are registered in Vertex AI Model Registry. Which approach should they take?

Medium
84

A retail company wants to implement a recommendation system using Recommendations AI. They need to generate personalized recommendations for users based on their browsing history and purchase behavior. Which THREE recommendation types are available in Recommendations AI?

Hard
85

A company has a large-scale ML system that uses Vertex AI Pipelines to retrain models weekly. The pipeline includes a custom training job and a batch prediction step. After moving to production, they observe that batch prediction jobs often fail with 'Quota exceeded' errors. The project has sufficient CPU quota. What is the most likely cause?

Hard
86

A company wants to implement a centralized model registry for governance. Which two features should they use? (Choose two.)

Medium
87

You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?

Hard
88

A data scientist wants to automate the retraining of a model when new data arrives in Cloud Storage. Which Google Cloud service is most appropriate for orchestrating this workflow?

Easy
89

A company uses Vertex AI Feature Store with an online store for low-latency serving. They observe high latency during peak hours. The feature values are small (< 1 KB each) and the workload is read-heavy. Which change would most effectively reduce latency?

Hard
90

A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?

Hard
91

A company has an existing TensorFlow model for fraud detection that they want to use for predictions in BigQuery. They want to call the model from SQL queries without moving data out of BigQuery. How should they deploy the model?

Hard
92

A junior engineer on your team has a trained scikit-learn model saved as a local joblib file and wants other teams to be able to discover it, view its evaluation metrics, and deploy it to a Vertex AI Endpoint. Which action should they take first?

Easy
93

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a data preprocessing step, a training step, and an evaluation step. The evaluation step must run only if the training step succeeds and the evaluation metric meets a threshold. The engineer wants to define this logic natively in the pipeline without writing a custom component that exits with a specific code. Which Vertex AI Pipelines feature should they use?

Medium
94

A logistics company wants to classify shipping documents into categories such as invoice, packing slip, and bill of lading. They have a small set of labeled documents (about 50 per category) and want to use a low-code approach. They need a model that can be trained quickly and deployed for online predictions. Which Google Cloud service should they use?

Medium
95

You are deploying a scikit-learn model to Vertex AI for online prediction. The model was trained on a dataset with numerical features and expects input data in a specific JSON format. You have created a custom container that serves the model using a Flask app. After deploying the model to a Vertex AI Endpoint, you send a prediction request with a JSON payload, but the response is an error indicating that the input format is invalid. What is the most likely reason for this error?

Easy
96

You need to run batch predictions on 10 TB of text data stored in BigQuery using a custom container model hosted in Vertex AI. What is the most cost-effective and simple approach?

Medium
97

You manage a Vertex AI Model Monitoring job on an Endpoint that serves an image classification model. The monitoring job reports feature skew for the input feature 'brightness' but no prediction drift. You want to determine whether the skew is caused by a change in the distribution of incoming images compared to the training data. Which monitoring configuration should you inspect first?

Medium
98

You need to perform batch predictions on 10 TB of data stored in BigQuery using Vertex AI. The model requires some preprocessing that cannot be expressed in SQL. What is the most scalable approach?

Medium
99

An ML engineer is training a very large PyTorch model on Vertex AI using a TPU v3 pod. The training is slower than expected, and the TPU utilization is low. What is the most likely cause?

Hard
100

A team wants to use Vertex AI Workbench for collaborative notebook development. They need a persistent environment that can be stopped and restarted without losing installed packages and data. Which instance type should they choose?

Easy
101

An ML engineer needs to monitor the online prediction latency of a Vertex AI Endpoint. Which metrics should they look at in Cloud Monitoring?

Easy
102

A company uses Vertex AI Model Monitoring on an Endpoint that serves a regression model. They configure monitoring for both feature skew and prediction drift with a 10% threshold. After a week, they receive an alert that prediction drift exceeds the threshold, but feature skew remains below threshold. They want to understand what this indicates about the model's performance. What should they conclude?

Hard
103

You are monitoring a production model that is experiencing gradual decay in AUC. Which THREE metrics should you set up alerts for to diagnose the root cause? (Choose three.)

Hard
104

A machine learning team wants to share features across multiple models to reduce training-serving skew and ensure consistency. Which Vertex AI service should they use?

Easy
105

A company wants to train a custom machine learning model on Vertex AI using a pre-built container for scikit-learn. They want to use spot VMs to reduce costs. However, the training job fails intermittently due to preemption. Which TWO actions should they take to ensure the training job completes successfully?

Medium
106

A team is using Cloud Composer to orchestrate ML workflows. They want to allow multiple data scientists to contribute DAGs without interfering with each other. What is the recommended approach?

Easy
107

Drag and drop the steps to set up a distributed training job on Vertex AI using a custom container in the correct order.

Medium
108

A company uses a custom container on Vertex AI Prediction. They want to send custom metrics from their prediction container to Cloud Monitoring. Which method should they use?

Hard
109

Drag and drop the steps to set up a feature store for ML features using Vertex AI Feature Store in the correct order.

Medium
110

You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?

Hard
111

An ML engineer is configuring Vertex AI Model Monitoring for drift detection on a deployed endpoint. Which TWO settings directly affect the frequency and accuracy of drift detection? (Choose 2)

Medium
112

A financial services company uses a custom deep learning model on Vertex AI to automatically approve or reject credit card transactions. The model is explainable using Vertex Explainable AI, and the company monitors feature attribution drift with thresholds defined per feature. Last week, the monitoring system flagged that the mean absolute attribution score for the 'transaction_amount' feature increased from 0.35 to 0.55. The overall model accuracy, measured on a daily batch of labeled transactions, has remained around 97%. The operations team is concerned about potential compliance issues due to changing model behavior. What should the data scientist do?

Medium
113

A data scientist wants to perform feature engineering on a large dataset stored in BigQuery before training a model. Which feature engineering tool is most appropriate?

Easy
114

A machine learning team needs to ensure that the same features used for training are used for serving in production to avoid training-serving skew. They use Vertex AI Feature Store. Which THREE actions should they take?

Hard
115

A team wants to run a Vertex AI pipeline that deploys a model, runs a smoke test against the endpoint, and automatically rolls back if the smoke test fails. They need the deployment step to be reversible and the endpoint to remain available during the update. Which approach should they use?

Medium
116

An organisation wants to use Document AI to process contracts but requires human review for high-risk clauses. Which feature should they enable?

Medium
117

A data science team needs to share features across multiple ML models while ensuring consistency between training and serving. Which approach best achieves this?

Medium
118

An ML engineer is designing a Vertex AI pipeline that includes a custom training component. The component must read training data from a Cloud Storage bucket and write the trained model to a Vertex AI Model Registry. The engineer wants to ensure the component can access these resources securely. Which two configurations should the engineer implement? (Choose two.)

Hard
119

A company wants to use Vertex AI Vizier to tune hyperparameters for a PyTorch model. They have a limited budget of 50 training jobs. The objective metric is validation accuracy, and they want to find the best configuration efficiently. Which algorithm should they choose?

Medium
120

A machine learning team is collaborating on a project using Vertex AI Experiments to track model training runs. They want to ensure that all team members can reproduce any experiment by using the same code, data, and environment. Which THREE actions should the team take?

Medium
121

A company wants to implement a document processing solution that extracts key information from invoices and receipts. They have limited ML expertise and want to use a pre-trained solution as much as possible. Which Google Cloud service should they use?

Easy
122

An ML engineer is using Vertex AI for distributed training of a PyTorch model across multiple nodes. The training job must use TPUs for high throughput. The engineer sets up the job configuration. Which THREE components are required for the training to work correctly? (Select 3)

Medium
123

A machine learning engineer is using Vertex AI Pipelines and wants to run a custom Python function as a component. They need to pass a dataset artifact from a previous component and output a model artifact. Which decorator should they use to define the component in the Kubeflow Pipelines SDK v2?

Easy
124

A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)

Easy
125

A company wants to implement a CI/CD pipeline for their ML models using Vertex AI. They need to automatically retrain the model when new data arrives, but only if the model performance on a validation set has degraded by more than 5% compared to the current production model. Which three services or components should they incorporate into the automated pipeline? (Choose three.)

Hard
126

A regulated enterprise must prove to auditors that a specific production prediction can be traced to the exact model version, training data, and pipeline execution that produced it. Their ML workflows run on Vertex AI Pipelines. Which two practices should the team adopt? (Choose two.)

Medium
127

A data science team is using a shared Cloud Storage bucket to store training data. Multiple team members are simultaneously uploading new data files, and occasionally the wrong version of a file is used in training, leading to inconsistent results. Which best practice should the team implement to ensure data version consistency?

Easy
128

A machine learning engineer needs to share a trained model with the product team for integration. The model is stored in Cloud Storage, and the product team’s service account needs read access. The engineer wants to follow the principle of least privilege. Which IAM configuration should be used?

Hard
129

Your team has a production ML model on Vertex AI that shows a gradual decline in accuracy over the past week. The model is retrained weekly using the latest data. Which monitoring approach should you implement to detect the issue earlier?

Medium
130

A team has a trained TensorFlow model running locally and wants to deploy it for low-latency online predictions on Google Cloud. Which service should they use?

Easy
131

You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?

Easy
132

A team is training a model using historical data and wants to avoid data leakage when joining feature values from a feature store. The features include time-varying data like user activity counts. Which retrieval method should they use when creating a training dataset?

Hard
133

An ML engineer is building a pipeline component that takes a dataset URI and a model URI as inputs, and outputs a classification metrics artifact. Which KFP SDK v2 type should the output artifact be annotated with?

Medium
134

An organization runs a Vertex AI pipeline that includes a model evaluation step. Team members want to reuse previously computed evaluation metrics when re-running the pipeline with unchanged code and hyperparameters. Which feature should they enable?

Medium
135

Which TWO are benefits of using Vertex AI Pipelines for ML workflow orchestration over deploying custom Airflow DAGs in Cloud Composer? (Choose TWO.)

Easy
136

A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?

Hard
137

You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?

Hard
138

A company uses Vertex AI Model Monitoring to detect training-serving skew. They have a categorical feature 'product_category' with high cardinality. The monitoring job alerts for skew, but the data scientists believe the model performance is still acceptable. Which THREE actions should the team take to investigate and resolve the alert?

Hard
139

An ML team is using Vertex AI Pipelines to automate model training and deployment. They want to reuse components across multiple pipelines. What is the best practice for managing component code?

Medium
140

A team has a pipeline that trains a model and then evaluates it. They want to conditionally deploy the model to a staging endpoint only if evaluation metrics exceed a threshold. Which KFP feature should they use?

Hard
141

The pipeline fails during the evaluate component with error "Model not found". What is the most likely cause?

Hard
142

Which THREE factors should be considered when choosing a compute option for serving a deep learning model in production on Google Cloud? (Choose three.)

Easy
143

A company wants to automatically retrain their model when data drift is detected. Which THREE components are needed to implement this pipeline?

Medium
144

A large e-commerce company uses Vertex AI Pipelines to orchestrate its recommendation model training. The pipeline has several parallel components: feature engineering, model training, and model evaluation. Recently, they noticed that the pipeline often fails due to resource exhaustion in the Vertex AI custom training job for the model training component. The training job consumes significant memory and occasionally exceeds the allocated memory limit, causing the pod to be OOMKilled. The team has already increased the memory to the maximum allowed for the chosen machine type. They need to prevent the pipeline from failing while still using the same machine type. Which approach should they take?

Hard
145

Which THREE components should you include in a comprehensive model monitoring dashboard for a production ML system?

Hard
146

A model deployed on a Vertex AI Endpoint uses an image model with XRAI explainability. The team notices that the prediction distributions are shifting over time. They want to monitor prediction drift. However, the explainability feature is not enabled. What must the engineer do to enable monitoring prediction drift?

Hard
147

Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?

Hard
148

You are preparing to deploy a trained scikit-learn model to Vertex AI for online prediction. You need to create a custom container that serves the model. Which two of the following steps are required to ensure the container works with Vertex AI? (Choose two.)

Medium
149

An e-commerce company uses a Vertex AI endpoint for product recommendations. Recently, the click-through rate (CTR) dropped significantly. Model monitoring shows no significant data drift or skew. Logs show increased latency but no errors. Which technique should the engineer use to diagnose the issue?

Hard
150

You are using Vertex AI continuous evaluation (model monitoring) for your deployed model. You receive an alert that the prediction distribution is significantly different from the training distribution. What should you do first?

Medium
151

Your organization uses Vertex AI Feature Store to serve features for a real-time fraud detection model. Multiple teams contribute features, and you need to ensure that feature values are consistent between training and serving. Which practice should you implement to prevent training-serving skew?

Hard
152

A company wants to monitor features in Vertex AI Feature Store for drift over time. Which two services should they use? (Choose two.)

Medium
153

A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?

Hard
154

A data scientist deployed a TensorFlow model for sentiment analysis to Vertex AI Prediction. The model expects input key 'text' but the client sends requests with key 'review_text'. Which step should the data scientist take to resolve the error without retraining the model?

Medium
155

An ML team wants to share feature definitions across multiple projects to reduce training-serving skew and ensure consistency. They currently store features in Cloud Storage and manually coordinate updates, leading to errors. Which Google Cloud service should they use to centrally manage and serve features for both training and online inference?

Medium
156

A company runs a Vertex AI endpoint that serves a model for real-time predictions. The endpoint uses a custom container that loads a 10 GB model into memory. During a traffic spike, the autoscaler adds new replicas, but each new replica takes 8 minutes to become ready because it must download the model from Cloud Storage. The team wants to reduce scale-up time. Which approach is most effective?

Hard
157

Your team manages a production ML pipeline on Google Cloud that trains a fraud detection model every 6 hours using new transaction data. The pipeline steps are: (1) Cloud Function triggered by new files in Cloud Storage to validate data, (2) Dataflow job for feature engineering, (3) Vertex AI CustomJob for training, (4) Cloud Function to deploy the model to a Vertex AI endpoint after evaluation. You notice that the pipeline sometimes fails during the Dataflow job step with an error: 'Workflow failed. Causes: The job encountered a system error. Please try again later.' The error occurs sporadically, and retrying the pipeline manually usually succeeds. The team needs a reliable automated solution. What should you do?

Medium
158

A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?

Medium
159

An ML engineer is authoring a Vertex AI Pipelines component that runs a custom Python script. The component must accept a GCS path to training data and output a model artifact. The engineer wants the component interface to be strongly typed and to automatically generate the component specification from the Python function. Which approach should the engineer use?

Medium
160

A data scientist has trained a TensorFlow model locally and wants to deploy it to Vertex AI for online predictions. The model accepts a single input tensor of shape (1, 224, 224, 3) and outputs a probability distribution over 10 classes. The data scientist wants to minimize deployment effort and ensure the model is served with low latency. What is the simplest way to deploy this model on Vertex AI?

Easy
161

An ML team is converting a prototype model to a production pipeline using Vertex AI. They want to ensure model versioning and lineage. Which two practices should they adopt? (Select TWO)

Easy
162

A team of data scientists and ML engineers is collaborating on a project using Vertex AI Workbench. They need to share notebooks and code, but want to avoid conflicts and maintain a history of changes. Which approach should they use?

Medium
163

The exhibit shows a Vertex AI PipelineJob submission command. The pipeline fails because the component cannot find the input data. What is the most likely cause?

Easy
164

You are configuring Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. The model uses a mix of numerical and categorical features. You want to ensure that the monitoring job effectively detects drift while minimizing false alerts. Which two actions should you take? (Choose two.)

Medium
165

Your team has deployed a tabular model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with training-serving skew detection. You configured the monitoring job to run hourly and store statistics in a Cloud Storage bucket. After the first run, you notice that the job did not produce any drift metrics. You want to determine the cause of the missing metrics. What should you do first?

Medium
166

A media company wants a low-code pipeline that ingests uploaded video files, detects scenes and on-screen text, and stores structured metadata for search. They prefer managed services and minimal custom code. Which TWO Google Cloud capabilities should they combine? (Choose two.)

Hard
167

A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?

Medium
168

A team uses Vertex AI Feature Store with an online store for low-latency serving. They need to support frequent updates to features (e.g., every minute) and require high write throughput (thousands of writes per second). Which online store type should they choose?

Hard
169

An ML team wants to monitor feature drift in their production model. Which Vertex AI Feature Store capability should they use?

Easy
170

Refer to the exhibit. A team configured Vertex AI Model Monitoring with skew detection for feature "income" with a threshold of 0.2. However, they have not received any alerts even though they suspect data drift. What is the most likely reason?

Medium
171

An ML team uses Vertex AI Pipelines and wants to automatically generate model cards documenting model purpose, evaluation results, and intended use. Which approach should they take?

Easy
172

A team trained a TensorFlow model locally and wants to deploy it to BigQuery ML for predictions without retraining. They have exported the SavedModel to Cloud Storage. Which statement is correct?

Hard
173

A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)

Medium
174

A data science team deploys a custom container on Vertex AI Prediction for a PyTorch model. After deployment, the model returns predictions that are consistently off by a constant factor. The model performed correctly during local testing. What is the most likely cause?

Medium
175

A data scientist wants to perform A/B testing between two model versions deployed on the same Vertex AI endpoint. They need to route 10% of traffic to the challenger model. Which approach should they use?

Hard
176

A company wants to set up end-to-end monitoring for a Vertex AI model. Which three components should they include?

Medium
177

An organization runs a batch prediction job on Vertex AI for a large dataset (10 TB). The job is configured to use a cluster of 100 n1-standard-16 machines. Midway through, the job fails with 'Out of memory' errors. What is the most effective mitigation strategy?

Hard
178

Your team is deploying a large model on edge devices and needs to reduce its size by 80% while maintaining reasonable accuracy. Which THREE techniques should they consider? (Choose 3.)

Hard
179

A data scientist wants to train a PyTorch model on Vertex AI using a pre-built container for GPU training. She needs to use 4 NVIDIA A100 GPUs on a single machine. Which machine configuration should she select?

Medium
180

Which Vertex AI service is designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale?

Easy
181

Refer to the exhibit. A machine learning engineer deployed a model on Vertex AI using this configuration. When testing the endpoint, the engineer receives a 400 error with the message: 'Invalid argument: Explanation metadata missing required field: `outputs`.' What is the most likely cause?

Medium
182

A data analyst wants to create a classification model directly in BigQuery using SQL. Which feature should they use?

Easy
183

A data science team deploys a large language model (LLM) on Vertex AI Prediction using an NVIDIA A100 GPU. The end-to-end latency is acceptable, but the cost is high due to low GPU utilization. The model is stateless and requests are independent. Which strategy would most effectively reduce cost per prediction?

Hard
184

An ML team is using Vertex AI Pipelines to orchestrate a training workflow. They need to pass a large dataset (500 GB) between two components. The first component preprocesses the data and writes the output to Cloud Storage. The second component trains a model using that preprocessed data. The team wants to minimize pipeline execution time and cost. Which two strategies should they use? (Choose two.)

Hard
185

A team wants to implement CI/CD for their ML pipeline using Cloud Build. They want to automatically compile and deploy the pipeline when code is pushed to the main branch. Which three steps should they include in the Cloud Build configuration? (Choose three.)

Hard
186

An engineer is designing a distributed training job on Vertex AI for a TensorFlow model that uses the MultiWorkerMirroredStrategy. They need to ensure proper communication between workers. Which environment variable must be set correctly for each worker?

Hard
187

Two teams train models in separate Vertex AI projects but must share the same curated feature set. The platform team wants a single authoritative definition of each feature so that online serving and offline training always return consistent values, while each team keeps its own model training pipeline. Which approach should the platform team implement?

Hard
188

You are deploying a scikit-learn model to Vertex AI for online predictions. The model requires a custom preprocessing step that transforms raw JSON input into a feature vector before calling predict. You want to minimize latency and avoid managing infrastructure. What should you do?

Medium
189

A machine learning engineer wants to monitor model performance on Vertex AI for a regression model. Which metric is most appropriate to track the average prediction error?

Easy
190

A company deploys a custom TensorFlow model to Vertex AI Endpoint for online predictions. After deployment, prediction latency is consistently high (over 500ms) even under low traffic. The model is CPU-only and the default machine type (n1-standard-2) is used. Which action will most likely reduce prediction latency?

Medium
191

A company trains a model using Vertex AI Training and then deploys it to Vertex AI Prediction. They notice that prediction requests fail with 'InvalidArgument: input tensor shape mismatch'. Which THREE are possible causes?

Hard
192

A data scientist wants to automatically generate model documentation that includes model purpose, training data, evaluation results, and intended use. Which tool should they use?

Easy
193

Refer to the exhibit. The team wants to automatically deploy the best-performing model version to production. They have set up a Cloud Function triggered by Model Registry events. Which alias should they use in the function to get the latest champion?

Hard
194

A data analyst wants to train a binary classification model in BigQuery ML on a dataset of 10 million rows with 50 features. They need to evaluate the model's performance on a held-out test set. Which sequence of SQL statements should they run?

Medium
195

Match each model evaluation metric to its use case.

Medium
196

You need to serve a TensorFlow model that has a cold start latency of 20 seconds. The model is used for a real-time application with unpredictable traffic, but occasional bursts require immediate responses. What is the best deployment strategy to minimize both cold start impact and cost?

Easy
197

Which THREE considerations are important when setting up a shared feature store in Vertex AI Feature Store for multiple teams?

Medium
198

A healthcare organization wants to build a model to predict patient readmission risk using structured electronic health record (EHR) data. They need to train a model using SQL in BigQuery, but they also want to leverage AutoML's ability to automatically search for the best architecture. Which approach should they take?

Hard
199

Which TWO practices are important when scaling a prototype ML model to production on Google Cloud? (Choose two.)

Medium
200

A company needs to reduce inference latency for their online prediction service on Vertex AI. Which two actions would help? (Choose 2)

Medium
201

A global retailer has deployed a real-time product recommendation model on Vertex AI Endpoints. The model is a large neural network that runs on a single node with 8 vCPUs and 30 GB memory. Over the past week, the p99 latency has increased from 200ms to 2 seconds, and the error rate has risen to 5%. Cloud Monitoring shows that the endpoint's CPU utilization is consistently near 100%, and memory is at 80%. The ML engineer suspects the model is too large for the node, but model size has not changed. Logs show no increase in request volume (steady at 50 QPS). There are no recent model updates. The engineer has tried to increase the node to 16 vCPUs, but latency decreased only slightly. What is the most likely root cause and the best first step to resolve it?

Hard
202

A media company wants to automatically moderate user-uploaded videos by detecting explicit content (e.g., violence, adult material). They need a solution that integrates with their video processing pipeline and scales to millions of videos. Which approach should they take?

Hard
203

You are using Vertex AI Model Monitoring to detect prediction drift on a deployed model that serves online predictions. You want to ensure that the monitoring job can correctly compute drift metrics for numerical features. Which two configurations are required for the monitoring job to compute drift for a numerical feature? (Choose two.)

Hard
204

A company has a TensorFlow model trained outside of Google Cloud and wants to use it for online predictions on Vertex AI. They have saved the model in SavedModel format. What is the most efficient way to deploy this model?

Hard
205

An ML team is using Vertex AI to train a deep learning model on a large dataset. To reduce costs, they want to use preemptible VMs for training jobs. However, training must complete within a bounded time. Which strategy should they use?

Medium
206

Your team trains models in a shared Vertex AI project. A data engineer accidentally overwrites a BigQuery training table that three production pipelines depend on, and nobody can tell which pipeline used which version of the data. You need to make dataset versions immutable and traceable so that any training run can be reproduced. What should you do?

Medium
207

A logistics company wants to classify shipping documents into categories (invoice, packing slip, bill of lading) using a custom model with minimal code. They have labeled training images. Which Google Cloud service is most appropriate?

Medium
208

An ML engineer needs to run batch predictions on tens of petabytes of data using a trained model. The data is stored in Cloud Storage. Which service should they choose?

Medium
209

A team is designing a ML pipeline that includes training, evaluation, and conditional deployment. They want to use Vertex AI Pipelines. Which THREE concepts should they use? (Choose three.)

Hard
210

A company wants to use Vertex AI JumpStart to deploy a pre-trained image classification model and later fine-tune it on their own data. Which TWO statements are true about Vertex AI JumpStart?

Easy
211

An ML engineer is building a Vertex AI pipeline that must run a custom training component for each of 12 hyperparameter combinations. The component is defined as a custom Python function (Lightweight Python component). The engineer wants each combination to run as a separate parallel task so the pipeline completes faster, and wants the pipeline to fail fast if any single trial fails. Which approach should the engineer take?

Medium
212

A machine learning team is building a feature engineering pipeline using Dataflow. They need to compute features from streaming data and store them in Vertex AI Feature Store for online serving. The features must be updated within 5 seconds of the event. Which TWO services should they combine? (Select 2)

Hard
213

A data scientist wants to share a trained model with the team for review before deployment. The model is stored in Vertex AI Model Registry. What is the recommended way to grant the team read access to the model?

Easy
214

A financial institution needs to extract structured data from scanned PDFs of loan applications, including text fields and tables. They require a human review step for high-risk applications. Which Google Cloud service and configuration should they use?

Hard
215

An ML engineer is monitoring a deployed model on Vertex AI Endpoints and wants to detect anomalies in the distribution of a categorical feature 'product_category' which has 50 possible values. The engineer configures Vertex AI Model Monitoring with training-serving skew detection. After a week, they receive an alert that the feature is skewed. Upon investigation, they find that a new product category was introduced in the serving data that was not present in the training data. What is the most likely reason for the alert?

Hard
216

A team has deployed a model on a Vertex AI Endpoint and enabled Vertex AI Model Monitoring for feature skew. They notice that the skew metric for a categorical feature with high cardinality is consistently high, even though the feature's distribution appears stable to the team. What is the most likely cause of this high skew metric?

Hard
217

You have an edge device with limited compute resources. You need to deploy a deep learning model for real-time inference. Which model compression technique should you apply to reduce the model size and latency with minimal accuracy loss?

Hard
218

A data scientist wants to log prediction inputs and outputs for model monitoring. Which Google Cloud service is best suited for this?

Easy
219

A company is migrating from an on-premises ML serving infrastructure to Vertex AI. They have multiple models that need to be served from the same endpoint with different traffic percentages. They also need to monitor prediction quality. Which THREE actions should they take? (Choose 3)

Hard
220

A company is building a document processing pipeline using Document AI to extract data from invoices. They want to ensure high accuracy and handle edge cases where the model may be uncertain. Which THREE steps should they include in their pipeline?

Hard
221

A retailer uses BigQuery ML to build a linear regression model for sales forecasting. The model's evaluation shows high RMSE. Which step should they take first?

Medium
222

An ML team wants to run a hyperparameter tuning job on Vertex AI using a pre-built pipeline component. Which component should they use?

Medium
223

Two teams independently develop two different versions of a model for the same use case. They both deploy to the same Vertex AI endpoint, causing conflicts. What is the best way to manage multiple model versions and avoid conflicts in a collaborative environment?

Hard
224

A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?

Medium
225

A company needs to extract entities (e.g., names, dates) from customer emails using a pre-trained model. Which service should they use?

Easy
226

A data scientist needs to train a large PyTorch model on a custom dataset using Vertex AI. The training script expects data from Cloud Storage and uses GPU acceleration. Which option correctly configures a custom training job with a pre-built container for PyTorch and attaches a single NVIDIA V100 GPU?

Medium
227

An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?

Medium
228

A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?

Easy
229

A data scientist wants to track machine learning experiments, including parameters, metrics, and artifacts, and compare runs. Which Vertex AI service should they use?

Easy
230

You are collaborating with a team of data scientists on a Vertex AI Workbench notebook that preprocesses data for a machine learning model. You need to ensure that all team members can work on the notebook simultaneously without overwriting each other's changes, and that the notebook's execution environment remains consistent across the team. (Choose two.)

Medium
231

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built Google Cloud Pipeline Components (GCPC) to train a custom model. Which component should the engineer use to submit a custom training job to Vertex AI?

Easy
232

A financial institution has deployed a fraud detection model on a Vertex AI Endpoint. The model uses both numerical and categorical features. They have enabled Vertex AI Model Monitoring with training data and configured drift thresholds. After several weeks, they notice that the model's recall has dropped significantly, but the overall prediction distribution remains stable. They suspect that a specific subgroup of transactions is being misclassified. Which approach should they use to identify the subgroup and diagnose the issue?

Hard
233

An ML engineer is designing a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it only if the evaluation metric meets a threshold. The pipeline must pass the evaluation metric from the evaluation component to a downstream conditional. Which Vertex AI Pipelines feature should the engineer use to implement this flow?

Medium
234

Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?

Hard
235

An ML engineer is configuring Vertex AI Model Monitoring for a deployed model that receives a mix of numerical and categorical features. The engineer wants to ensure that the monitoring job can detect both data drift and training-serving skew. Which two configurations are required to enable both types of detection? (Choose two.)

Hard
236

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to deploy the model to an endpoint only if the evaluation metric exceeds a threshold defined at pipeline submission time. The threshold must be changeable without recompiling the pipeline. Which mechanism should the engineer use?

Medium
237

A data science team wants to share a set of engineered features across multiple projects and teams to reduce training-serving skew and ensure consistency. They need low-latency serving (single-digit milliseconds) for online predictions and also need to retrieve historical feature values for training. Which approach should they take?

Medium
238

You are an ML engineer at a fintech company. You have a prototype credit risk model built using XGBoost that achieves high accuracy on historical data. The model is trained on a dataset with 500,000 rows and 50 features. The company wants to deploy this model to production to score loan applications in real-time. The production environment must handle a peak load of 100 requests per second with a latency under 200ms. You have decided to use Vertex AI for deployment. After deploying the model as a Vertex AI endpoint with a single n1-standard-4 machine, you notice that latency exceeds 500ms at peak load and some requests time out. You have verified that the model prediction itself (excluding network overhead) takes about 50ms on average. What should you do to meet the latency and throughput requirements?

Medium
239

An ML engineer has deployed a tabular binary classification model to a Vertex AI Endpoint. After enabling Vertex AI Model Monitoring with training-serving skew detection, the engineer notices that the feature 'customer_age' is flagged as skewed. The feature values at serving time appear to be shifted upward by about 10 years compared to the training data. Which of the following is the most likely root cause?

Medium
240

Which TWO actions should be taken to ensure reproducibility of ML experiments when collaborating across teams on Vertex AI?

Hard
241

You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?

Medium
242

A data scientist wants to evaluate the performance of a BigQuery ML classification model on a test dataset. Which function should they use?

Easy
243

A company has a model serving predictions on Vertex AI Endpoints and wants to monitor for prediction drift. They enable Vertex AI Model Monitoring but also need to see a confusion matrix over time. How should they set up the confusion matrix monitoring?

Medium
244

You are designing a distributed training job on Vertex AI for a PyTorch model using DataDistributedParallel (DDP). You have 4 nodes, each with 4 GPUs. What is the total number of workers that should be configured in the TF_CONFIG equivalent for PyTorch?

Medium
245

A data scientist uses Vertex AI Pipelines to orchestrate an ML workflow. They want to reuse a component from Google's curated repository. What is the recommended way to incorporate it?

Hard
246

Which TWO actions are appropriate when you detect that a production model's prediction distribution has shifted significantly from the training distribution?

Easy
247

Which Vertex AI service is used to track the lineage of ML pipeline components, artefacts, and executions?

Easy
248

You need to preprocess a large dataset (terabytes) for training a TensorFlow model. The preprocessing includes scaling and bucketizing features, and the same transformations must be applied during serving. Which tool should you use?

Hard
249

A team is using Vertex AI AutoML to train a forecasting model. They need to retrain the model weekly and only if the new week's data significantly changes the data distribution. What is the most efficient way to achieve this?

Hard
250

A company runs a Vertex AI pipeline that uses a container component to preprocess data. The component downloads a large file from a public URL and saves the output to Cloud Storage. The pipeline fails intermittently with a 'timeout' error. Which THREE steps should the team take to improve reliability? (Choose three.)

Hard
251

You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?

Hard
252

You have an online prediction model that is showing increasing prediction latency. You have already verified that the request rate and input data size are unchanged. Which of the following should you investigate next?

Easy
253

A marketing team wants to automatically categorize customer feedback emails into topics such as 'billing', 'technical support', or 'general inquiry'. They have a dataset of 5,000 labeled emails and want to build a custom model with minimal coding effort. Which Google Cloud service should they use?

Easy
254

A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?

Medium
255

A company uses Vertex AI Prediction with a custom container for a TensorFlow model. They notice that after deploying a new model version, requests still go to the old version. What is the most likely cause?

Hard
256

A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)

Hard
257

A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?

Medium
258

A retail company wants to build a low-code ML solution to predict customer lifetime value (CLV) using historical transaction data stored in BigQuery. They have limited ML expertise and want to use BigQuery ML. Which two steps are necessary to train and evaluate a model using BigQuery ML? (Choose two.)

Medium
259

You are monitoring a machine learning pipeline that runs on Vertex AI Pipelines. The pipeline occasionally fails with a 'ResourceExhausted' error when attempting to read data from BigQuery. Which action should you take to resolve this issue?

Medium
260

An ML engineer is building a monitoring dashboard for a Vertex AI pipeline that includes training, evaluation, and batch prediction. Which THREE components should be included to provide comprehensive observability? (Select THREE.)

Hard
261

A data science team collaborates using Vertex AI Workbench user-managed notebooks. They want to version control their notebook code and share it with team members. Which TWO tools should they use? (Choose 2)

Medium
262

You have a very large language model that does not fit on a single GPU. You need to train it efficiently across multiple GPUs on a single machine. Which approach should you use?

Hard
263

What is the primary purpose of Vertex AI Edge Manager?

Easy
264

To enable collaboration on notebook-based experiments across teams, what is the recommended approach in Google Cloud?

Easy
265

You are deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that involves scaling numerical features and one-hot encoding categorical features. You have packaged the preprocessing and model into a single Python script that uses a custom prediction routine. You need to ensure that the online prediction service can handle varying input formats and provide low-latency responses. What should you do?

Hard
266

A team is deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that scales numerical features using statistics computed from the training set. The preprocessing must be identical between training and serving. The team wants to minimize latency and ensure consistency. What should they do?

Hard
267

Your team trains a scikit-learn model locally and uploads it to Vertex AI Model Registry. A colleague needs to deploy it to a Vertex AI Endpoint for online prediction with a prebuilt container. The model artifacts are stored in a Cloud Storage bucket. Which deployment approach should they use?

Medium
268

A large organization uses a multi-project setup with a central data lake. Different teams manage their own models. To enable cross-team sharing of features, they want to use Vertex AI Feature Store. What is the best practice to manage access?

Hard
269

Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?

Easy
270

A retail company wants to build a product recommendation system using BigQuery ML for their e-commerce platform. The data includes customer purchase history, product metadata, and clickstream logs. The ML engineer needs to minimize manual feature engineering and leverage pre-built solutions. Which approach should the engineer take?

Medium
271

A team uses Vertex AI Pipelines with a custom training component that reads data from a BigQuery table. They need to ensure that a new pipeline run uses a specific snapshot of the training data for reproducibility. Which approach should they take?

Medium
272

A machine learning engineer notices that the online prediction latency for a custom TensorFlow model deployed on Vertex AI has increased significantly over the past week. Cloud Monitoring shows that the CPU utilization of the endpoints remains below 40%, but the number of concurrent requests has doubled. What is the most likely cause of the latency increase?

Medium
273

You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?

Medium
274

You have a TensorFlow training script that runs on a single machine. To speed up training on Vertex AI with 8 GPUs on a single machine, which strategy should you use?

Easy
275

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a custom component for data validation. The component takes a dataset URI and outputs a validation report. The engineer wants to fail the pipeline immediately if the validation report indicates that the data is invalid, without running subsequent steps. How should the engineer implement this?

Hard
276

An organisation wants to monitor fairness of their loan approval model across demographic subgroups. They have predictions stored in BigQuery along with ground truth. Which GCP service can evaluate model performance for each subgroup and identify disparities?

Medium
277

A retail company wants to forecast daily sales for each of its 500 stores for the next 90 days. They have three years of historical daily sales data stored in BigQuery, including promotions, holidays, and store attributes. The data science team has minimal ML expertise and wants to use SQL to build and deploy the model with minimal coding. Which approach should they use?

Medium
278

What is the purpose of the 'importer' component in Vertex AI Pipelines?

Easy
279

A company wants to classify customer support emails into categories like 'billing', 'technical', or 'account'. They have labeled email text data. Which AutoML solution should they use?

Easy
280

You are training a scikit-learn random forest on a dataset that fits in memory using Vertex AI custom training. The prototype notebook took 15 minutes, but the Vertex AI job takes over an hour and occasionally fails with a resource error. You want the production job to complete reliably without changing the model or preprocessing. What should you do?

Medium
281

In a Vertex AI Pipeline, a component produces a Metrics artifact that includes an evaluation metric. The engineer wants to use this metric value as a condition to decide whether to deploy the model. However, the metric value is stored in the artifact's metadata and not directly as a pipeline parameter. How can the engineer pass the metric value to a downstream conditional task?

Hard
282

You are deploying a prototype model to Vertex AI for online prediction. The model was trained with a custom preprocessing step that must run on raw JSON input before inference. You need the endpoint to return predictions with minimal latency and to support rolling updates of new model versions without downtime. (Choose two.)

Medium
283

An ML engineer is monitoring a Vertex AI Feature Store used for online serving. Which metrics are most important to track for ensuring low-latency online serving?

Easy
284

A hospital's radiology department wants to build a model that flags possible pneumonia on chest X-rays. They have 8,000 labeled DICOM studies in a Cloud Storage bucket and no in-house data science staff. They need a managed service that can ingest the images, train a classifier, and provide an endpoint for their viewing software. What should they do?

Medium
285

A data science team uses Vertex AI Workbench and wants to share notebooks with version history. Which service should they use?

Easy
286

An organization is deploying a mission-critical model on Vertex AI Endpoints. They need to ensure high availability and meet a strict SLO of 99.9% uptime. Which THREE steps should they take? (Choose 3)

Hard
287

Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?

Easy
288

A data scientist has trained a model using Vertex AI Training and wants to deploy it to a Vertex AI Endpoint for online predictions. Which orchestration service should be used to automate the deployment step after training completes?

Easy
289

Your team is preparing to hand a trained model to a separate operations team that will deploy it to a Vertex AI endpoint. The operations team needs to understand the model's input schema, the training run that produced it, and which alias currently points to production. Which two Vertex AI resources should you share with them to provide this information? (Choose two.)

Medium
290

A machine learning team uses Vertex AI Pipelines to orchestrate training workflows. They want to share pipeline runs and artifacts with stakeholders who do not have Google Cloud accounts. What should they do?

Easy
291

A marketing team wants to build a model that predicts customer lifetime value (CLV) using historical transaction data. They are comfortable with spreadsheets but have no coding experience. They need a low-code solution that automatically handles feature engineering and model selection. Which Google Cloud service should they use?

Easy
292

A marketing team wants to use a pre-built natural language processing (NLP) model from Vertex AI Model Garden to analyze customer feedback. They need to extract sentiment from text data stored in Cloud Storage. The team has no experience with model serving infrastructure. Which deployment option minimizes operational overhead?

Easy
293

An organization wants to implement continuous training for a model that serves predictions via Vertex AI Endpoints. Which approach best automates the retrain-deploy cycle?

Easy
294

Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)

Easy
295

A user receives the error "Deployment failed due to insufficient memory. Please use a machine type with higher memory." when deploying an AutoML model. What should they do?

Easy
296

You are deploying a custom model to Vertex AI for online prediction. The model requires a preprocessing step that normalizes input features. You want to ensure that the same preprocessing is applied during both training and serving to avoid training-serving skew. What should you do?

Medium
297

A data scientist is defining a Vertex AI pipeline and needs to include a step that imports a pre-existing model from Cloud Storage into the pipeline as an artifact. Which Kubeflow Pipelines SDK v2 component should they use?

Easy
298

A data science team uses TFX to train and deploy a model on Vertex AI. They want automated monitoring for pipeline health. Which set of metrics should they monitor to quickly detect issues in the training pipeline?

Medium
299

You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?

Hard
300

Which THREE actions are best practices for managing ML models in production on Google Cloud? (Choose 3)

Medium
301

You are training a TensorFlow model on Vertex AI using a custom container with a single Tesla T4 GPU. You notice that training is slower than expected, and GPU utilization is consistently below 20%. Profiling shows that the input pipeline is the bottleneck. Which change should you make to improve GPU utilization?

Hard
302

Which TWO metrics should you monitor to detect data drift in a batch prediction pipeline?

Medium
303

A machine learning team is deploying a PyTorch model on Vertex AI Prediction for real-time inference. The model was trained with preprocessing that includes tokenization and normalization. They want to embed the preprocessing logic in the model to reduce prediction latency and avoid additional service calls. Which approach should they take?

Hard
304

A company uses Vertex AI AutoML to train a vision model, but the model has low accuracy. What should they do first?

Medium
305

An ML engineer is using Vertex AI distributed training for a TensorFlow model that uses the MirroredStrategy. They notice that the training throughput drops significantly when moving from a single GPU to multiple GPUs on the same machine. What is the most likely cause?

Hard
306

You are using DVC for data versioning in an ML project on Google Cloud. Your training data is stored in Cloud Storage. You want to track a new version of the dataset after preprocessing. Which DVC command should you use to register the changes?

Medium
307

An ML engineer is running a Vertex AI pipeline that includes a data validation component and a training component. The engineer wants the pipeline to stop before training if data validation fails, but wants the validation component to record its result as an output artifact for later inspection. Which combination of pipeline features should the engineer use?

Hard
308

An ML team is optimizing an inference model for deployment on edge devices. They need to reduce the model size and improve latency while maintaining accuracy as much as possible. Which two techniques should they use? (Choose TWO.)

Medium
309

An ML team is scaling a prototype to production. The data pipeline currently reads from Cloud Storage and transforms data with a custom Python script. They need to handle higher throughput and add monitoring. Which approach should they take?

Medium
310

An ML engineer is setting up Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. They want to receive alerts when either feature skew or prediction drift exceeds a threshold. Which two configurations are required to enable these alerts? (Choose two.)

Medium
311

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced and that artifacts are tracked. Which two of the following practices should they follow? (Choose two.)

Medium
312

You have deployed a regression model that predicts house prices. Over the past month, the model's predictions have been consistently too high. You suspect data drift in the input features. Which monitoring metric should you prioritize to confirm this?

Medium
313

Match each Google Cloud AI/ML service to its primary purpose.

Medium
314

A data science team uses Vertex AI Experiments to compare multiple model training runs. They want to capture and compare hyperparameters, metrics, and code versions for each run. Which TWO steps should they take?

Medium
315

A team has trained a sentiment analysis model using PyTorch on Vertex AI Training. They now want to deploy it for online predictions with low latency. Which TWO actions should they take? (Choose 2)

Medium
316

A large e-commerce company deploys a recommendation model on Vertex AI with autoscaling enabled. During Black Friday, traffic spikes rapidly. The autoscaler adds new instances, but new instances take several minutes to become ready (cold start). As a result, many requests time out. What should they do to mitigate this issue?

Hard
317

A data scientist needs to scale a prototype deep learning model to train on a massive dataset using multiple GPUs. Which three strategies are essential for efficient distributed training? (Select THREE)

Medium
318

A machine learning engineer wants to use Vertex AI Vizier to tune three hyperparameters: learning rate (log scale), number of layers (integer), and optimizer (categorical). They have 50 parallel trials available. Which parameter specification types should they define?

Easy
319

An ML engineer needs to deploy a model to an endpoint and gradually shift traffic from the previous version (champion) to a new version (challenger) for A/B testing. How should they configure the endpoint?

Medium
320

A company is deploying a machine learning model for real-time inference on Vertex AI. Which TWO practices improve serving performance and reliability?

Easy
321

You have trained a scikit-learn model and saved it as a joblib file in Cloud Storage. You need to deploy this model to Vertex AI for online predictions with minimal effort. What should you do?

Easy
322

Which TWO are best practices for deploying models to Vertex AI Prediction? (Choose 2.)

Easy
323

An ML team is using Population Stability Index (PSI) to monitor feature drift on a Vertex AI Endpoint. The PSI value for a feature is 0.25, which exceeds the alert threshold of 0.2. The feature has high SHAP importance. The team wants to automatically retrain the model. What is the correct end-to-end setup?

Hard
324

A company uses Vertex AI Feature Store for feature engineering. They need to ensure point-in-time correctness to avoid data leakage during training. Which feature retrieval method should they use?

Hard
325

A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?

Medium
326

An ML engineer is building a continuous training pipeline that retrains a model when new data arrives. The pipeline should also detect skew between training and serving data. Which TWO Google Cloud services should they use? (Choose two.)

Medium
327

You have a prototype ML model that you want to scale to production on Vertex AI. The model is a Python function that performs simple data preprocessing and then calls a pre-trained scikit-learn model. You need to deploy this as a batch prediction job that runs weekly on a large dataset stored in BigQuery. What is the most efficient way to accomplish this?

Easy
328

An ML team wants to monitor their recommendation model for fairness. Which TWO metrics should they track to detect potential bias? (Select TWO.)

Easy
329

A company deploys a classification model on Vertex AI for loan approval. After a month, they notice the precision has dropped significantly. What should they do first?

Medium
330

A company needs to serve a model with strict latency requirements (<100ms). They are using Vertex AI Prediction with CPU. During testing, latency is 150ms. What should they do?

Easy
331

A data scientist needs to forecast daily sales for the next 30 days using historical sales data stored in BigQuery. They want to use BigQuery ML. Which model type should they choose?

Medium
332

A team is monitoring a model and observes that the error rate (prediction failures) has increased. They have enabled request/response logging on the Vertex AI Endpoint. How can they set up a metric and alert for prediction error rate?

Medium
333

An ML engineer has a prototype that trains a TensorFlow model on a single CPU machine using Vertex AI custom training. The job now needs to train on a larger dataset and must use multiple GPUs on one machine. The training script already uses tf.distribute.MirroredStrategy. What change is required to scale the job?

Easy
334

A company wants to monitor the cost of their Vertex AI prediction endpoint. They are charged per hour per replica and per request for GPU instances. Which approach should they use to track these costs?

Easy
335

A company uses Vertex AI Pipelines to train and deploy models. The pipeline has a step that runs a custom container. The step fails intermittently with a timeout error. Which approach should be taken to robustly handle this?

Hard
336

A company wants to predict customer churn using a dataset with 10,000 rows and 20 features. They have no ML expertise. Which low-code solution should they use?

Easy
337

A team is using Vertex AI Experiments to compare different hyperparameters. They want to automatically record the hyperparameters. What is the correct way?

Medium
338

A team monitors features in Vertex AI Feature Store for drift. They want to set up automated alerts when a feature's distribution deviates significantly from the baseline. Which feature monitoring configuration should they use?

Hard
339

A data scientist wants to use a pre-trained ResNet model from Keras Applications and fine-tune it on a small custom dataset. Which approach should they take to avoid overfitting?

Easy
340

Your team owns a Vertex AI Model Registry entry that other teams depend on for production serving. A new retrained model shows better offline metrics, and you need to roll it out gradually to a small percentage of live traffic while keeping the ability to revert instantly if quality degrades. What should you do?

Medium
341

Which THREE should be considered when setting up an automated retraining pipeline using Vertex AI Pipelines and Cloud Composer? (Choose THREE.)

Hard
342

You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

Easy
343

A company deploys a model on Vertex AI Endpoints for real-time inference. They notice latency spikes during peak hours. Which action is most effective to reduce latency without sacrificing accuracy?

Easy
344

A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?

Medium
345

A company runs a Vertex AI Pipeline that includes a custom component for hyperparameter tuning. The component uses a large search space and runs many trials. The pipeline is taking too long to complete, and the team wants to reduce the execution time without sacrificing model quality. They have already optimized the training code. Which Vertex AI Pipelines feature should they use to speed up the tuning component?

Hard
346

A team is using Vertex AI Model Registry to manage models. They need to ensure that when a new model version is registered, it is automatically evaluated for fairness and bias before being deployed. Which two Google Cloud services should they integrate to achieve this? (Choose two.)

Medium
347

An engineer is using TensorFlow Transform (tf.Transform) to preprocess training data. They want to ensure that the same preprocessing logic is applied during inference without code duplication. Which approach should they take?

Medium
348

A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?

Easy
349

An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?

Medium
350

A team wants to implement CI/CD for their ML models using Cloud Build. They have a pipeline that trains a model and deploys it. What is the best practice for triggering the pipeline when a new commit is pushed to the source repository?

Medium
351

A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?

Hard
352

You need to serve a large embedding model for similarity search with low latency. The model was trained to generate 256-dimensional embeddings. You plan to use Vertex AI Vector Search. Which index type should you choose to balance accuracy and performance for a dataset with 10 million vectors?

Medium
353

What is the most likely cause of the error?

Medium
354

You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?

Easy
355

You are deploying a scikit-learn model for online predictions. The model size is 200 MB. You want to minimize latency and cost. Which serving option should you choose?

Medium
356

A team uses Vertex AI Feature Store for storing features. They want to share feature definitions with other teams in a collaborative manner. What is the best way to collaborate on feature definitions?

Easy
357

You are deploying a model to a Vertex AI Endpoint and need to reduce inference latency for a real-time application. Which two actions should you take? (Choose two.)

Medium
358

You are training a scikit-learn random forest model on a large tabular dataset using a Vertex AI custom training job. The training script reads a CSV file from a Cloud Storage bucket and writes the trained model artifact to the same bucket. You need to ensure the training job can access the Cloud Storage bucket without embedding long-lived credentials in the container. What should you do?

Medium
359

A team deploys a PyTorch model on Vertex AI for online predictions. They notice that after deployment, the latency increases over time, especially during peak hours. The model is served using a custom container. What is the most likely cause?

Medium
360

A data scientist wants to track the lineage of a dataset used in a training run. Which Vertex AI feature should they use?

Easy
361

A company has multiple teams working on different models. They want to enforce consistent data preprocessing steps across all teams. Which approach should they take?

Hard
362

A company has deployed a model for image classification and wants to monitor for feature drift using XRAI attributions. However, they notice that the XRAI attribution maps are too large and are causing high latency in the monitoring pipeline. What is the most effective way to reduce the overhead of explainability monitoring for image models?

Hard
363

An organization wants to use Cloud Composer (Airflow) to orchestrate a machine learning workflow that includes running a Vertex AI Pipeline, followed by a BigQuery job, and then a Dataflow pipeline. What is the primary advantage of using Cloud Composer for this orchestration?

Easy
364

A financial institution wants to use Natural Language API for sentiment analysis on customer feedback, but the domain-specific language (e.g., 'bullish', 'bearish') is not correctly classified. They have 200 labeled examples. Which approach minimizes coding effort while improving accuracy?

Hard
365

A machine learning engineer is deploying a TensorFlow model on an edge device with limited memory and compute. The model needs to perform inference with low latency. The engineer has a trained float32 model. Which model compression technique should be applied first to reduce the model size and improve inference speed without significant accuracy loss?

Hard
366

You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?

Medium
367

Your team has deployed a model on Vertex AI endpoints and you are planning an A/B test to compare a new challenger model (v2) against the current champion (v1). The test should measure business metrics such as click-through rate. Which THREE steps should you take to set up the A/B test correctly? (Choose 3 correct answers)

Hard
368

You need to create a reproducible snapshot of a BigQuery table as of a specific timestamp for ML model training. The snapshot should be queryable without copying the entire dataset. Which BigQuery feature should you use?

Hard
369

A data science team deploys a regression model to predict house prices. After one month, the mean absolute error (MAE) on the serving data increases by 20% compared to the test set. Which monitoring strategy should the team implement first to diagnose the issue?

Easy
370

You are setting up feature monitoring in Vertex AI Feature Store to detect drift in a numerical feature. The monitoring job should run daily and alert if the Jensen-Shannon divergence exceeds 0.1. Which configuration should you use?

Medium
371

A data scientist uses Vertex AI Workbench to train a model and then deploys it to an endpoint. They want to automate the retraining and redeployment pipeline when new data arrives. Which service should they use?

Medium
372

A company needs to run batch predictions on 10 TB of data stored in Cloud Storage. The predictions should be written to BigQuery. Which approach should they use?

Medium
373

A media company wants to build a real-time recommendation system for articles. They have a large user base (10M+) and frequent updates to user interactions. They need to handle cold-start users and new articles. Which architecture on Vertex AI is most suitable?

Hard
374

What does the `ML.PREDICT` command do in BigQuery ML?

Easy
375

You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?

Medium
376

A data engineer wants to orchestrate a complex workflow that includes running a Vertex AI pipeline, then a BigQuery job, and finally a Dataflow pipeline. The workflow must handle dependencies, retries, and monitoring. Which Google Cloud service is most suitable for this orchestration?

Easy
377

An ML engineer needs to run batch predictions on 10 TB of data stored in BigQuery using a TensorFlow model. The predictions must be written to BigQuery. Which service should they use?

Medium
378

A team uses Vertex AI Workbench managed notebooks. They want to version control their notebook files and collaborate using Git. What is the best way to integrate Git?

Medium
379

A company has a TensorFlow model for image classification that must run on edge devices with limited memory. They need to reduce the model size without significant accuracy loss. Which technique should they use?

Medium
380

You want to use Vertex AI JumpStart to quickly deploy a pre-built foundation model for text summarization. Which action is required?

Easy
381

A company needs to perform sentiment analysis on streaming social media data. Which architecture should they use?

Medium
382

When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?

Easy
383

An ML engineer is preparing to train a large model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket as a set of TFRecord files. The engineer wants to optimize the training job to reduce cost and improve performance. Which two actions should the engineer take? (Choose two.)

Medium
384

A team uses Cloud Build to automatically trigger a Vertex AI pipeline when changes are pushed to the model code repository. They have a cloudbuild.yaml file that builds a container image and submits the pipeline. However, they want to run the pipeline only if the commit includes changes to the 'training/' directory. Which Cloud Build configuration option should be used to filter the trigger?

Medium
385

A data engineer wants to use BigQuery ML to train a model for predicting customer churn (binary classification) using a large dataset. They want the model to be automatically tuned. Which model type should they choose?

Medium
386

A financial institution needs to deploy a fraud detection model with strict latency <100ms per prediction and high throughput (1000 predictions/sec). The model is a deep neural network. Which architecture on Google Cloud meets these requirements?

Hard
387

A company has deployed a model to a Vertex AI Endpoint and wants to receive an alert when the model's prediction latency exceeds a threshold. They have configured Cloud Monitoring to track the endpoint's latency metrics. They now need to create a notification channel to send alerts to their on-call team. Which Cloud Monitoring resource should they use to define the condition that triggers the alert?

Easy
388

A data analyst wants to train a binary classification model on a BigQuery table without moving data out of BigQuery. They have limited ML expertise. Which approach should they take?

Easy
389

Which TWO of the following are low-code machine learning solutions on Google Cloud?

Easy
390

Which TWO actions are recommended for collaborating on machine learning models using Vertex AI Model Registry?

Medium
391

A company wants to implement continuous delivery (CD) for ML models, where a model is automatically deployed to a staging environment and only promoted to production after passing an evaluation gate. Which combination of GCP services is BEST suited for orchestrating this CD pipeline?

Medium
392

A company wants to classify support ticket text into categories. They have labeled historical tickets. Which Google Cloud service allows them to train a custom classification model with no code?

Easy
393

You are running a distributed training job on Vertex AI using PyTorch and the DistributedDataParallel (DDP) strategy across 4 nodes, each with 8 GPUs. You notice that the training loss is not decreasing as expected and the job occasionally hangs. You suspect a communication issue between nodes. Which of the following should you check first?

Hard
394

Refer to the exhibit. An alert policy is configured to trigger when prediction latency exceeds 500 ms for 5 consecutive minutes. The team is experiencing many false positive alerts during brief latency spikes. Which adjustment would most effectively reduce false positives while still detecting prolonged latency issues?

Hard
395

A company needs to serve a model for real-time predictions with a strict latency SLA of 100ms at the 99th percentile. The model is lightweight and traffic patterns are highly variable with occasional spikes. Which deployment strategy best meets the SLA while controlling cost?

Easy
396

A company uses Vertex AI for AutoML training. Which THREE are best practices for managing model versions?

Medium
397

A financial services company uses Document AI to process loan applications. They want to ensure that any documents the model cannot process with high confidence are reviewed by a human before finalizing the decision. Which Document AI feature should they enable?

Hard
398

A machine learning engineer is monitoring a deployed churn prediction model that has shown a gradual decline in accuracy over the past month. The engineer wants to diagnose the root cause of the performance degradation. Which TWO actions should the engineer take? (Choose two.)

Medium
399

You need to perform a large-scale feature computation on streaming data from Pub/Sub, transforming raw events into features, and writing results to Vertex AI Feature Store for online serving. Which Google Cloud architecture is most appropriate?

Hard
400

A financial services firm wants to predict loan default risk using a dataset with 30,000 labeled examples and 25 numeric and categorical features. Their team includes SQL analysts but no Python developers, and they want to minimize operational overhead. They decide to use BigQuery ML. Which model type should they use to achieve the best predictive performance while keeping the solution low-code?

Medium
401

A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?

Medium
402

A company needs to maintain an audit trail of model changes for compliance. Multiple teams will be updating models. What is the best approach to track who created, modified, or deployed each model version?

Hard
403

An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?

Medium
404

A company needs to forecast product demand for the next 12 months using historical sales data. They want to use BigQuery ML with minimal coding. Which model type is most suitable?

Medium
405

A company deploys a custom ML model on Vertex AI to predict customer churn. The model retrains weekly, and predictions are served via a Vertex AI endpoint. After a recent retraining, the monitoring dashboard shows a sudden increase in prediction requests but a decrease in predicted churn probabilities. The model's accuracy on the validation set remains stable. What is the most likely cause of the observed behavior?

Medium
406

A team wants to deploy two versions of a model (v1 and v2) on Vertex AI Endpoint to conduct an A/B test. They need to split traffic so that 10% of requests go to v2. Which configuration achieves this?

Medium
407

Which THREE practices improve collaboration when using Cloud Composer for ML pipelines?

Medium
408

Your team uses a Vertex AI Pipeline that reads from a BigQuery table, trains a model, and registers it. A teammate wants to know which BigQuery table snapshot was used for a specific registered model version so they can reproduce the training data exactly. Which action should they take?

Hard
409

You are an MLOps engineer at a retail company. Your team has deployed a demand forecasting model to a Vertex AI Endpoint. The model uses 12 numerical features and outputs a single numeric value representing predicted units sold. You have configured Vertex AI Model Monitoring with a training dataset that includes the full feature schema and prediction distribution. After a week, you observe that the feature 'promotion_flag' has a Jensen-Shannon divergence of 0.15, while all other features remain below 0.05. The model's prediction distribution has also shifted. Which action should you take first to diagnose the cause of the drift?

Medium
410

A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?

Medium
411

A recommendation system model is updated daily via a retraining pipeline. After each update, the online prediction latency increases significantly for about 30 minutes before returning to normal. What is the most likely cause and solution?

Hard
412

A team uses Vertex AI Pipelines to automate training and deployment. They need to ensure that only models that pass a set of quality checks (e.g., accuracy > 0.9, latency < 100ms) are deployed to production. How should they implement this?

Hard
413

A team is collaborating on a Vertex AI model using Vertex AI Model Registry. They need to ensure that model versions are properly managed and that deployments are reproducible. Which TWO practices should they follow? (Choose two.)

Hard
414

You are an ML engineer at a logistics company. The company uses a Vertex AI Pipeline with BigQuery ML to train a model that predicts delivery delays based on weather, traffic, and historical order data. The pipeline runs daily and includes steps: (1) data extraction from BigQuery, (2) feature engineering using Dataflow, (3) model training with BigQuery ML (logistic regression), (4) model evaluation, and (5) conditional deployment to a Vertex AI Endpoint if accuracy > 0.85. Recently, the pipeline has been failing at step 5 with the error: "Vertex AI Endpoint creation failed: Quota limit of 1 endpoint per region exceeded." The company has already created one endpoint in the same region for another model. The pipeline is configured to create a new endpoint each time a model is deployed. The engineer needs to fix this with minimal changes to the pipeline code. Which course of action should the engineer take?

Hard
415

An organization wants to trigger a Vertex AI pipeline whenever new data arrives in a Cloud Storage bucket. Which approach should they use?

Medium
416

A developer wants to add text translation to a mobile app. They need to translate user-generated content into multiple languages, and latency is critical. Which pre-built API should they use?

Easy
417

A company wants to build a model to predict housing prices using BigQuery ML. They have a dataset with features like area, number of bedrooms, and location. Which TWO model types are appropriate for this regression task?

Medium
418

A data scientist needs to retrieve training data from Vertex AI Feature Store that exactly matches the feature values as they were at a specific historical timestamp to avoid label leakage. Which feature view configuration should they use?

Medium
419

A team of data scientists and ML engineers is collaborating on a shared feature store in Vertex AI Feature Store. They need to ensure that feature definitions are versioned and that changes are reviewed before being used in production pipelines. Which TWO practices should they implement?

Medium
420

A logistics company uses Vertex AI AutoML Tables to predict delivery delays based on order attributes, weather data, and traffic data. The model is retrained weekly using a Vertex AI Pipeline that runs a BigQuery query to get training data, then triggers AutoML training. Recently, the pipeline fails with the error 'Dataset not found' when the AutoML training step starts. The BigQuery query runs successfully and outputs a table. Which is the most likely cause?

Hard
421

A data science team is deploying a PyTorch model for real-time inference using Vertex AI Endpoints. The model requires a custom container with specific CUDA drivers and Python packages. They have created a Docker image and pushed it to Artifact Registry. The pipeline should automatically retrain the model every week and deploy the new version if it passes validation. However, the deployment step fails intermittently with the error 'The container image is not compatible with the machine type.' What is the most likely cause?

Hard
422

You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?

Hard
423

A global e-commerce company uses BigQuery ML to forecast daily sales for 10,000 products. They use a time-series model with a horizon of 7 days. Recently, forecasts for a specific product category have been consistently too high. They suspect the model is not capturing a new seasonal pattern. Which action should they take first to diagnose the issue?

Hard
424

A team is using Delta Lake on Dataproc for their data lake with ACID transactions. They want to version data for ML experiments and roll back to a previous version if needed. Which Delta Lake feature should they use?

Medium
425

An ML engineer has a prototype scikit-learn model that must be served on Vertex AI. The model requires a custom preprocessing step that cannot be expressed in a scikit-learn Pipeline. The engineer wants to package the model with this preprocessing logic and deploy it to a Vertex AI Endpoint for online predictions. Which approach should they take?

Easy
426

A marketing agency uses Vertex AI AutoML Vision to classify social media images into brand logos and generic content. They have 5,000 images per class. The model achieves 95% accuracy on validation set, but in production it misclassifies many images that contain logos in unusual angles or lighting. They have limited ML expertise and want to improve robustness. Which action should they take?

Easy
427

A team is training a large TensorFlow model that requires more memory than a single GPU provides. They have access to multiple GPUs on a single machine. Which distributed training strategy should they use to split the model layers across GPUs?

Medium
428

A marketing team wants to build a model that predicts whether a customer will click on an ad, using a dataset in BigQuery. They have limited ML expertise and want to avoid writing complex code. They decide to use BigQuery ML with a logistic regression model. Which SQL statement should they use to create the model?

Easy
429

An ML pipeline must run a set of preprocessing tasks for each data shard in parallel. Which KFP SDK features should they use to implement this? (Choose two.)

Medium
430

What is the primary benefit of using a centralised model registry in MLOps?

Easy
431

Your ML pipeline uses Vertex AI Feature Store to serve features for online predictions. You need to monitor the freshness of features in the online store. Which approach is most effective?

Medium
432

A data science team wants to version control their datasets along with code using Git. They need a tool that integrates with Git and tracks changes to large data files. Which tool should they use?

Medium
433

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to use a low-code solution that can automatically handle feature engineering and model selection. They also need to deploy the model for online predictions with low latency. Which Google Cloud service should they use?

Hard
434

Which THREE actions should be taken to automate a machine learning pipeline using Cloud Build and Vertex AI?

Medium
435

A company uses Vertex AI Vector Search (Matching Engine) for a product recommendation system. The product embeddings are updated hourly. Which index update method should they use to ensure low latency for new items?

Medium
436

An ML team is deploying a model to Vertex AI for the first time. Which THREE are best practices for scaling from prototype to production?

Easy
437

A team has developed a prototype of a recommendation model using a small dataset on a single VM. They need to scale to a larger dataset for production training. They plan to use Vertex AI training with a custom container. What is the best practice for handling the increased data volume?

Easy
438

An ML engineer is designing a pipeline that should run only when new training data arrives in a Cloud Storage bucket. Which event-driven approach should they use to trigger the Vertex AI Pipeline?

Medium
439

Your company deploys batch prediction jobs using Vertex AI Batch Prediction. You need to monitor the jobs for failures and performance. What is the recommended approach?

Easy
440

A retail company has deployed a machine learning model using Vertex AI Endpoints to predict inventory demand. The model was trained on data from the past two years and has been in production for six months. The team has enabled Vertex AI Model Monitoring to track prediction drift with an alert threshold of 0.2. Last week, they received an alert that the prediction drift score reached 0.35, exceeding the threshold. The engineer checks the monitoring dashboard and sees that the distribution of predictions has shifted noticeably compared to the training data. The engineer also notices that the model's accuracy metrics, computed from weekly ground truth data, have remained within acceptable range. What should the engineer do first?

Easy
441

A company is using Vertex AI Vizier for hyperparameter tuning of a model with 5 integer hyperparameters, each with a range of 10-100. They have a budget of 50 trials and want to maximize the chance of finding the best configuration. Which Vizier algorithm should they use?

Medium
442

An ML engineer wants to containerize a custom training script and use it as a component in a Vertex AI Pipeline. The component should accept a dataset URI and a learning rate parameter, and output a trained model artifact. Which approach should the engineer use to define the component?

Medium
443

You are scaling a prototype recommendation model to production on Vertex AI. The model is trained with a custom container and uses a large embedding table that must be updated frequently. You need to serve predictions with low latency and also support online updates to the embeddings without redeploying the model. Which Vertex AI feature should you use?

Hard
444

A team deploys a model using Vertex AI and wants to monitor for concept drift. What should they track?

Medium
445

Which TWO are best practices for building ML pipelines on Vertex AI Pipelines?

Easy
446

You are deploying a scikit-learn model to Vertex AI for online prediction. The model expects a JSON payload with a single feature vector. You need to ensure the endpoint can handle bursts of traffic up to 1000 requests per second while maintaining low latency. What should you do?

Easy
447

Which TWO are best practices when deploying AutoML models to production?

Medium
448

A company is deploying a Vertex AI pipeline that trains a model and then runs a custom evaluation component. The evaluation component must only run if the training component succeeds and the model's accuracy exceeds a threshold. The pipeline must also support retries for transient errors in the training component. The engineer needs to configure the pipeline to meet these requirements. Which two actions should the engineer take? (Choose two.)

Hard
449

A team is using Vertex AI Pipelines to automate their ML workflow. They want to ensure that pipeline runs are reproducible and that artifacts are tracked. Which feature should they use?

Easy
450

An ML engineer needs to track the costs incurred by Vertex AI prediction endpoints. Which tool should they use to set budget alerts and monitor spending?

Easy
451

You need to set up monitoring for a Vertex AI model that serves predictions in real-time. The model is expected to have a latency SLA of under 100ms. Which metric should you configure an alert on to ensure the SLA is met?

Easy
452

An ML engineer is authoring a Vertex AI pipeline where a custom training component must read a dataset from a BigQuery table and write the trained model to a Cloud Storage bucket. The engineer wants the component to be reusable across projects and environments without hardcoding project IDs or bucket names. Which design should the engineer use?

Hard
453

A company has deployed a TensorFlow model on Vertex AI Prediction for real-time inference. They notice that during peak hours, the prediction latency increases significantly, and some requests time out. The model requires GPU acceleration. Which action should they take to reduce latency and avoid timeouts?

Easy
454

A data scientist has a TensorFlow 2.x model trained on a single GPU. They want to scale training to multiple GPUs on a single Vertex AI machine without code changes. Which strategy should they use?

Medium
455

A startup wants to add sentiment analysis to their customer feedback app without any labeled data or custom model training. Which Google Cloud service should they use?

Easy
456

A machine learning engineer wants to deploy a trained model to Vertex AI for online predictions. Which Vertex AI resource is required to serve the model and provide an endpoint URL?

Easy
457

An ML engineer is building a Vertex AI Pipeline that includes a data validation component. The component should fail the pipeline if the input data does not meet certain statistical thresholds. The engineer wants to ensure that the pipeline stops immediately and does not proceed to training if validation fails. Which mechanism should the engineer use in the component?

Medium
458

A startup is deploying a scikit-learn model to Vertex AI for online predictions. They want to minimize the effort required to containerize the model and ensure it can handle HTTP requests. What should they do?

Easy
459

Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

Easy
460

A team is using Vertex AI Model Monitoring to detect prediction drift on a deployed model. They have configured the monitoring job to run every 6 hours. After a week, they notice that the drift metric has been consistently high but no alerts have been triggered. They have verified that the alerting policy in Cloud Monitoring is correctly configured. What is the most likely cause?

Medium
461

The exhibit shows part of a Vertex AI Pipeline definition. The pipeline fails at the training step with an error: 'Missing required input: train_data'. What is the most likely cause?

Medium
462

A data science team is using Vertex AI Pipelines to orchestrate their ML workflows. They want to ensure that each pipeline run is reproducible and that artifacts are versioned. Which Vertex AI feature should they use to track and manage pipeline artifacts?

Easy
463

Refer to the exhibit. What is the purpose of this query?

Medium
464

A financial services company wants to extract text and structured data from scanned loan application forms. They need a fully managed, low-code solution that can handle various form layouts and requires minimal machine learning expertise. Which Google Cloud service should they use?

Easy
465

Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?

Easy
466

You have a Python training script that reads a 500 GB CSV dataset from a Cloud Storage bucket. You submit a Vertex AI custom training job using a pre-built container, specifying a machine with 16 vCPUs and 60 GB RAM. The job fails after a few minutes with an out-of-memory error. You need to scale the prototype to handle this dataset without changing the model architecture. What should you do?

Medium
467

A company has a large dataset of labeled images (e.g., different species of plants). They want to train a custom image classification model with minimal effort and no prior ML experience. Which Google Cloud service should they use?

Medium
468

An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?

Hard
469

A data analyst wants to build a binary classification model using a low-code ML solution on Google Cloud. The dataset is stored in BigQuery and contains 500,000 rows with 20 features, including categorical and numerical columns. The analyst has minimal coding experience and needs to deploy the model as an API endpoint for real-time predictions. Which two Google Cloud services should the analyst use to accomplish this task with minimal code? Choose two options.

Easy
470

A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?

Easy
471

A company uses Cloud Composer to orchestrate their ML pipelines. They notice that tasks are being queued but not executed, causing delays. What is the most likely cause?

Easy
472

You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?

Easy
473

Which THREE actions should be taken to manage model versions effectively?

Hard
474

A data scientist has deployed a classification model on a Vertex AI Endpoint and wants to monitor for feature drift in the serving data compared to the training data. Which Vertex AI service should be used?

Easy
475

You are preparing a Vertex AI custom training job that uses a custom Docker image built on top of the PyTorch pre-built training container. The image must be able to read training data from a Cloud Storage bucket without embedding credentials in the image. You want the job to run on a single NVIDIA T4 GPU. Which approach should you take?

Medium
476

A company serves a PyTorch model using a custom container on Vertex AI Prediction. They notice that after a few hours, the endpoint returns 502 errors. The logs show 'Out of memory' errors. The container has a memory limit of 4GB, and the model loads a 3GB vocabulary file. What is the most likely cause and best fix?

Hard
477

You are fine-tuning a pre-trained BERT model from Hugging Face on a custom text classification dataset using Vertex AI Training. You want to speed up training by using mixed precision. What should you do?

Medium
478

Your team manages multiple ML models on Vertex AI. You need to implement a centralized monitoring solution to track model performance over time. Which TWO approaches should you consider? (Choose two.)

Medium
479

An ML engineer is troubleshooting why a Vertex AI Endpoint is returning high prediction latency. They have enabled request/response logging and see that some requests take >1 second while most are fast. Which THREE actions should they take to diagnose the issue?

Hard
480

A financial services firm wants relationship managers to query a natural-language interface such as 'show me customers likely to churn next quarter' and receive results from their BigQuery data warehouse. They want to minimize the code they maintain and keep data in BigQuery. Which approach best fits?

Hard
481

A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?

Medium
482

An organization needs to implement MLOps with standardized pipeline templates across multiple teams. Which Vertex AI feature should they use to create reusable pipeline components?

Hard
483

A financial services company uses Vertex AI Pipelines to train and deploy models for fraud detection. The ML team consists of data scientists who develop models and ML engineers who deploy them. They use a CI/CD pipeline with Cloud Build to build and push Docker images to Artifact Registry, then trigger Vertex AI Pipelines. Recently, the team noticed that a model deployed to production was trained on a dataset that had not been approved by the data governance team. Upon investigation, they found that a data scientist accidentally used an unapproved version of the training data by specifying a Cloud Storage path that was not the latest approved dataset. The company needs to enforce that only approved datasets are used in training jobs. Which approach should they take?

Hard
484

A team of ML engineers is collaborating on a project using Vertex AI. They want to ensure that only approved models are deployed to production. Which approach should they use?

Medium
485

A company deploys an online prediction model serving 100 requests per second. They are optimizing for both latency and throughput. Which monitoring strategy should they use?

Easy
486

An ML engineer is using Vertex AI Model Monitoring on a deployed model that predicts customer churn. The model's input features include a mix of numerical and categorical features. The engineer wants to detect changes in the distribution of a specific categorical feature, 'contract_type', which has 20 possible values. Which approach should be used to effectively monitor drift for this feature?

Hard
487

Which TWO of the following are best practices for managing data in a collaborative machine learning environment on Google Cloud?

Medium
488

A company deploys a model on Vertex AI Prediction with autoscaling enabled. They notice that during a traffic spike, new instances take several minutes to become available, causing high latency. What is the best solution?

Medium
489

A company is using Vertex AI Pipelines to orchestrate a training workflow. They want to implement a CI/CD process where the pipeline is automatically triggered when a new version of the training code is pushed to a GitHub repository. They also want to ensure that the pipeline uses the latest code. Which two actions should they take? (Choose two.)

Medium
490

Drag and drop the steps to create and deploy a custom ML model on Vertex AI using a container in the correct order.

Medium
491

A team wants to implement continuous training for their ML model. The pipeline should be triggered when new training data arrives in a Cloud Storage bucket. Which combination of services should they use?

Medium
492

A company wants to monitor fairness of a model by evaluating performance metrics across demographic subgroups. They have ground truth labels stored in BigQuery. Which Vertex AI service should they use?

Easy
493

A team uses Vertex AI Workbench notebooks for collaborative model development. They want to ensure that code changes are version-controlled, that multiple data scientists can work on the same notebook without conflicts, and that the environment is reproducible across team members. Which approach should they take?

Medium
494

An ML team wants to implement data versioning for large datasets stored in Google Cloud Storage. They need to track changes over time and reproduce previous data states. Which tool is most appropriate?

Medium
495

You need to run a batch prediction job on Vertex AI for a large dataset stored in BigQuery. The model expects CSV input. Which input format should you specify for the batch prediction job?

Easy
496

A team has a prototype image classification model trained on a small dataset using TensorFlow Keras on a single GPU. They need to train on a larger dataset (1 million images) using a distributed strategy on Vertex AI with 8 GPUs. They implement a MirroredStrategy for data parallelism. During the first few epochs, the training speed does not improve significantly compared to a single GPU, and GPU utilization is low. The data is stored as JPEG files in Cloud Storage, and the input pipeline uses tf.data with map to decode images. What is the most likely cause?

Medium
497

A company uses Vertex AI Predictions with a custom container that invokes an external API for feature enrichment. The prediction response time is highly variable. The engineer wants to monitor the external API's contribution to latency. What should the engineer do?

Hard
498

A company wants to build a product recommendation engine for their e-commerce website. They have historical purchase data and user interaction logs. They want a managed service that can quickly generate personalized recommendations without building custom models. Which service should they use?

Medium
499

A financial services company wants to detect fraudulent transactions in real-time. They have a trained XGBoost model that runs on a single Compute Engine instance. The current solution processes about 100 transactions per second, but they need to scale to 10,000 transactions per second. Which approach should they take?

Medium
500

An ML engineer is designing a Vertex AI Pipeline that includes a hyperparameter tuning step. The tuning step runs multiple trials and outputs the best model. The engineer wants to ensure that the pipeline can resume from the tuning step if it fails later, without re-running all the tuning trials. What should the engineer do?

Hard
501

Which TWO actions are best practices when scaling a prototype ML model to production in Google Cloud?

Easy
502

A team is building a fraud detection model that requires joining real-time transaction features with historical user features. They need to ensure that the training data does not use future information (data leakage). Which Vertex AI Feature Store capability should they use?

Medium
503

A team uses Vertex AI Pipelines and wants to track lineage of artifacts and executions. Which three resources should they use? (Choose three.)

Hard
504

An ML engineer is using Vertex AI Vizier to tune hyperparameters for a PyTorch model. They want to maximise the chance of finding the global optimum within a fixed trial budget of 50 trials. Which algorithm should they select?

Medium
505

A real-time recommendation model deployed on Vertex AI Endpoints is experiencing increased latency, especially during peak hours. The model is hosted on a single machine with 4 CPUs. Which set of actions should you take to diagnose and resolve the issue?

Hard
506

Your team is training a very large transformer model that does not fit on a single GPU. They are using Vertex AI custom training with PyTorch. Which distributed training approach should they use?

Hard
507

An MLOps team needs to automatically retrain a model when new training data becomes available. They use Vertex AI Pipelines. What is the recommended way to trigger the pipeline?

Medium
508

Which TWO statements about Vertex AI Feature Store are correct? (Choose 2)

Easy
509

An ML team is using Vertex AI Online Prediction and wants to receive alerts when the 99th percentile latency exceeds 500ms for more than 5 minutes. What is the best practice to set up this alert in Cloud Monitoring?

Easy
510

A data engineering team needs to compute rolling window features (7-day average, 30-day sum) from a high-volume stream of e-commerce events stored in BigQuery. They must output the features to Vertex AI Feature Store for online serving. Which approach is MOST cost-effective and scalable?

Hard
511

You have a champion model serving 100% traffic on a Vertex AI endpoint. You want to deploy a challenger model and gradually shift 10% of traffic to it for A/B testing. What is the correct approach?

Medium
512

Which THREE are best practices for implementing CI/CD for ML pipelines on Google Cloud? (Choose THREE.)

Medium
513

A company uses BigQuery ML to train a boosted tree classifier on a large dataset. After training, they want to understand which features most influence predictions. Which BigQuery ML function should they use?

Hard
514

A data science team has deployed a model on Vertex AI and wants to automatically detect when the distribution of a specific feature shifts significantly from the training data. Which service should they use?

Easy
515

A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

Medium
516

You are deploying a custom PyTorch model to a Vertex AI Endpoint for real-time inference. The model artifact is stored in a Cloud Storage bucket. Your security team requires that the model be served from a container that runs as a non-root user and has no network access except to the Vertex AI prediction service. Which deployment configuration should you use?

Medium
517

An MLOps team is implementing a CI/CD pipeline for a TensorFlow model on Vertex AI. The model training job takes 2 hours and produces a SavedModel. The team wants to automatically trigger a new pipeline run whenever a change is pushed to the 'main' branch of their source repository. The pipeline should include training, evaluation, and if metrics exceed a threshold, deploy the model to a Vertex AI endpoint. Which trigger configuration should they use?

Medium
518

You have a trained XGBoost model that you want to deploy on Vertex AI for online prediction. The model expects input features in a specific order and requires a custom preprocessing step that normalizes numerical features using statistics computed during training. You need to ensure that the same preprocessing is applied at serving time. What should you do?

Medium
519

You are training a large tabular model on Vertex AI using a custom training job. The dataset is stored in BigQuery and is several terabytes in size. Training reads the same data for many epochs, and reading directly from BigQuery each epoch is slow and expensive. You want to maximize training throughput while keeping the data accessible to the training container. What should you do?

Medium
520

An ML team wants to use Vertex AI Hyperparameter Tuning to tune a custom training job. They have a budget of 50 trials and want to use an algorithm that balances exploration and exploitation. Which algorithm should they choose?

Easy
521

You want to use a pre-trained model from TensorFlow Hub for image classification, but you need to adapt it to classify your own custom categories with a small dataset. Which Vertex AI approach is most appropriate?

Easy
522

You want to reduce training costs by using preemptible VMs on Vertex AI for a fault-tolerant distributed training job that uses checkpointing. Which machine type should you choose in the worker pool configuration?

Medium
523

A company wants to use ML to predict customer churn. They have user activity logs in Cloud Storage, account data in BigQuery, and want an automated pipeline. Which pipeline architecture on Google Cloud should they use?

Hard
524

A retail team must run nightly batch predictions over 20 TB of Parquet data stored in Cloud Storage using a custom PyTorch model registered in Vertex AI Model Registry. They want the job to finish within a fixed maintenance window and prefer not to manage the underlying compute. Which configuration should they use?

Medium
525

After setting up model monitoring on Vertex AI for a classification model, the engineer sees a high number of anomaly alerts for the "age" feature. Upon investigation, the age distribution in recent predictions is similar to training data. What might be the cause?

Hard
526

A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?

Medium
527

An ML engineer is creating a Vertex AI Pipeline that includes a component to train a model. The component requires a machine type with a GPU. The engineer wants to specify the machine type and GPU type for the component. Which approach should the engineer use?

Easy
528

You are deploying a pre-trained BERT model for inference on edge devices. The model must be under 500 MB and inference latency under 50 ms. Which approach should you take?

Medium
529

You are using tf.Transform to preprocess data at scale. Which TWO services are required to run tf.Transform on Google Cloud? (Choose 2)

Medium
530

A company has a CI/CD pipeline that retrains a model every time new training data is available. They want to automatically deploy the new model to production only if it passes a set of evaluation tests on a staging environment. Which approach best implements this?

Hard
531

Your team is serving a large language model on Vertex AI using a custom container. The endpoint experiences intermittent 502 errors during traffic spikes. The autoscaling configuration uses a CPU utilization target of 60% and the model is deployed on n1-standard-4 instances. The model requires significant memory. Which combination of changes is most likely to resolve the issue?

Hard
532

A company deploys a model on Vertex AI Endpoint and expects high traffic spikes during promotional events. The current configuration uses manual scaling with 2 replicas. Which autoscaling configuration should they use to handle spikes while minimizing cost during normal traffic?

Medium
533

A company has deployed a model to a Vertex AI Endpoint and wants to receive an email notification whenever Vertex AI Model Monitoring detects feature drift above a configured threshold. They have already set up the monitoring configuration with a training dataset baseline. What should they do next to enable email alerts?

Easy
534

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built AutoML Tables component to train a classification model. The pipeline also includes a conditional step that deploys the model to an endpoint only if the evaluation metrics exceed a threshold. Which KFP feature should be used to implement the conditional deployment?

Hard
535

A team has deployed a model on Vertex AI Prediction and wants to monitor for data drift. Which TWO metrics should they use to detect drift in numerical features?

Easy
536

A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?

Medium
537

A team is responsible for monitoring the health of a Vertex AI pipeline that runs daily. Which THREE resources should they use to gain visibility into pipeline performance and failures? (Choose 3.)

Medium
538

A data science team uses a shared Cloud Storage bucket to store training datasets. They notice that some team members accidentally overwrite existing datasets, causing issues with reproducibility. Which approach best prevents accidental overwrites while maintaining collaboration?

Medium
539

A company wants to run batch predictions on millions of records stored in BigQuery. They need to preprocess the data (e.g., feature engineering) before feeding it to the model. Which approach is most scalable and cost-effective?

Medium
540

You are responsible for maintaining an ML pipeline that runs daily on Vertex AI Pipelines. The pipeline preprocesses data, trains a model, and deploys it to an endpoint. Recently, the pipeline has been failing at the deployment step because the endpoint already exists and the deploy step tries to create a new endpoint instead of updating the existing one. The pipeline code is written using the Kubeflow Pipelines SDK. You need to modify the pipeline to resolve this issue with minimal changes. What should you do?

Easy
541

A machine learning engineer is training a model on Vertex AI using a custom container. The training job uses a large dataset stored in BigQuery. The engineer wants to minimize data transfer costs and maximize training speed. Which of the following approaches is most efficient?

Hard
542

A travel booking company has a real-time recommendation system that suggests hotels and flights to users. The model is served using TensorFlow Serving on a Google Kubernetes Engine (GKE) cluster with auto-scaling enabled. The cluster uses n1-standard-4 machine types. The team has set up Cloud Monitoring dashboards and alerts. Last week, during a major holiday promotion, the team noticed that the model's inference latency P99 increased from 150 ms to 450 ms over a 30-minute period, while the request throughput increased from 500 to 1,200 requests per second. CPU utilization across the cluster rose to 95%, but memory utilization remained at 60%. The model version and the serving infrastructure configuration have not changed since the last deployment. Which action should the team take to mitigate the latency issue?

Hard
543

An MLOps engineer needs to collect ground truth labels for a deployed classification model to compare predictions against actuals. Where should the engineer store the ground truth data to enable Vertex AI model quality monitoring?

Easy
544

A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?

Medium
545

A fraud detection team trains a model on data stored in BigQuery. They want to ensure that the model can be reproduced exactly one year later, including the specific data version and training code. They use Vertex AI Pipelines for orchestration. Which practice should they implement?

Medium
546

A company has a prototype ML model that predicts equipment failure. They want to deploy it to production using Vertex AI. The model must be retrained weekly with new data. They also need to monitor for data drift and model performance. Which THREE components should they include in their MLOps pipeline? (Choose 3)

Hard
547

You want to use Vertex AI Vizier for hyperparameter tuning. You have 2 categorical parameters and 3 continuous parameters. Which algorithm is best suited for this mixed parameter space?

Easy
548

A team is scaling their prototype inference model to handle high-throughput requests with low latency. They use a custom container on Vertex AI Prediction. They notice that latency spikes occur under heavy load. What is the most effective strategy?

Hard
549

An engineer needs to compile a Kubeflow Pipeline defined in Python to a JSON format that can be run on Vertex AI Pipelines. Which command should they use?

Medium
550

A media company uses a Vertex AI Endpoint to serve a video recommendation model. They have enabled Vertex AI Model Monitoring for prediction drift. After a major news event, they observe a significant increase in prediction drift alerts, but the model's recommendations remain relevant and user engagement is stable. They want to reduce unnecessary alerts without losing the ability to detect true model degradation. What should they do?

Hard
551

A data science team wants to build a machine learning pipeline on Vertex AI Pipelines that preprocesses data, trains a model, and evaluates it. They need to ensure that components can be reused across multiple pipelines and that outputs from one component can be passed as inputs to another. Which approach should they take?

Medium
552

Match each feature engineering technique to its description.

Medium
553

A data engineering team wants to orchestrate an ML pipeline that includes data preprocessing in Dataflow, AutoML training, and model deployment. They want to minimize operational overhead. Which approach is best?

Hard
554

A team just moved a model from prototype to production using Vertex AI. They notice prediction errors for certain inputs that were not present in training data. What should they do to detect such issues automatically?

Easy
555

A retail company wants to build a recommendation system to show 'frequently bought together' items. Which Recommendations AI model type should they use?

Easy
556

An ML team wants to automatically retrain a model when data drift is detected. They have set up a Cloud Monitoring alert on drift. What service should they use to trigger a retraining pipeline in response to the alert?

Easy
557

A healthcare company has deployed a diagnostic model on Vertex AI Endpoints. They use Vertex AI Model Monitoring to detect drift in features and predictions. The model's input features include patient age, which is a numerical feature. The monitoring job has been running for a month and has generated several alerts for age drift. However, the model's performance has not degraded. The team wants to reduce false alerts without missing critical drifts. What should they do?

Medium
558

An ML engineer is training a PyTorch model on Vertex AI using a custom training job. The dataset is stored in a Cloud Storage bucket with 500,000 small JPEG images. The training job is configured with an n1-standard-8 machine and a single NVIDIA T4 GPU. The engineer observes that GPU utilization is very low (around 15%) and training is slow. The model code is not the bottleneck. What is the most likely cause and the best solution?

Medium
559

A team is building a batch prediction pipeline that processes raw data from Cloud Storage, performs complex preprocessing, and then runs predictions using a large model. The preprocessing step is compute-intensive and the prediction step is I/O-bound. Which TWO Google Cloud services should they combine to optimize cost and performance? (Choose 2)

Hard
560

Drag and drop the steps to set up a batch prediction job using Vertex AI in the correct order.

Medium
561

A machine learning engineer needs to create a pipeline that runs a custom container component on Vertex AI. The container expects a Cloud Storage path as input and outputs a model artifact. Which component type should they define using the Kubeflow Pipelines SDK v2?

Medium
562

A retail company wants to forecast daily sales for inventory planning. They have 3 years of historical sales data with clear weekly and yearly seasonality. Which approach should they use?

Easy
563

A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?

Medium
564

A data science team is configuring Vertex AI Model Monitoring for a deployed model. They want to detect both feature skew and feature drift. Which TWO configurations must they set?

Medium
565

Refer to the exhibit. The team notices that the pipeline fails to read data from the specified Cloud Storage path. What is the most likely issue?

Easy
566

Refer to the exhibit. An engineer notices no drift alerts but the model performance has degraded. What is the likely cause?

Hard
567

A team wants to serve a large PyTorch model (3 GB) for online predictions with low latency. Which THREE actions should they take?

Medium
568

A company runs a Vertex AI Pipeline that trains a model and then deploys it to a Vertex AI Endpoint. The pipeline uses a conditional deployment step based on the model's evaluation metric. The team wants to ensure that if the evaluation metric falls below a threshold, the pipeline fails and no deployment occurs. Which approach should they use?

Hard
569

A data science team wants to build a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it if the accuracy exceeds 0.9. They want to use the Kubeflow Pipelines SDK v2. Which construct allows them to conditionally execute the deployment step based on the evaluation metric?

Medium
570

Which TWO of the following can be used as input sources for Vertex AI batch prediction jobs? (Choose 2)

Easy
571

A company deploys a model on Vertex AI Prediction for real-time inference. Users report intermittent high latency during peak hours. The model is deployed on a single machine type with `min_replica_count=1` and `max_replica_count=5`. Autoscaling is enabled based on CPU utilization. What is the most likely cause of the latency spikes?

Easy
572

You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?

Medium
573

What is the primary benefit of using pipeline caching in Vertex AI Pipelines?

Easy
574

An ML team is building a feature pipeline with Dataflow that reads from BigQuery, computes features, and writes to Vertex AI Feature Store. They need to ensure that features are available for both training and serving with low latency. Which Feature Store option should they use?

Medium
575

You have deployed a model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with a monitoring frequency of every 24 hours. The model serves predictions with a feature called 'transaction_amount' that has a skewed distribution. After several days, you notice that the drift metrics for this feature are consistently below the threshold, but you suspect that the feature distribution has actually shifted. You want to improve the sensitivity of drift detection for this feature. What should you do?

Hard
576

A manufacturing company wants to predict equipment failure using sensor data stored in BigQuery. They have limited ML expertise and want to use AutoML Tables. The data includes timestamps, numerical sensor readings, and a boolean 'failure' column. The dataset is highly imbalanced with only 1% failure cases. Which of the following is the most effective approach to handle the imbalance in AutoML Tables?

Medium
577

A media streaming company uses a recommendation model deployed on Vertex AI Endpoints. The model predicts whether a user will click on a recommended item. They have set up Vertex AI Model Monitoring with a training dataset and configured drift detection for both features and predictions. Recently, they observed that the prediction drift metric (Jensen-Shannon divergence) for the 'click' prediction has exceeded the threshold, but feature drifts are within normal ranges. What is the most likely cause of this prediction drift?

Hard
578

A logistics company has a Vertex AI Model Monitoring job that detects feature drift on a deployed route optimization model. The model uses 50 features, and the monitoring job is configured to monitor all features. The MLOps team notices that the monitoring job is incurring high costs and taking a long time to complete. They want to reduce monitoring overhead while still detecting significant drift on the most important features. What should they do?

Hard
579

An engineer wants to use BigQuery ML to explain predictions from a trained boosted tree classifier for a specific set of input rows. Which function should they use?

Hard
580

You are designing a distributed training job for a very large neural network that does not fit on a single machine. You need to split the model across multiple devices. Which TWO techniques can you use?

Medium
581

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline must run on a schedule every day at 2:00 AM UTC. The engineer wants to use a fully managed Google Cloud service to trigger the pipeline. Which service should the engineer use?

Easy
582

Your model serving endpoint on Vertex AI is experiencing increased memory usage after a recent update. The model was converted from TensorFlow to TF Lite for faster inference. You notice that the endpoint's instances occasionally get killed due to out-of-memory (OOM) errors. What is the most likely cause?

Hard
583

You are deploying a scikit-learn model to a Vertex AI endpoint for real-time inference. Prediction requests arrive as JSON payloads containing a single instance per request, and the model's predict method expects a pandas DataFrame with named columns. You want to avoid writing a custom container. Which approach should you take?

Medium
584

You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?

Medium
585

You are defining a Python function component in KFP SDK v2. Which decorator should you use?

Easy
586

A retail company has a model deployed on a Vertex AI Endpoint that predicts customer lifetime value. The ML team wants to monitor the model for feature drift without setting up a full Vertex AI Model Monitoring job. They need a lightweight solution that compares live prediction requests to a reference distribution and sends alerts when drift exceeds a threshold. What should they do?

Easy
587

An engineer is training a model on Vertex AI using a custom container. The training job fails with an error indicating that the container exited with a non-zero status. The engineer wants to debug the issue. What is the best way to access the logs?

Medium
588

Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?

Hard
589

Drag and drop the steps to set up model monitoring for drift detection on Vertex AI in the correct order.

Medium
590

A team develops a pipeline that trains a model and evaluates it. They want to pass the test accuracy (a float) from the evaluation component to a subsequent deployment component. Which KFP SDK type should the evaluation component output be annotated with?

Medium
591

Drag and drop the steps to perform a hyperparameter tuning job on Vertex AI in the correct order.

Medium
592

An ML engineer trained a model and registered it in Vertex AI Model Registry. They want to assign the alias 'champion' to the best-performing version for production deployment. Which gcloud command should they use?

Hard
593

A company wants to use Vertex AI for hyperparameter tuning. Which three components are required to configure a hyperparameter tuning job? (Choose THREE.)

Easy
594

An ML engineer is using Vertex AI Training to fine-tune a large image classification model on a dataset stored in Cloud Storage. The training job uses a custom container and runs on a single NVIDIA V100 GPU. The engineer notices that GPU utilization is consistently low (around 20%) and training is slow. The data is stored as many small JPEG files. What should the engineer do to improve GPU utilization and training speed?

Medium
595

You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?

Medium
596

A company is building a document processing pipeline for invoices. They need to extract key fields (invoice number, date, total amount) and allow human review for invoices over $10,000. Which TWO Google Cloud services/features should they combine?

Hard
597

You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?

Easy
598

A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?

Easy
599

A company is serving a model for their e-commerce website. They expect traffic to be low at night and very high during flash sales. They want to minimize costs while ensuring availability during spikes. Which autoscaling configuration should they use?

Easy
600

A company has a prototype ML model that works well on historical data, but when deployed to production, the model performance degrades over time. The data distribution shifts gradually. Which strategy should they implement to maintain model accuracy?

Easy
601

A data analyst wants to train a linear regression model to predict house prices using only SQL queries on BigQuery. Which BigQuery ML model type should they use?

Easy
602

A developer needs to transcribe phone calls with high accuracy for a call center analytics application. The audio is in English and has background noise. Which Speech-to-Text model should they choose?

Medium
603

A company has deployed a fraud detection model on Vertex AI Prediction. After three months, the model's accuracy has degraded, and the business is losing money due to undetected fraud. What should the team implement to proactively detect such issues?

Easy
604

Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?

Medium
605

A company uses Vertex AI Pipelines to orchestrate ML workflows. After a pipeline run, they want to query the lineage of a particular model artifact to find out which dataset and hyperparameters were used to produce it. Which API method should they use?

Hard
606

An ML engineer is building a Vertex AI pipeline that includes a component to train a model. The component takes a long time to run and occasionally fails due to transient errors (e.g., network timeouts). The engineer wants to automatically retry the component a few times if it fails. How should the engineer configure this?

Medium
607

You need to run a custom training job on Vertex AI using a pre-built container for scikit-learn. Which container image should you specify?

Easy
608

Match each ML acronym to its definition.

Medium
609

A machine learning team uses Vertex AI Pipelines for model training. They want to implement a conditional step that runs additional evaluation if the model accuracy exceeds 0.9, otherwise it runs a data augmentation component. Which two Kubeflow Pipelines SDK v2 constructs can they use to achieve this? (Choose two.)

Medium
610

You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?

Hard
611

A team is operationalizing a machine learning pipeline using Vertex AI. They want to automatically track experiment runs, log model parameters and metrics, and store model artifacts for reproducibility. They also need to capture lineage between pipeline components (e.g., which dataset and hyperparameter tuning job produced a model). Which TWO services should they use together to achieve this? (Choose two.)

Hard
612

Match each ML model interpretability method to its description.

Medium
613

An ML engineer has a model trained in Vertex AI and wants to deploy it to an endpoint with autoscaling and traffic splitting for canary testing. They have the model artifact stored in Vertex AI Model Registry with alias 'champion'. What is the correct sequence of steps?

Medium
614

You are training a PyTorch model on Vertex AI using a custom container. The training script uses DistributedDataParallel (DDP) with NCCL backend across 4 nodes, each with 8 GPUs. You notice that training throughput is low and GPUs are underutilized. After profiling, you find that the data loading is the bottleneck. You need to improve throughput without changing the model. What should you do?

Hard
615

You have a prototype model trained on a single machine using scikit-learn. You now need to scale training to a larger dataset that does not fit in memory on one machine. You want to use Vertex AI training with minimal changes to your existing scikit-learn code. What should you do?

Medium
616

A data scientist wants to use AutoML to classify images of retail products into categories. There are 50 categories and the dataset has 100,000 labelled images. Which Vertex AI AutoML service is most appropriate?

Medium
617

A team is building ML pipelines with Vertex AI. They want to reuse standard pipeline components across teams and enforce governance. What approach should they take?

Medium
618

A team deploys a real-time model using a custom container on Vertex AI Prediction. The container is large (5 GB) and cold starts are causing latency spikes. The endpoint is configured with `min_replica_count=0` to reduce cost. The team wants to keep the cost low while reducing cold starts. What is the best approach?

Hard
619

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?

Easy
620

Your organization uses Vertex AI Pipelines for training. A compliance auditor asks you to prove which dataset version and which preprocessing code commit produced a model that is currently deployed. You need to retrieve this information programmatically for a specific model version. Which approach should you use?

Hard
621

An ML engineer needs to trigger a Vertex AI Pipeline on a recurring schedule, every 24 hours, to retrain a model with the latest data. Which approach should they use to set up this schedule?

Easy
622

A retail company has deployed a demand forecasting model on a Vertex AI Endpoint. The model uses 20 numeric features. The MLOps team wants Vertex AI Model Monitoring to detect training-serving skew for each feature and receive alerts when skew exceeds a threshold. They have enabled Model Monitoring for the endpoint and configured a monitoring frequency of every 24 hours. However, after several days, no skew metrics appear in the Vertex AI console. What is the most likely cause?

Medium
623

A company trains a model using features from Vertex AI Feature Store. They notice training-serving skew because the feature values used at training time differ from those served online. How should they address this?

Hard
624

A data analyst wants to use low-code ML to analyze text data. Which TWO Google Cloud services are appropriate?

Easy
625

A machine learning engineer wants to manage multiple model versions and facilitate collaboration across teams. The goal is to track model lineage, versioning, and approvals. Which Vertex AI service should they use?

Easy
626

A team is implementing CI/CD for ML using Cloud Build. They want to trigger a training pipeline in Vertex AI whenever a new model code is pushed to the main branch of the repository. Which Cloud Build configuration should they use to achieve this?

Medium
627

You are building a CI/CD pipeline for an ML model using Cloud Build. When code is pushed to the main branch, you want to automatically build a training image, run a Vertex AI pipeline, and if the model evaluation passes, deploy it to a staging endpoint. Which two components are essential for this CI/CD pipeline?

Medium
628

You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?

Easy
629

An ML engineer is monitoring a Vertex AI Endpoint and notices a spike in 5xx error rates. Which TWO metrics should they examine to diagnose the issue? (Choose 2)

Easy
630

Your team trains a model on a Vertex AI Workbench notebook and logs hyperparameters, metrics, and a confusion matrix. Your manager asks you to ensure that anyone in the organization can reproduce the exact training run and compare it with other runs without manually digging through notebook cells. Which Vertex AI component should you use to record this information?

Medium
631

A company deploys a training pipeline on Vertex AI using custom containers. The pipeline includes a hyperparameter tuning job that uses Bayesian optimization. After several runs, they observe that the tuning job is not converging and the search space is large. They want to reduce the number of trials while still finding good hyperparameters. Which strategy should they use?

Hard
632

Which THREE of the following are recommended practices for model governance and lineage in Vertex AI?

Hard
633

Which TWO options are best practices for building ML pipelines on Vertex AI?

Easy
634

You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?

Medium
635

A data scientist needs to train a time-series forecasting model on historical sales data stored in BigQuery to predict future demand. The data has strong seasonal patterns. Which BigQuery ML model type should they use?

Easy
636

An ML engineer has set up Vertex AI Model Monitoring on an endpoint with a sampling rate of 0.1 (10%). They notice that the monitoring job runs hourly but the reported drift metrics seem inconsistent. What is the most likely cause?

Medium
637

You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?

Medium
638

Which TWO are best practices for implementing a low-code ML solution using Vertex AI AutoML? (Choose 2)

Hard
639

An e-commerce company uses a recommendation model that suggests products based on user browsing history. The model was trained on data from the past year and has high accuracy on the test set. However, after deployment, the click-through rate (CTR) on recommendations is much lower than expected. Which three steps should the data scientist take to diagnose and improve the model? (Choose THREE)

Hard
640

A team is training a large recommendation model on Vertex AI using a custom container. They need to log training metrics and visualize them in Vertex AI TensorBoard. The training code is written in PyTorch and runs on multiple worker nodes. Which of the following is the correct way to enable TensorBoard logging?

Hard
641

A company uses Vertex AI Matching Engine for a product recommendation system. They need to update the index with new product embeddings every hour, but the index is used for online queries with low latency. Which index update strategy should they use?

Hard
642

An organization wants to trigger a Vertex AI pipeline whenever a new commit is pushed to the main branch of their Cloud Source Repository. The pipeline should retrain and evaluate the model. Which service should they use to detect the push event and start the pipeline?

Medium
643

A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)

Hard
644

A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?

Hard
645

A team deploys a model on Vertex AI that uses a custom prediction routine (CPR) with a dependency on a native library. The container crashes with 'ImportError: libcudart.so.11.0: cannot open shared object file'. How should they resolve this?

Medium
646

A data scientist runs a BigQuery ML prediction query and gets a region mismatch error. The model is in the US region, but the new_data table is in the EU region. What is the simplest way to resolve this?

Easy
647

You are using Vertex AI Training to train a model and then automatically deploy the best candidate to a Vertex AI Prediction endpoint via the Vertex AI Model Registry. However, after deployment, you notice that the endpoint returns predictions for the new model, but they are significantly different from the evaluation metrics computed during training. The training scripts used TensorFlow with a serving input function. What is the most likely issue and how would you fix it?

Easy
648

Your team has deployed a text classification model on Vertex AI Endpoints. You notice that the model's latency has increased significantly over the last week, but the request rate has remained stable. Which of the following is the most likely cause?

Hard
649

You are collaborating on a Vertex AI Feature Store implementation. A data engineer updates a feature's values in the offline store, but the online store still serves the old values for several hours. The online store is configured with a feature value TTL of 24 hours and uses batch ingestion. What is the most likely cause of the stale online values?

Hard
650

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to build a custom model with minimal coding. They also need to integrate the model into an existing application that expects a REST API. Which Google Cloud service should they use to train and deploy the model with the least effort?

Hard
651

A company wants to build a recommendation system that suggests products to users based on their past interactions. They have user-item interaction data in BigQuery and want a low-code solution that can generate recommendations for all users. Which approach should they use?

Hard
652

Drag and drop the steps to set up a BigQuery ML linear regression model for forecasting in the correct order.

Medium
653

A company wants to implement a central model governance strategy using Vertex AI. They need to track model lineage, store evaluation metrics, and manage model versions across teams. Which THREE Vertex AI services should they use? (Choose 3)

Medium
654

A media company wants to automatically transcribe and analyze customer support calls to identify common issues. They need a low-code solution that provides both transcription and sentiment analysis. Which Google Cloud service should they use?

Medium
655

Which THREE factors are critical when designing a model serving architecture for a global user base with strict latency SLAs? (Choose 3.)

Hard
656

A data scientist has trained a scikit-learn model locally and wants to deploy it to Vertex AI for online predictions with low latency. The model is a small RandomForestClassifier (100 MB). What is the recommended way to deploy this model?

Easy
657

Refer to the exhibit. A data scientist notices that predictions from a deployed model are taking longer than expected. Which Cloud Monitoring metric should be inspected first to identify the bottleneck?

Easy
658

A team uses Vertex AI Metadata to track pipeline runs. They need to identify all artifacts that were generated by a particular pipeline execution. Which API method should they use?

Hard
659

A company deploys a TensorFlow model on Vertex AI Prediction with a single node. During peak hours, inference latency increases. What should they do first to reduce latency?

Easy
660

An ML engineer needs to monitor a deployed model for data drift. They want to compare the distribution of incoming predictions against a baseline distribution. Which Vertex AI service should they use?

Easy
661

A company runs a Vertex AI Pipeline that includes a hyperparameter tuning step followed by a training step. The tuning step outputs the best hyperparameters. The engineer wants the training step to use these hyperparameters and to ensure that the training step only runs if tuning succeeds. Which approach should the engineer take?

Hard
662

Your team trains models in a shared Vertex AI project, and multiple engineers run pipelines against the same BigQuery training tables. A reviewer needs to reproduce the exact dataset used to train a model six weeks ago, but the source tables have been overwritten many times since. Which BigQuery capability should you have used to make each training snapshot reproducible?

Medium
663

You are fine-tuning a BERT model from Hugging Face Transformers on Vertex AI. You want to minimise cost for a short experiment. Which compute configuration should you use?

Easy
664

You are training a TensorFlow model on Vertex AI using a custom container. The training job uses a single node with 4 GPUs and a global batch size of 1024. You notice that the training is slower than expected and GPU utilization is low. You suspect the input pipeline is the bottleneck. Which of the following should you do to improve training throughput?

Hard
665

Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?

Medium
666

A team is building a CI/CD pipeline for ML using Cloud Build. The pipeline trains a model and deploys it to Vertex AI. Recently, a change in the data processing step caused the model to be trained with a different data version, leading to a failed deployment because the model was invalid. How should the team prevent this in the future?

Hard
667

A team is monitoring a production ML system that includes multiple models and data processing pipelines. They want to set up a comprehensive alerting strategy that minimizes false positives while ensuring critical issues are promptly addressed. Which approach is the most effective?

Hard
668

Your team uses Vertex AI Pipelines to automate the training and deployment of a recommendation model. The pipeline includes a step that evaluates the model and only deploys it if the evaluation metric exceeds a threshold. You need to ensure that the pipeline's artifacts, including the evaluation metrics and the deployed model, are tracked and can be traced back to the pipeline run for auditing. What should you do?

Hard
669

A machine learning engineer is exporting a trained model from Vertex AI Training to the Model Registry. Which artifact should they upload as the model artifact?

Easy
670

A team wants to collect ground truth labels for their model deployed on Vertex AI Endpoint to perform model quality monitoring. They have a process that generates actual outcomes within 24 hours of prediction. What is the recommended approach for storing these labels?

Medium
671

A retail company wants to predict customer churn using historical purchase data stored in BigQuery. The data includes customer demographics, transaction history, and support interactions. The team is comfortable writing SQL and wants to avoid moving data to a separate environment. Which approach should they take?

Medium
672

Drag and drop the steps to set up data lineage tracking for ML pipelines using Vertex AI Experiments in the correct order.

Medium
673

Which THREE are key capabilities of Vertex AI Feature Store?

Medium
674

A data engineer wants to compute feature aggregates over a large dataset stored in BigQuery and write the results to Vertex AI Feature Store. The pipeline must handle both batch and streaming data. Which Google Cloud service should they use?

Easy
675

A team uses Vertex AI Pipelines for continuous training triggered by model drift. They want to monitor the pipeline execution cost and optimize resource usage. Which THREE metrics should they track? (Choose 3)

Hard
676

You are serving a model on a Vertex AI endpoint that requires a GPU. The model is used for interactive predictions with a strict latency SLO. You notice that during peak hours, some requests time out because the endpoint's autoscaler is slow to add GPU replicas. Which action should you take to meet the SLO?

Hard
677

An ML engineer is building a Vertex AI pipeline that trains a model and then evaluates it. The evaluation component must compare the new model's accuracy against a fixed threshold and, if the model passes, trigger a downstream deployment component. The engineer wants to avoid running the deployment component when the model fails. Which approach should the engineer use?

Medium
678

An ML engineer is training a TensorFlow model on Vertex AI using a custom training job with a single worker and multiple GPUs. The training script uses tf.distribute.MirroredStrategy. After a few epochs, the job fails with a NCCL timeout error. The engineer confirms the GPUs are healthy and the batch size is reasonable. What should they do to resolve the error?

Hard
679

A media company is serving a video recommendation model on a Vertex AI Endpoint. The model receives a mix of requests: some require only a few features, while others require many features from a feature store. The team wants to reduce average latency and cost without retraining the model. Which TWO strategies should they use? (Choose two.)

Medium
680

A team uses Vertex AI Pipelines. They need to ensure that only certain team members can deploy models to production. What is the best approach?

Medium
681

A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?

Easy
682

Your organization uses Vertex AI Model Registry to manage models. A data scientist has trained a new model version and wants to ensure that only approved models are deployed to production. You need to implement a workflow where a model must be reviewed and approved by a designated approver before it can be deployed. What should you do?

Hard
683

You are preparing to scale a prototype ML model to production on Vertex AI. The model is trained with a custom training job, and you want to ensure that the training is reproducible and that you can compare different runs. Which two practices should you follow? (Choose two.)

Medium
684

You are scaling a prototype ML model to production on Vertex AI. The model is a TensorFlow model that you want to train on a large dataset using distributed training across multiple nodes. You need to minimize training time and ensure the job can recover from node failures. Which approach should you take?

Medium
685

A machine learning engineer notices that the Vertex AI Prediction endpoint's error rate has increased over the past week. The model was retrained with new data and redeployed. Which step should the engineer take first to diagnose the issue?

Medium
686

Which TWO practices help ensure reproducible ML experiments?

Easy
687

A company is training a large neural network on Vertex AI and training jobs keep failing with 'Out of memory' errors. The VM uses a standard n1-standard-4 machine with 15 GB RAM. Which action should they take first?

Medium
688

A fraud-detection model is deployed on a Vertex AI endpoint and must respond within 30 ms for 95% of requests. During testing, the team sees that p95 latency is dominated by feature retrieval from an external online store, not by model inference. They want to reduce latency without retraining the model. What should they do first?

Hard
689

A retail company wants to build a customer churn prediction model using BigQuery ML. The data is stored in BigQuery tables and includes customer demographics, purchase history, and support interactions. The data scientist wants to experiment with different model types quickly without moving data to another environment. Which approach should they use?

Medium
690

A team is using Vertex AI Pipelines to orchestrate a multi-step ML workflow. They need to pass a large dataset (several terabytes) between two components: a preprocessing component and a training component. The preprocessing component outputs a preprocessed dataset that the training component consumes. The team wants to minimize data transfer time and cost. What is the most efficient way to pass the data between these components?

Hard
691

A team is building a feature pipeline for an ML model. They need to compute aggregate features over a sliding time window from streaming data. Which Google Cloud service is most appropriate for this task?

Easy
692

A retail company wants to forecast daily sales for the next 30 days based on historical sales data. They have two years of daily sales records with no missing values. They want to use a low-code solution on Google Cloud that automatically handles seasonality and trends. Which service should they use?

Easy
693

An organization has multiple ML pipelines running on Vertex AI. They want to centralize monitoring and alerting for pipeline failures, including root cause analysis. Which combination of services should they use?

Hard
694

A team is troubleshooting a Vertex AI Pipelines run that keeps failing at the model evaluation step. The pipeline includes steps: data preprocessing, training, evaluation, and deployment. Which THREE actions should they take to diagnose the issue?

Hard
695

You are using Vertex AI Matching Engine for similarity search. Your index has 10 million embeddings of 512 dimensions. The query latency requirement is under 10ms for 99th percentile. Which index type should you choose?

Medium
696

You are using Vertex AI Vector Search with an approximate nearest neighbor index. You need to update the index with new data every hour. The updates must be available for queries immediately. Which update method should you use?

Medium
697

A healthcare provider needs to extract structured information from incoming PDF forms (e.g., patient intake forms). They want to automate data extraction without writing custom models. Which Google Cloud service should they use?

Medium
698

Match each ML pipeline component to its description.

Medium
699

You are using Vertex AI Vector Search for a product recommendation system. Your index is updated with new embeddings every hour. To minimize query latency while keeping the index fresh, what should you do?

Hard
700

You are training a model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket in the us-central1 region, and the training job runs in the us-central1 region. You notice that the training job takes significantly longer than expected due to data loading. You want to improve data loading performance without changing the model architecture. What should you do?

Medium
701

You have a TensorFlow model that you want to deploy on Vertex AI for online prediction. The model requires a custom preprocessing step that transforms raw input features into the format expected by the model. You need to ensure that the same preprocessing is applied both during training and serving, and that it is maintained as part of the model artifact. What should you do?

Medium
702

An ML engineer manages a Vertex AI Endpoint serving a fraud detection model. Compliance requires that every prediction be logged with its input features for audit, but the team also wants to minimize storage costs. They decide to enable request-response logging on the endpoint. Which configuration should they use to meet both requirements?

Hard
703

Refer to the exhibit. A team deploys a model using Cloud Run. They notice that after scaling up, the new instances take about 90 seconds to become ready and serve requests. They want to reduce this startup time. Which configuration change is most likely to help?

Easy
704

A retail company wants to forecast monthly sales for each of its 500 stores using historical sales data. They have two years of daily sales data per store and want to use BigQuery ML to build a forecasting model. They need to account for seasonality and trends. Which BigQuery ML model type should they use?

Medium
705

A retail company wants to generate product recommendations on their website using Google Cloud. They have historical transaction data and need a managed service that provides personalized recommendations like 'frequently bought together'. Which service should they use?

Medium
706

An organisation uses Delta Lake on Dataproc to manage a data lake for ML training. They need ACID transactions for concurrent reads and writes. Which file format does Delta Lake use as the underlying storage?

Medium
707

An e-commerce company uses a recommendation model deployed on Vertex AI Endpoints. The model's latency increases gradually over two weeks, causing timeouts. The model is served using a custom container. What is the most likely root cause and corrective action?

Medium
708

A team deployed a model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring for skew detection. They notice that training-serving skew metrics are only produced when the endpoint receives traffic, but they want to ensure the skew is computed correctly even during periods of low traffic. Which configuration should they adjust to ensure skew detection remains statistically valid without generating excessive false positives?

Medium
709

Match each Google Cloud storage option to its best use case.

Medium
710

A team wants to enforce governance and compliance for all ML models across the organisation. They need a centralised repository that tracks model versions, deployment history, and evaluation metrics. Which service should they use?

Easy
711

You are fine-tuning a large language model (LLM) from Vertex AI Model Garden using a custom dataset. You need to minimize training cost while maintaining reasonable throughput. Which THREE strategies should you combine?

Hard
712

You need to run batch predictions on a large dataset stored in BigQuery using a Vertex AI model. The dataset contains 10 million rows, and each prediction takes about 100ms. You want to minimize cost and execution time. What should you do?

Medium
713

An MLOps engineer is setting up monitoring for a deployed model on Vertex AI Endpoints. Which TWO actions are required to enable Vertex AI Model Monitoring for feature skew and drift? (Choose two.)

Medium
714

A company uses Vertex AI Pipelines to train and deploy models. They want to automatically generate model documentation that includes model details, intended use, and evaluation results. What should they use?

Medium
715

A financial services firm serves a fraud-detection model on a Vertex AI endpoint that consumes features from a Vertex AI Feature Store online store. During a load test, prediction latency is acceptable, but the firm discovers that the model's feature values in production drift from the values used at training time because the training pipeline read from a BigQuery table with different transformation logic. The team wants the serving path to use the same feature definitions as training so online and offline values match. Which approach should they take?

Hard
716

An ML engineer is scaling a prototype to production using Vertex AI Pipelines. The pipeline includes data validation, preprocessing, training, and deployment steps. They want to ensure that the pipeline can be reproduced and audited. What is the best practice?

Medium
717

A data scientist wants to define a lightweight Python function component in Vertex AI Pipelines using Kubeflow Pipelines SDK v2. Which decorator should be applied to the function to make it a pipeline component?

Easy
718

A marketing team needs to build a model that predicts whether a customer will respond to a promotional email. They have a BigQuery table with 2 million rows and 30 features, and they want to avoid writing any Python code. They require an explainable model and the ability to generate predictions directly in SQL. Which approach should they use?

Medium
719

A company is experiencing high prediction costs on Vertex AI Endpoints. They want to monitor and optimize costs. Which THREE actions should they take? (Choose 3)

Hard
720

A machine learning engineer is training a TensorFlow model on Vertex AI using distributed training with the MultiWorkerMirroredStrategy. The training job uses 4 workers with 4 GPUs each. The engineer notices that the training is not scaling linearly. What is the most likely cause?

Medium
721

A research team is training a very large Transformer model that does not fit into the memory of a single GPU. They have access to multiple GPUs on a single machine and want to split the model layers across GPUs. Which distributed training strategy should they use?

Hard
722

A company uses Vertex AI Pipelines to train models on a daily schedule. The pipeline includes a component that runs a BigQuery query to extract features. The team wants to ensure that if the BigQuery component fails due to transient network errors, the pipeline automatically retries it. How can they configure retries in Vertex AI Pipelines?

Medium
723

Which API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints?

Easy
724

A model serving team is experiencing high latency in production. Which TWO actions should they take to diagnose the root cause? (Choose 2.)

Medium
725

You are deploying a model to a Vertex AI Endpoint that will serve predictions to a global user base. You want to minimize latency for users in different regions while ensuring high availability. What should you do?

Medium
726

A pipeline includes a component that produces a model artifact. The team wants to automatically detect skew between the training data distribution and the serving data distribution. Which three best practices should they implement? (Choose three.)

Hard
727

A company wants to log all prediction requests and responses from a Vertex AI Endpoint to BigQuery for auditing and debugging. How can they achieve this?

Easy
728

You have a prototype model trained on a small sample of data. You now want to train on the full dataset using Vertex AI, but the dataset is stored in BigQuery and is several terabytes. You need to minimize data movement and avoid exporting the full dataset to Cloud Storage. What should you do?

Medium
729

A model deployed on Vertex AI Prediction is returning high latency for real-time requests. The model is a small TensorFlow model. Which troubleshooting step should the team take first?

Medium
730

A data science team has trained a TensorFlow model on-premises using a large dataset. When they try to deploy the model to Vertex AI for online predictions, the deployed model fails to start with a ‘MemoryError’. The model artifact is 2 GB, and the machine type is n1-standard-4 (15 GB RAM). What is the most likely cause?

Hard
731

An ML engineer notices that predictions are taking longer than expected under moderate traffic. Reviewing the endpoint configuration, what is the most likely cause of the high latency?

Medium
732

An ML pipeline runs on Vertex AI and includes a component that uses a third-party library not available in the default Python environment. The team wants to avoid building a custom container image. Which approach should they use?

Hard
733

A company needs to extract text from scanned invoices and parse key fields like invoice number and total amount. Which Document AI processor should they use?

Easy
734

An organization wants to deploy a TensorFlow model on edge devices such as smartphones and IoT devices for offline inference. Which format should they export the model to?

Medium
735

A company has deployed a machine learning model that uses a large input tensor. They notice that the prediction latency varies significantly between requests of the same size. Cloud Monitoring shows that the serving endpoint's CPU utilization is consistently below 50%, but memory utilization fluctuates between 70% and 95%. What is the most likely cause?

Hard
736

A data science team wants to share engineered features across multiple projects while ensuring low-latency serving for online predictions. Which Google Cloud service should they use to store and serve these features?

Easy
737

You need to run a distributed training job on Vertex AI using TensorFlow with MirroredStrategy on a single machine with 4 GPUs. Which training configuration should you use?

Medium
738

An ML engineer is preparing to train a large recommendation model on Vertex AI. The model uses a custom training loop in PyTorch and requires a multi-node cluster with 8 A100 GPUs per node. The engineer wants to minimize training time and ensure the job can recover from a node failure without restarting from scratch. Which combination of Vertex AI features should the engineer use?

Hard
739

You are scaling a prototype ML model to production on Vertex AI. The model is trained with a custom training job and you want to ensure reproducibility and traceability of each training run. Which two practices should you implement? (Choose two.)

Medium
740

A company uses Vertex AI Model Monitoring. Which two configuration options can be set to reduce false positive drift alerts?

Medium
741

Match each MLOps practice to its description.

Medium
742

You need to query a Vertex AI Vector Search index for nearest neighbours. The index is deployed on an endpoint. Which API method should you use to perform the query?

Medium
743

You are deploying a model to a Vertex AI Endpoint that must serve predictions with a strict 50 ms latency SLA. The model is a custom container that loads a large model file from Cloud Storage at startup. You notice that the first few predictions after a new deployment are slow, and sometimes the endpoint scales up and the new replicas also have slow first predictions. What should you do to reduce this cold-start latency?

Hard
744

A Vertex AI Endpoint hosts a model that must serve predictions with a strict 99th percentile latency under 100 ms. The model is a large TensorFlow model that processes images. During load testing, you observe that p99 latency spikes to 300 ms when batch size exceeds 1. You need to meet the latency SLO while maintaining reasonable throughput. What should you do?

Hard
745

An organization wants to implement central governance for ML models across teams. Which TWO services should they use together to achieve model versioning, lineage, and deployment management? (Select 2)

Medium
746

A team is building a CI/CD pipeline for an ML model. They want to automatically trigger a Vertex AI pipeline for retraining whenever new training data arrives in a Cloud Storage bucket, but only if a specific Pub/Sub notification is published by a data ingestion process. Which approach meets these requirements with minimal operational overhead?

Hard
747

A company wants to build a recommendation system that suggests products to users based on their past purchase history. They have a large dataset of user-item interactions in BigQuery and want to use a low-code approach. They decide to use BigQuery ML's matrix factorization model. Which SQL statement correctly creates such a model?

Medium
748

A media company wants to transcribe audio files from customer support calls into text for analysis. The audio is in English with clear speech and no background noise. They want a quick solution with no ML model training. Which Google Cloud service should they use?

Easy
749

You have a custom model deployed on a Vertex AI endpoint that receives online prediction requests. The model expects input features in a specific order, but clients sometimes send features in a different order. You want to ensure that the endpoint consistently receives correctly ordered features without modifying every client. What should you do?

Medium
750

Which TWO actions are recommended to detect and mitigate data drift in a production ML system on Vertex AI?

Hard
751

An ML team trains a model using a dataset stored in a BigQuery table. They want to ensure that the exact data snapshot used for training is recorded and can be reproduced later for auditing. Which approach should they take?

Medium
752

A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?

Medium
753

A data analyst wants to use BigQuery ML to train a linear regression model (LINEAR_REG) to predict house prices. They have a table with features like square footage, number of bedrooms, and location. Which TWO statements about the training process are correct?

Easy
754

A team is using Vertex AI Explainability with a deployed model. They need to generate explanations for image classification predictions. Which explanation method should they configure in the ExplanationSpec?

Hard
755

A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?

Easy
756

A data engineer is troubleshooting a Vertex AI Endpoint that serves a large BERT model. After deployment, many prediction requests fail with 'Out of Memory' errors. The machine type is n1-standard-8 (30 GB memory) with no accelerator. Which action will most likely resolve the issue?

Hard
757

An ML engineer needs to monitor the error rate of prediction jobs on a Vertex AI Endpoint. Where can they view the number of failed prediction requests over time?

Easy
758

Your PyTorch training script uses DistributedDataParallel (DDP) across 4 vertices each with 4 GPUs (16 GPUs total). You submit a Vertex AI custom training job. How should you configure the worker pool spec?

Medium
759

An ML team is moving from a prototype Jupyter notebook to a production training pipeline. They want to ensure reproducibility. Which approach should they take?

Easy
760

Your team is using Vertex AI Pipelines to train a model weekly. You want to monitor the pipeline for failures and receive a notification when a pipeline run fails. You have configured the pipeline to send logs to Cloud Logging. What should you do to receive an alert on pipeline failure?

Medium
761

A data scientist trains an XGBoost model on Vertex AI with a custom container. The model performs well on a held-out test set but fails to generalize in production. They suspect data leakage between training and validation. What is the best practice to prevent this?

Medium
762

A hospital wants to build a system that automatically transcribes doctors' dictated notes into text and then identifies key medical terms such as diagnoses and medications. They have no ML expertise and want to use Google Cloud's pre-trained APIs. Which combination of services should they use?

Easy
763

A team wants to implement automated model documentation that captures training data, feature importance, evaluation metrics, and intended use. Which Vertex AI feature supports this?

Hard
764

A model deployed on Vertex AI Prediction repeatedly exits with code 137. What is the most likely cause?

Medium
765

A data scientist trained a model on a single GPU but needs to train on multiple GPUs for a larger dataset. They observe that training time does not decrease linearly with additional GPUs. Which common issue is most likely?

Medium
766

An ML engineer is troubleshooting a Vertex AI Model Monitoring setup on a deployed model. The monitoring configuration uses a training dataset baseline and monitors several numerical features. After several days, the engineer notices that drift scores are being computed, but no alerts have fired even though one feature's distribution has shifted dramatically. The monitoring configuration specifies a drift threshold of 0.3, and the observed drift score for that feature is 0.45. What is the most likely explanation?

Hard
767

A retail company wants to build a demand forecasting model for thousands of product SKUs. They have historical sales data in BigQuery and limited ML expertise. They want to minimize coding and automatically handle seasonality and promotions. Which approach should they use?

Easy
768

A team deploys a model using Vertex AI Endpoint with automatic scaling. They observe that during traffic spikes, new instances take a long time to become ready, causing high latency for some requests. What should they configure to reduce this startup time?

Medium
769

An engineer needs to perform sentiment analysis on customer reviews. They have a large volume of text and need a solution that requires minimal customisation. Which option is most efficient?

Medium
770

A retail company wants to predict customer churn using their transaction history and customer demographics. They have limited ML expertise and want to use a managed service on Google Cloud. Which service should they use?

Easy
771

A financial services company has deployed a credit risk ML model on Vertex AI. They want to monitor the model for fairness across demographic groups to ensure no biased outcomes. Which TWO actions should they take as best practices? (Choose TWO.)

Medium
772

A company uses BigQuery to store feature data for ML training. A data engineer notices that a Vertex AI Training job is failing with 'Access Denied' errors when reading from a BigQuery table. The training job uses a custom service account that has been granted the 'bigquery.dataViewer' role on the dataset. What is the most likely cause of the failure?

Medium
773

Which TWO statements are true about canary deployments for Vertex AI endpoints?

Hard
774

You are using TensorFlow Transform (tf.Transform) to preprocess data for a model that will be deployed on Vertex AI. What is the primary benefit of using tf.Transform over Dataflow alone?

Medium
775

A machine learning engineer wants to define a lightweight pipeline component that runs custom Python code without building a container image. Which KFP SDK feature should they use?

Easy

Frequently asked questions

What does the scenario questions domain cover on the PMLE exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 775 scenario questions questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.