Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 226–300

775 questions total · 11pages · All types, answers revealed

Page 3

Page 4 of 11

Page 5
226
MCQmedium

A data scientist needs to train a large PyTorch model on a custom dataset using Vertex AI. The training script expects data from Cloud Storage and uses GPU acceleration. Which option correctly configures a custom training job with a pre-built container for PyTorch and attaches a single NVIDIA V100 GPU?

A.Use a custom container built from PyTorch base image and specify accelerator_count=1 in the machine spec
B.Use the pre-built container 'us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-12:latest' and in worker_pool_specs set machine_type='n1-standard-4', accelerator_type='NVIDIA_TESLA_V100', accelerator_count=1
C.Use the AI Platform Training service with gcloud ai-platform jobs submit training and --scale-tier BASIC_GPU
D.Create a training pipeline with AutoML and select GPU runtime
AnswerB

The PyTorch GPU pre-built container supplies the training runtime, while worker_pool_specs declares machine_type, accelerator_type='NVIDIA_TESLA_V100' and accelerator_count=1, attaching exactly one V100. This satisfies both the custom PyTorch training and single-GPU constraints without building a custom image.

Why this answer

Vertex AI custom training jobs use worker_pool_specs to define the machine type, container image, and accelerator configuration. The correct configuration uses the pre-built PyTorch GPU container and specifies machine_type='n1-standard-4', accelerator_type='NVIDIA_TESLA_V100', and accelerator_count=1 in the worker pool spec. This is the documented Vertex AI pattern for attaching a single V100 GPU.

Exam trap

PMLE often tests the exact API field names and the distinction between Vertex AI custom training and legacy AI Platform Training — candidates pick 'custom container' or 'AI Platform Training' because they sound plausible, but only the worker_pool_specs configuration with the pre-built PyTorch GPU image and correct accelerator_type is valid on Vertex AI.

How to eliminate wrong answers

Option A is wrong because while a custom container is valid, the answer omits the required worker_pool_specs structure and the correct accelerator_type string — 'accelerator_count=1 in the machine spec' is not a valid Vertex AI API field; the field is accelerator_count inside worker_pool_specs.machine_spec. Option C is wrong because AI Platform Training (the legacy service) uses --scale-tier BASIC_GPU, which does not let you specify a V100 or a pre-built PyTorch container in the Vertex AI manner — it's a different, deprecated service. Option D is wrong because AutoML does not support custom PyTorch training scripts or GPU runtime selection; AutoML is for tabular/image/text models with managed training, not custom code.

227
MCQmedium

An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?

A.The model is not deployed correctly.
B.The endpoint is configured with the wrong machine type.
C.The CPU utilization is below the target threshold, so the autoscaler does not add replicas.
D.The endpoint is using GPU which cannot autoscale.
AnswerC

Vertex AI autoscaling adds replicas only when average CPU utilisation exceeds the 60% target; if observed utilisation stays below that threshold, the scaler holds at minReplicas=1 regardless of request volume. The spike therefore did not drive per-replica CPU high enough to trigger scale-out, making this the likely cause.

Why this answer

Vertex AI's autoscaler uses CPU utilization as a metric to decide when to add replicas. If the CPU utilization remains below the 60% target threshold, the autoscaler will not trigger scale-up, even during a traffic spike. The endpoint is configured with minReplicas=1 and maxReplicas=3, but without exceeding the target, it stays at the minimum.

Exam trap

The trap here is that candidates assume any traffic spike automatically triggers scaling, but Vertex AI's autoscaler only scales based on the configured metric (CPU utilization), not request volume directly.

How to eliminate wrong answers

Option A is wrong because the model being deployed correctly is unrelated to autoscaling behavior; a misdeployment would typically cause prediction failures or errors, not a failure to scale. Option B is wrong because the machine type affects performance and cost, but does not directly prevent the autoscaler from adding replicas when CPU utilization exceeds the target. Option D is wrong because GPU-enabled endpoints can autoscale; Vertex AI supports autoscaling for GPU instances, though GPU metrics may require custom configuration.

228
MCQeasy

A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?

A.Create a Vertex AI batch prediction job with BigQuery source and BigQuery destination.
B.Use Dataflow to process the data and call the model via Vertex AI online prediction.
C.Export BigQuery data to CSV in GCS, then create a batch prediction job with GCS source.
D.Create a Cloud Function to iterate over BigQuery rows and call the endpoint.
AnswerA

Native BigQuery source and destination lets Vertex AI read and write directly without exporting 50 TB to Cloud Storage, avoiding costly intermediate copies. This satisfies the efficiency constraint by keeping data in place and streaming results back to BigQuery.

Why this answer

Vertex AI batch prediction natively supports BigQuery as both input source and output destination, allowing the service to read the 50 TB directly from BigQuery and write predictions back without exporting data. This avoids data movement, leverages BigQuery's scalability, and is the most efficient, fully managed approach for large-scale batch inference with a custom container.

Exam trap

PMLE often tests whether candidates default to exporting data to GCS out of habit — the trap is missing that Vertex AI batch prediction supports BigQuery natively, making export steps unnecessary and inefficient.

How to eliminate wrong answers

Option B is wrong because using Dataflow to call online prediction introduces per-row network calls, latency, and cost overhead, and online prediction endpoints are not designed for 50 TB batch workloads. Option C is wrong because exporting 50 TB to CSV in GCS adds significant time, storage cost, and an unnecessary data movement step when BigQuery source is natively supported. Option D is wrong because a Cloud Function iterating over BigQuery rows is not scalable, has execution time and memory limits, and would be prohibitively slow and expensive for 50 TB.

229
MCQeasy

A data scientist wants to track machine learning experiments, including parameters, metrics, and artifacts, and compare runs. Which Vertex AI service should they use?

A.Vertex AI Metadata
B.Vertex AI Experiments
C.Vertex AI Feature Store
D.Vertex AI Model Registry
AnswerB

Vertex AI Experiments logs parameters, metrics, and artifacts for each run and provides comparison across runs within a experiment. This directly satisfies the need to track training metadata and evaluate multiple runs side by side.

Why this answer

Vertex AI Experiments is the service designed to track ML experiments, including parameters, metrics, and artifacts, and to compare runs. It integrates with Vertex AI Metadata to log experiment lineage and provides a UI for comparing runs. This directly matches the data scientist's requirement.

Exam trap

PMLE often tests the overlap between Vertex AI Experiments and Vertex AI Metadata — candidates may pick Metadata, but Experiments is the user-facing service for tracking and comparing runs.

How to eliminate wrong answers

Option A is wrong because Vertex AI Metadata stores metadata about artifacts and executions but is not the primary service for tracking and comparing experiment runs. Option C is wrong because Vertex AI Feature Store manages and serves ML features, not experiment tracking. Option D is wrong because Vertex AI Model Registry manages model versions and deployment, not experiment parameters and metrics.

230
Multi-Selectmedium

You are collaborating with a team of data scientists on a Vertex AI Workbench notebook that preprocesses data for a machine learning model. You need to ensure that all team members can work on the notebook simultaneously without overwriting each other's changes, and that the notebook's execution environment remains consistent across the team. (Choose two.)

Select 2 answers
A.Store the notebook in a Cloud Source Repositories repository and use Git for version control.
B.Use Vertex AI Workbench's managed notebooks, which automatically sync all user changes to a central repository without requiring Git.
C.Enable the notebook's built-in real-time collaboration feature, which allows multiple users to edit the same notebook simultaneously.
D.Schedule regular exports of the notebook to a shared Cloud Storage bucket, and have team members import the latest version before making changes.
E.Configure the notebook to use a custom container image stored in Artifact Registry, and share the image with the team.
AnswersA, E

Using Cloud Source Repositories with Git provides version control, allowing multiple team members to work on the notebook concurrently by branching and merging. It tracks changes and resolves conflicts, ensuring that no work is overwritten. This is a standard collaborative practice for code and notebooks, and it integrates with Vertex AI Workbench.

Why this answer

Git-based version control via Cloud Source Repositories enables concurrent work with conflict resolution, while a shared custom container image ensures a consistent environment. Together, they allow team members to collaborate effectively without overwriting changes and with identical dependencies. Other options either lack robust version control or misstate Vertex AI Workbench capabilities.

Exam trap

The trap here is assuming that Vertex AI Workbench has built-in real-time collaboration or automatic syncing, when in fact you must use external tools like Git and Artifact Registry for effective teamwork.

231
MCQeasy

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built Google Cloud Pipeline Components (GCPC) to train a custom model. Which component should the engineer use to submit a custom training job to Vertex AI?

A.HyperparameterTuningJob
B.CustomJob
C.BatchPredictionJob
D.ModelDeploy
AnswerB

The CustomJob component submits a custom training job to Vertex AI, letting the pipeline run bespoke training code within a managed job rather than a pre-built trainer. It satisfies the requirement to launch a custom training workload as a pipeline step.

Why this answer

The CustomJob component is the correct choice because it is the pre-built GCPC component specifically designed to submit a custom training job to Vertex AI. It allows the engineer to specify a custom container image or a Python training script, along with machine configuration and hyperparameters, directly within a Vertex AI pipeline. Other components serve different purposes, such as hyperparameter tuning, batch predictions, or model deployment.

Exam trap

The trap here is that candidates may confuse HyperparameterTuningJob with CustomJob because both involve training, but HyperparameterTuningJob is for multi-trial optimization, not a single training run, and ModelDeploy is a distractor that does not exist as a GCPC component.

How to eliminate wrong answers

Option A is wrong because HyperparameterTuningJob is used for optimizing hyperparameters across multiple trials, not for submitting a single custom training job. Option C is wrong because BatchPredictionJob is for running batch predictions on a trained model, not for training. Option D is wrong because ModelDeploy is not a standard GCPC component; the correct component for deploying a model to an endpoint is ModelDeployer or a similar deployment component, and ModelDeploy does not exist in the GCPC library.

232
MCQhard

A financial institution has deployed a fraud detection model on a Vertex AI Endpoint. The model uses both numerical and categorical features. They have enabled Vertex AI Model Monitoring with training data and configured drift thresholds. After several weeks, they notice that the model's recall has dropped significantly, but the overall prediction distribution remains stable. They suspect that a specific subgroup of transactions is being misclassified. Which approach should they use to identify the subgroup and diagnose the issue?

A.Use Vertex AI Model Monitoring's feature attribution analysis to identify which features contribute most to the misclassifications.
B.Export the serving data and predictions to BigQuery, then use slice-based evaluation to compare performance metrics across different subgroups.
C.Configure Vertex AI Model Monitoring to monitor the 'transaction_amount' feature with a lower threshold to detect subtle drifts that might affect recall.
D.Retrain the model immediately using the most recent data to improve recall across all subgroups.
AnswerB

Exporting serving data and predictions to BigQuery allows you to perform slice-based analysis, such as comparing recall across different subgroups defined by categorical features. This approach can identify which subgroup is experiencing performance degradation. Vertex AI Model Monitoring itself does not provide subgroup-level performance metrics, so this manual analysis is necessary to diagnose the issue.

Why this answer

Exporting serving data and predictions to BigQuery enables slice-based evaluation, which can reveal performance disparities across subgroups. Since Vertex AI Model Monitoring does not provide subgroup-level metrics, this manual analysis is necessary to identify the problematic subgroup and diagnose the recall drop.

Exam trap

The trap here is assuming that Vertex AI Model Monitoring automatically provides subgroup performance metrics, when in fact it focuses on drift and requires exporting data for deeper analysis.

233
MCQmedium

An ML engineer is designing a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it only if the evaluation metric meets a threshold. The pipeline must pass the evaluation metric from the evaluation component to a downstream conditional. Which Vertex AI Pipelines feature should the engineer use to implement this flow?

A.Use the Vertex AI Model Registry to compare the new model's metric with the deployed model's metric and let the registry automatically deploy if better.
B.Write the metric to a Cloud Storage file and use a Cloud Function to trigger a separate pipeline for deployment.
C.Use the output of the evaluation component as an input parameter to a condition in the pipeline definition, referencing the component's output parameter.
D.Store the metric in a Vertex AI Experiment run and then use the Experiment's metric in the condition.
AnswerC

Vertex AI Pipelines allows a component to expose outputs as parameters that can be consumed by downstream conditions. The evaluation component can output a metric value as a pipeline parameter, and the condition can compare this value to a threshold. This is the standard way to implement conditional logic based on runtime values in a pipeline.

Why this answer

In Vertex AI Pipelines, component outputs can be promoted to pipeline parameters, which can then be used in conditions to control downstream execution. The evaluation component should output the metric as a parameter, and the condition compares it to the threshold. This maintains a single, traceable pipeline run and avoids external services.

Exam trap

The trap here is assuming that Vertex AI Experiments or Model Registry can directly control pipeline flow based on metrics, when in fact they are for tracking and storage, not runtime decision-making.

234
MCQhard

Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?

A.Use a smaller quantized version of the model and serve it with a batch size of one.
B.Implement continuous batching and paged attention in the serving container.
C.Increase maxReplicaCount and rely on autoscaling to handle long prompts.
D.Enable response streaming and return tokens as they are generated.
AnswerB

Continuous batching allows the server to add new requests to an in-flight batch at each decoding step, improving GPU utilization and throughput. Paged attention manages the key-value cache in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together they lower per-token latency and increase the number of requests served per GPU, which directly addresses long-prompt inference cost and latency.

Why this answer

For autoregressive LLM serving, the key bottlenecks are GPU underutilization during decoding and memory fragmentation in the key-value cache. Continuous batching keeps the GPU busy by admitting new requests into the current batch, while paged attention stores the KV cache in fixed-size blocks to avoid fragmentation and support more concurrent sequences. These techniques reduce per-token latency and improve throughput, making them the right optimization for long prompts on a cost-conscious deployment.

Exam trap

The trap here is confusing perceived latency improvements from streaming with actual compute reductions, when the real gains come from batching and KV-cache memory management.

235
Multi-Selecthard

An ML engineer is configuring Vertex AI Model Monitoring for a deployed model that receives a mix of numerical and categorical features. The engineer wants to ensure that the monitoring job can detect both data drift and training-serving skew. Which two configurations are required to enable both types of detection? (Choose two.)

Select 2 answers
A.Enable skew detection and specify the training dataset used to train the model.
B.Set the monitoring objective to 'drift' and provide a training dataset for baseline.
C.Set the monitoring frequency to at least once per hour to capture both types of drift.
D.Enable drift detection and specify a baseline dataset (e.g., a sample of recent serving data).
E.Configure alerting via Cloud Monitoring to notify when either drift or skew exceeds thresholds.
AnswersA, D

Training-serving skew detection requires comparing the feature distributions of the training data to the serving data. Therefore, you must enable skew detection and provide the training dataset as the baseline. This allows Vertex AI Model Monitoring to compute the training distribution and detect any significant differences in serving data, which is essential for identifying skew.

Why this answer

To detect both data drift and training-serving skew in Vertex AI Model Monitoring, you must enable each detection type and provide the appropriate baseline. Skew detection requires the training dataset to compare against serving data. Drift detection requires a baseline dataset, often a sample of recent serving data.

These two configurations are independent and both necessary to achieve the desired monitoring.

Exam trap

The trap here is conflating drift and skew detection configurations, assuming that one baseline or objective can enable both, when each requires its own setup.

236
MCQmedium

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to deploy the model to an endpoint only if the evaluation metric exceeds a threshold defined at pipeline submission time. The threshold must be changeable without recompiling the pipeline. Which mechanism should the engineer use?

A.Set the threshold as an environment variable in the pipeline's service account and read it inside the deployment component.
B.Hardcode the threshold in the evaluation component and use a dsl.Condition on the evaluation component's output.
C.Define the threshold as a pipeline input parameter and use a dsl.Condition on that parameter to gate the deployment component.
D.Use a Vertex AI Model Registry alias and configure the endpoint to only serve models whose alias matches the threshold value.
AnswerC

A pipeline input parameter is part of the pipeline's runtime interface, so the threshold can be supplied at submission time without recompiling. Wrapping the deployment component in dsl.Condition on that parameter creates the conditional gate. This is exactly the pattern Vertex AI Pipelines supports for runtime-configurable branching.

Why this answer

Vertex AI Pipelines distinguishes compile-time structure from runtime parameters. Making the threshold a pipeline input parameter allows it to be supplied at submission time without recompiling, and dsl.Condition on that parameter creates the deployment gate. The other options either bake the threshold into the pipeline definition, misuse Model Registry aliases, or rely on an unsupported environment variable mechanism that the compiler cannot see.

Exam trap

The trap here is confusing a dsl.Condition on a component output with a condition on a runtime parameter, when only the latter allows the threshold to change without recompiling.

237
MCQmedium

A data science team wants to share a set of engineered features across multiple projects and teams to reduce training-serving skew and ensure consistency. They need low-latency serving (single-digit milliseconds) for online predictions and also need to retrieve historical feature values for training. Which approach should they take?

A.Use Vertex AI Feature Store to define features once, serve online predictions from the online store, and retrieve historical features from the offline store for training.
B.Create a shared BigQuery dataset where each team writes features; serve predictions by querying BigQuery synchronously.
C.Store features in Cloud Storage Parquet files and load them into BigQuery for training; serve predictions from a custom microservice that reads from Cloud Storage.
D.Use Cloud Memorystore (Redis) to store the latest feature values for low-latency serving; each team independently computes features and pushes to Redis.
AnswerA

Feature Store decouples feature engineering from consumption: one definition feeds both the low-latency online store for serving and the offline store for point-in-time historical retrieval, eliminating training-serving skew. Reusing a Cloud Storage copy cannot meet single-digit millisecond online latency.

Why this answer

Vertex AI Feature Store is a fully managed, purpose-built feature store that centralizes feature definitions and provides both an online store for low-latency serving and an offline store for historical retrieval. The online store is optimized for single-digit millisecond reads, while the offline store (backed by BigQuery) allows point-in-time correct feature retrieval for training, which directly addresses training-serving skew. By defining features once and reusing them across projects, it ensures consistency between training and serving.

Exam trap

PMLE often tests the misconception that a data warehouse like BigQuery can serve online predictions with low latency, or that a simple key-value store like Redis suffices for both online and offline needs, ignoring the need for point-in-time correct historical retrieval and centralized feature definitions.

How to eliminate wrong answers

Option B is wrong because BigQuery is not designed for synchronous low-latency online serving; querying it per prediction would introduce hundreds of milliseconds to seconds of latency and lacks a built-in online store. Option C is wrong because Cloud Storage is not a low-latency serving layer, and a custom microservice reading Parquet files would not provide the required single-digit millisecond performance or point-in-time correctness for training. Option D is wrong because Redis alone does not provide historical feature retrieval for training, and independent feature computation by each team reintroduces training-serving skew and lacks centralized governance.

238
MCQmedium

You are an ML engineer at a fintech company. You have a prototype credit risk model built using XGBoost that achieves high accuracy on historical data. The model is trained on a dataset with 500,000 rows and 50 features. The company wants to deploy this model to production to score loan applications in real-time. The production environment must handle a peak load of 100 requests per second with a latency under 200ms. You have decided to use Vertex AI for deployment. After deploying the model as a Vertex AI endpoint with a single n1-standard-4 machine, you notice that latency exceeds 500ms at peak load and some requests time out. You have verified that the model prediction itself (excluding network overhead) takes about 50ms on average. What should you do to meet the latency and throughput requirements?

A.Change the machine type to a GPU-accelerated machine like n1-standard-4 with a T4 GPU.
B.Prune the model to reduce size and improve prediction speed.
C.Enable autoscaling with a minimum of 2 replicas and use a larger machine type (e.g., n1-standard-8) to handle more concurrent requests.
D.Switch from online prediction to batch prediction using Vertex AI Batch Prediction.
AnswerC

Autoscaling to at least two replicas distributes the 100 requests per second across instances, while the larger n1-standard-8 machine provides more CPU for concurrent inference. Together these cut the 500ms latency below the 200ms target.

Why this answer

The latency bottleneck is not the model inference time (50ms) but the inability of a single n1-standard-4 machine to handle 100 concurrent requests per second without queuing. By enabling autoscaling with a minimum of 2 replicas and upgrading to n1-standard-8, you increase both the number of concurrent requests the endpoint can process and the CPU/memory resources per replica, reducing queue wait times and keeping total latency under 200ms. This directly addresses the throughput and latency requirements without changing the model or switching to batch processing.

Exam trap

The trap here is that candidates assume latency issues are always due to model inference speed (leading them to choose GPU or model pruning), when in fact the bottleneck is often the lack of horizontal scaling to handle concurrent requests under load.

How to eliminate wrong answers

Option A is wrong because adding a GPU (e.g., T4) does not reduce latency for XGBoost inference; XGBoost is CPU-optimized and GPU acceleration typically adds overhead for tree-based models, making latency worse. Option B is wrong because pruning the model (e.g., reducing tree depth or number of trees) would only marginally improve the 50ms prediction time, but the primary issue is queuing due to insufficient replicas to handle 100 requests per second, not the raw inference speed. Option D is wrong because batch prediction is designed for offline, asynchronous processing and cannot meet the real-time requirement of under 200ms latency per request; it also does not solve the concurrency problem for online scoring.

239
MCQmedium

An ML engineer has deployed a tabular binary classification model to a Vertex AI Endpoint. After enabling Vertex AI Model Monitoring with training-serving skew detection, the engineer notices that the feature 'customer_age' is flagged as skewed. The feature values at serving time appear to be shifted upward by about 10 years compared to the training data. Which of the following is the most likely root cause?

A.The endpoint is receiving malformed requests that contain incorrect age values.
B.The model's predictions are biased, causing the monitoring system to misinterpret the feature values.
C.The monitoring configuration uses an incorrect sampling rate, leading to a false positive alert.
D.The training dataset contained a different feature distribution than the serving data because the model was trained on historical data from an older customer base.
AnswerD

Training-serving skew occurs when the feature distribution in production differs from the training data. If the model was trained on an older customer base, the 'customer_age' distribution would naturally be lower than the current serving population, causing the skew alert. This is a classic example of data drift due to changing population.

Why this answer

Training-serving skew detection compares the distribution of feature values in production against the training baseline. A systematic upward shift in 'customer_age' indicates that the serving population is older than the training population. This is a common scenario when a model is trained on historical data and deployed later, as demographic shifts occur.

Exam trap

The trap here is assuming that model monitoring skew alerts are caused by model bias or data quality issues, when they actually detect distribution differences between training and serving data.

240
Multi-Selecthard

Which TWO actions should be taken to ensure reproducibility of ML experiments when collaborating across teams on Vertex AI?

Select 2 answers
A.Lock dependency versions in a container image used for training
B.Share notebooks via Colab Enterprise with real-time editing
C.Version control datasets using DVC or Vertex AI ML Metadata
D.Allow each team to use their own preferred environment
E.Always use random seeds for all random operations
AnswersA, C

Locking dependency versions inside a container image guarantees identical library and runtime environments across teams, eliminating drift from differing package versions. This directly satisfies the reproducibility constraint by ensuring every training run executes against the same immutable image, so results can be recreated regardless of which team member or machine triggers the pipeline.

Why this answer

Option A is correct because locking dependency versions inside a container image used for training guarantees that every team member executes the same libraries and framework versions, eliminating environment drift as a source of non-reproducible results on Vertex AI. Option C is correct because versioning datasets with DVC or tracking them through Vertex AI ML Metadata pins the exact data lineage and artifact versions, so any experiment can be re-run against the identical input data. Options B, D, and E do not ensure reproducibility: real-time notebook sharing (B) is a collaboration feature, not a versioning mechanism; letting each team use its own environment (D) introduces dependency divergence; and always using random seeds (E) is misstated, since reproducibility requires fixing a consistent seed value, not merely using random seeds.

Exam trap

The trap here is that candidates often think 'always use random seeds' is a safe blanket rule, but in practice, seeds must be explicitly set and logged per run, and some operations (e.g., certain GPU kernels) are inherently non-deterministic, making this option an oversimplification that is not a guaranteed action for reproducibility.

241
MCQmedium

You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?

A.Use the component's temporary directory to share data between containers.
B.Pass the data as a serialized Python object in memory.
C.Write the data to Cloud Storage and pass the GCS URI as an artifact.
D.Store the data in a BigQuery table and pass the table reference.
AnswerC

Writing data to Cloud Storage and passing the GCS URI as an artifact avoids embedding large payloads in pipeline metadata, satisfying the large-dataset constraint. KFP passes lightweight artifact references between components rather than the data itself.

Why this answer

In KFP SDK v2, passing large datasets between components is best done by writing the data to Cloud Storage and passing the GCS URI as an artifact. This approach leverages KFP's built-in artifact tracking, ensures data persistence across container restarts, and avoids memory or disk limitations of ephemeral containers. The artifact is automatically serialized and passed as an input/output parameter, enabling efficient, scalable data exchange.

Exam trap

Google often tests the misconception that temporary directories are shared between containers in a pod, but in KFP each component runs in its own container with isolated storage, making Cloud Storage the correct choice for durable, cross-component data sharing.

How to eliminate wrong answers

Option A is wrong because a component's temporary directory is ephemeral and not shared between containers; each container runs in its own isolated filesystem, so data written there is lost after the component finishes. Option B is wrong because passing a serialized Python object in memory is limited by the container's memory capacity and cannot handle large datasets; KFP does not support in-memory object passing between components. Option D is wrong because storing data in a BigQuery table and passing the table reference is overkill for intermediate pipeline data; it introduces unnecessary latency, cost, and complexity compared to using Cloud Storage artifacts, which are the standard for KFP artifact passing.

242
MCQeasy

A data scientist wants to evaluate the performance of a BigQuery ML classification model on a test dataset. Which function should they use?

A.ML.PREDICT
B.ML.FEATURE_IMPORTANCE
C.ML.EVALUATE
D.ML.TRAIN
AnswerC

ML.EVALUATE computes standard classification metrics such as precision, recall, accuracy, F1 score, log loss and ROC AUC against a labelled dataset. It is the BigQuery ML function designed for assessing an already-trained model's performance, unlike ML.PREDICT, which only returns predictions.

Why this answer

ML.EVALUATE is the correct function because it computes classification metrics (e.g., precision, recall, accuracy, F1 score, ROC AUC) directly on a trained BigQuery ML model using a provided test dataset or evaluation input. This is the dedicated function for assessing model performance after training, aligning with the task of evaluating a classification model on held-out test data.

Exam trap

Google often tests the distinction between prediction (ML.PREDICT) and evaluation (ML.EVALUATE), trapping candidates who confuse generating outputs with measuring performance, especially when the question mentions 'evaluate performance' but the candidate fixates on 'predict' as the primary ML function.

How to eliminate wrong answers

Option A is wrong because ML.PREDICT is used to generate predictions (class labels or probabilities) on new data, not to compute evaluation metrics like accuracy or precision. Option B is wrong because ML.FEATURE_IMPORTANCE is used to retrieve feature weights or importance scores from a trained model (e.g., for interpretability), not to evaluate overall model performance on a test set. Option D is wrong because ML.TRAIN is used to initiate the training process of a BigQuery ML model, not to evaluate an already trained model on test data.

243
MCQmedium

A company has a model serving predictions on Vertex AI Endpoints and wants to monitor for prediction drift. They enable Vertex AI Model Monitoring but also need to see a confusion matrix over time. How should they set up the confusion matrix monitoring?

A.Use Cloud Monitoring to create a custom dashboard with a confusion matrix chart
B.Export predictions to Cloud Storage and run a Dataflow job to compute confusion matrices
C.Upload ground truth data to BigQuery and use Vertex AI Model Monitoring's model quality monitoring
D.Enable Vertex AI Explainable AI and configure it to output confusion matrices
AnswerC

Model quality monitoring computes confusion matrices, but it requires labelled ground truth to compare against predictions. Uploading ground truth to BigQuery supplies that reference data, satisfying the stem's need for confusion matrix metrics over time on the Vertex AI Endpoint.

Why this answer

Vertex AI Model Monitoring's model quality monitoring requires ground truth labels to compute metrics like confusion matrices, precision, recall, and F1. By uploading ground truth to BigQuery and configuring the monitoring job to join predictions with labels, Vertex AI can calculate and display confusion matrices over time. Prediction drift monitoring alone cannot produce a confusion matrix because it lacks label data.

Exam trap

PMLE often tests whether candidates confuse prediction drift monitoring (no labels needed) with model quality monitoring (labels required) — the confusion matrix is only computable with ground truth.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring dashboards display metrics, not computed confusion matrices — Vertex AI does not emit a confusion matrix metric that Cloud Monitoring can chart directly. Option B is wrong because exporting to Cloud Storage and running Dataflow is a manual, custom pipeline that reinvents functionality Vertex AI Model Monitoring already provides natively. Option D is wrong because Explainable AI produces feature attributions (e.g., SHAP values), not confusion matrices — it explains individual predictions, not aggregate classification performance.

244
MCQmedium

You are designing a distributed training job on Vertex AI for a PyTorch model using DataDistributedParallel (DDP). You have 4 nodes, each with 4 GPUs. What is the total number of workers that should be configured in the TF_CONFIG equivalent for PyTorch?

A.4
B.8
C.1
D.16
AnswerD

DataDistributedParallel assigns one worker process per GPU, so 4 nodes multiplied by 4 GPUs each yields 16 workers. Configuring 16 matches the total GPU count and ensures every device participates in the distributed training job.

Why this answer

In PyTorch DDP on Vertex AI, the number of workers equals the total number of processes across all nodes, which is nodes × GPUs per node. With 4 nodes and 4 GPUs each, the total is 16 workers, so the TF_CONFIG-equivalent worker count should be 16. Each GPU runs one DDP process, and all 16 processes participate in the all-reduce synchronization.

Exam trap

PMLE often tests the worker-count calculation — candidates multiply nodes by GPUs incorrectly or forget that each GPU runs its own DDP process, leading them to pick the node count instead of the total process count.

How to eliminate wrong answers

Option A is wrong because 4 corresponds only to the number of nodes, ignoring the 4 GPUs per node that each run a separate DDP process. Option B is wrong because 8 would correspond to 2 GPUs per node or 2 nodes × 4 GPUs, which does not match the stated topology. Option C is wrong because 1 would mean a single-process job, which contradicts the distributed 4×4 GPU configuration described.

245
MCQhard

A data scientist uses Vertex AI Pipelines to orchestrate an ML workflow. They want to reuse a component from Google's curated repository. What is the recommended way to incorporate it?

A.Import the component from Google Cloud Build
B.Use the 'aiplatform' Python SDK to define the component
C.Use prebuilt components from the Google Cloud Pipeline Components repository
D.Copy the component code into the pipeline definition
AnswerC

Google Cloud Pipeline Components are published, versioned Vertex AI pipeline components that can be loaded and inserted directly into a pipeline definition, giving tested implementations without writing component code. Referencing them from the curated repository is the documented reuse path, satisfying the reuse requirement.

Why this answer

Google provides a curated set of prebuilt components in the Google Cloud Pipeline Components repository, which are designed to be directly imported and used within Vertex AI Pipelines. These components encapsulate common ML tasks (e.g., model training, deployment) and are maintained by Google, ensuring compatibility and reducing custom code. Using them is the recommended approach to avoid reinventing the wheel and to leverage Google's best practices.

Exam trap

The trap here is that candidates may confuse the 'aiplatform' SDK (used for direct API calls) with the pipeline components SDK, or assume that copying code is acceptable for reusability, when Google specifically recommends using the curated prebuilt components to ensure compatibility and reduce maintenance overhead.

How to eliminate wrong answers

Option A is wrong because Google Cloud Build is a CI/CD service for building and testing code, not a repository for reusable pipeline components; importing a component from Cloud Build would not provide the curated, prebuilt component logic needed for Vertex AI Pipelines. Option B is wrong because the 'aiplatform' Python SDK is used to interact with Vertex AI services (e.g., creating datasets, jobs) but does not define or import prebuilt pipeline components; defining a component from scratch would bypass the curated repository. Option D is wrong because copying component code into the pipeline definition defeats the purpose of reusability and maintainability, and it is not the recommended method; the curated repository provides versioned, tested components that should be referenced rather than duplicated.

246
Multi-Selecteasy

Which TWO actions are appropriate when you detect that a production model's prediction distribution has shifted significantly from the training distribution?

Select 2 answers
A.Immediately roll back to the previous model version
B.Increase logging for future predictions
C.Retrain the model using the most recent data
D.Investigate the cause of the shift before taking corrective action
E.Reduce the traffic to the model to minimize impact
AnswersC, D

Retraining on recent data lets the model learn the new input-output relationship underlying the shifted distribution, restoring calibration. This is appropriate once drift is confirmed, though pairing it with root-cause investigation avoids retraining on transient or corrupted data.

Why this answer

Option C is correct because when a production model's prediction distribution has drifted from the training distribution, the underlying data distribution has likely changed, so retraining the model on the most recent data is the standard remediation to restore alignment between the model and current conditions. Option D is correct because before applying any fix, you must investigate the root cause of the shift — whether it is genuine data drift, a data pipeline/schema change, or a feature-engineering bug — since the appropriate corrective action depends on the cause. Option A is not appropriate as an automatic first step because rolling back only helps if the previous version is still valid for the current data distribution, and it may not address the underlying drift.

Option B is not appropriate because simply increasing logging does not correct the shifted predictions, though it may support diagnosis. Option E is not appropriate because reducing traffic does not fix the drift and may harm service availability without addressing the model's degraded relevance.

Exam trap

Google Cloud often tests the misconception that immediate rollback or traffic reduction is the correct first action, when in fact the proper response is to investigate the cause before taking corrective action like retraining.

247
MCQeasy

Which Vertex AI service is used to track the lineage of ML pipeline components, artefacts, and executions?

A.Vertex AI Metadata
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Experiments
AnswerA

Vertex AI Metadata provides the lineage tracking required, storing artefacts, executions and contexts as nodes within a managed metadata graph. It records relationships between pipeline components and their outputs, satisfying the stem's demand for tracing component, artefact and execution provenance across ML workflows.

Why this answer

Vertex AI Metadata is the service that tracks the lineage of ML pipeline components, artefacts, and executions. It provides a centralized repository to store and manage metadata about ML workflows, enabling reproducibility, auditing, and collaboration. By recording relationships between components, artefacts, and executions, it allows you to trace the provenance of models and data.

Exam trap

PMLE often tests the distinction between Vertex AI services that sound similar, such as confusing Metadata with Experiments or Model Registry, because all deal with tracking aspects of ML workflows.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Registry is a central repository for managing the lifecycle of ML models, including versioning and deployment, but it does not track pipeline component lineage. Option C is wrong because Vertex AI Feature Store is a centralized repository for organizing, storing, and serving ML features, focusing on feature management rather than lineage tracking. Option D is wrong because Vertex AI Experiments helps track and compare experiment runs, including parameters and metrics, but it does not provide comprehensive lineage tracking for pipeline components and artefacts.

248
MCQhard

You need to preprocess a large dataset (terabytes) for training a TensorFlow model. The preprocessing includes scaling and bucketizing features, and the same transformations must be applied during serving. Which tool should you use?

A.Dataflow with Apache Beam and tf.Transform
B.Dataproc with Spark ML
C.Vertex AI Feature Store
D.BigQuery ML
AnswerA

Dataflow with Apache Beam and tf.Transform applies identical scaling and bucketising transformations across terabyte-scale preprocessing and serving. tf.Transform exports the transform graph so training and prediction use consistent feature engineering, avoiding training-serving skew that separate pipelines would introduce.

Why this answer

tf.Transform is a library built on Apache Beam that allows you to define preprocessing functions (like scaling and bucketizing) once and then apply them consistently in both training and serving. Dataflow provides a scalable, fully managed runner for Apache Beam pipelines, making it ideal for processing terabyte-scale datasets. This combination ensures that the exact same transformations are used during training and serving, avoiding training-serving skew.

Exam trap

PMLE often tests the distinction between tools that can handle large-scale preprocessing and those that are meant for feature storage or model training, so candidates might incorrectly choose Vertex AI Feature Store or BigQuery ML thinking they handle preprocessing, but the key is the need for consistent transformations during serving with TensorFlow.

How to eliminate wrong answers

Option B is wrong because Spark ML on Dataproc does not provide a built-in mechanism to export and apply the same transformations during serving; it would require custom code to replicate the preprocessing logic, increasing the risk of skew. Option C is wrong because Vertex AI Feature Store is designed for online serving of features and does not handle large-scale preprocessing or transformation of raw data; it stores precomputed features. Option D is wrong because BigQuery ML is for building and training models directly in BigQuery using SQL, and while it can do some preprocessing, it does not provide a way to apply the same transformations during serving outside of BigQuery, and it is not designed for TensorFlow model serving.

249
MCQhard

A team is using Vertex AI AutoML to train a forecasting model. They need to retrain the model weekly and only if the new week's data significantly changes the data distribution. What is the most efficient way to achieve this?

A.Use a scheduled pipeline that always retrains
B.Use Cloud Monitoring alerts on data drift to trigger retraining
C.Use Vertex AI Model Monitoring to detect drift and trigger a pipeline
D.Use Cloud Functions on schedule to compare distributions
AnswerC

Model Monitoring continuously compares incoming prediction data against the training baseline and fires alerts when distribution shifts, so drift detection can automatically trigger the retraining pipeline. This satisfies the weekly, change-conditional requirement without wasteful scheduled retraining.

Why this answer

Vertex AI Model Monitoring can be configured to detect data drift on the model's input features, and when drift exceeds a threshold, it can trigger a Cloud Function or a Vertex AI pipeline to retrain the model. This approach avoids unnecessary retraining when the data distribution has not changed significantly, which is more efficient than always retraining. The integration with Cloud Functions or Pub/Sub allows for a serverless, event-driven retraining pipeline that only runs when needed.

Exam trap

Google Cloud often tests the distinction between infrastructure monitoring (Cloud Monitoring) and model-specific monitoring (Vertex AI Model Monitoring), and candidates mistakenly choose Cloud Monitoring because they think it can detect data drift, but it lacks the statistical algorithms needed for feature-level distribution comparison.

How to eliminate wrong answers

Option A is wrong because a scheduled pipeline that always retrains ignores the requirement to retrain only when the new week's data significantly changes the data distribution, leading to wasted compute resources and potential model instability from unnecessary retraining. Option B is wrong because Cloud Monitoring alerts are designed for infrastructure and application metrics (e.g., CPU, latency), not for detecting data drift in model features; data drift detection requires feature-level statistical analysis, which is not a built-in capability of Cloud Monitoring. Option D is wrong because using Cloud Functions on a schedule to compare distributions would require custom code to compute statistical tests (e.g., KS test) and manage state, which is less efficient and more error-prone than using Vertex AI Model Monitoring's managed drift detection service.

250
Multi-Selecthard

A company runs a Vertex AI pipeline that uses a container component to preprocess data. The component downloads a large file from a public URL and saves the output to Cloud Storage. The pipeline fails intermittently with a 'timeout' error. Which THREE steps should the team take to improve reliability? (Choose three.)

Select 3 answers
A.Make the component idempotent by checking for existing output before processing.
B.Increase the component's timeout setting.
C.Reduce the size of the file being downloaded.
D.Implement retries with exponential backoff in the component.
E.Increase the machine type for the component.
AnswersA, B, D

Idempotency lets a retried component detect that output already exists in Cloud Storage and skip re-downloading, so repeated attempts after a timeout do not duplicate work or waste the timeout window. This directly addresses the intermittent timeout failures.

Why this answer

Option A is correct because making the component idempotent (e.g., checking whether the output already exists in Cloud Storage before re-downloading and reprocessing) prevents redundant work on retries and ensures a re-run after a transient timeout does not corrupt or duplicate results. Option B is correct because the intermittent 'timeout' error indicates the component's configured timeout is too short for the large download plus preprocessing; raising the timeout setting gives the component enough time to finish on slow runs. Option D is correct because implementing retries with exponential backoff inside the component lets it recover from transient network or service failures when downloading from the public URL, which is the typical cause of intermittent timeouts.

Option C is not appropriate because reducing the file size changes the data being processed and is not a reliability fix the team can simply apply to the existing pipeline. Option E is not appropriate because increasing the machine type addresses CPU or memory constraints, not a timeout caused by download duration or transient network failures.

Exam trap

Google exams often test the distinction between fixing the symptom (increasing timeout) and addressing the root cause (idempotency and retries), leading candidates to overlook that idempotency and retries together with a reasonable timeout form the most robust solution.

251
Multi-Selecthard

You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?

Select 3 answers
A.Store batch prediction results in BigQuery.
B.Create a Vertex AI batch prediction job with input from GCS (TFRecord files).
C.Use Dataflow to read images and write TFRecord files to GCS.
D.Store batch prediction results in GCS.
E.Use Cloud Functions to convert images to TFRecord.
AnswersB, C, D

A batch prediction job reading TFRecord files directly from Cloud Storage matches the model's expected input format, avoiding costly conversion of 100 TB. Vertex AI distributes the job across managed workers, meeting the two-hour window cost-effectively.

Why this answer

Option B is correct because Vertex AI batch prediction natively accepts TFRecord input files from Cloud Storage, which matches the custom TensorFlow model's expected format and lets the managed service handle the large-scale inference job within the 2-hour window. Option C is correct because Dataflow is a scalable, cost-effective managed service for reading 100 TB of images from GCS and transforming them into TFRecord files, a task Cloud Functions cannot handle at this volume. Option D is correct because Vertex AI batch prediction writes its output to Cloud Storage, which is the appropriate and cost-effective destination for large result sets from 100 TB of input.

Option A is not appropriate because BigQuery is not the default or cost-effective sink for massive batch prediction outputs and would add unnecessary loading overhead. Option E is not appropriate because Cloud Functions has execution time, memory, and concurrency limits that make it unsuitable for converting 100 TB of images to TFRecord.

Exam trap

Google often tests the misconception that Cloud Functions can handle large-scale data processing tasks, but the trap here is that Cloud Functions have strict timeout and memory limits, making them unsuitable for converting 100 TB of images to TFRecord format.

252
MCQeasy

You have an online prediction model that is showing increasing prediction latency. You have already verified that the request rate and input data size are unchanged. Which of the following should you investigate next?

A.Check if the model was recently updated to a larger version
B.Check the monitoring dashboard configuration
C.Check if the feature engineering logic was changed
D.Check the geographic location of the endpoint
AnswerA

Latency rising while request rate and payload size stay constant points to increased per-inference compute. A larger model version, with more parameters or layers, raises inference time. Checking recent deployment history isolates this cause before investigating infrastructure or network factors.

Why this answer

If request rate and input data size are unchanged, increased prediction latency often points to a change in the model itself. A larger model (e.g., deeper neural network, more parameters) requires more computation per inference, directly increasing latency. This is a common root cause when monitoring ML pipelines, as model version updates can silently alter performance characteristics.

Exam trap

Google Cloud often tests the distinction between network-level latency (e.g., geographic location) and compute-level latency (e.g., model size), tempting candidates to pick the geographic option when the root cause is model-related.

How to eliminate wrong answers

Option B is wrong because the monitoring dashboard configuration only affects how metrics are displayed or alerted, not the underlying latency of predictions. Option C is wrong because feature engineering logic changes would alter input data size or structure, but the question states input data size is unchanged. Option D is wrong because the geographic location of the endpoint affects network latency, not the model's prediction latency (which is server-side compute time).

253
MCQeasy

A marketing team wants to automatically categorize customer feedback emails into topics such as 'billing', 'technical support', or 'general inquiry'. They have a dataset of 5,000 labeled emails and want to build a custom model with minimal coding effort. Which Google Cloud service should they use?

A.Dialogflow CX
B.Vision API
C.Natural Language API
D.AutoML Natural Language
AnswerD

AutoML Natural Language allows training custom text classification models with minimal coding, using labeled data. It handles preprocessing, training, and deployment automatically. With 5,000 labeled emails, it can achieve high accuracy and is ideal for teams with limited ML expertise. It integrates with other Google Cloud services and provides a user-friendly interface, making it the best fit for this low-code scenario.

Why this answer

AutoML Natural Language is purpose-built for custom text classification with minimal coding. It allows training on labeled data to recognize specific categories like billing or technical support. The other services either provide only pre-trained general models or are designed for different modalities, making AutoML Natural Language the correct choice for this low-code custom classification task.

Exam trap

The trap here is confusing the pre-trained Natural Language API with AutoML Natural Language; the former cannot be customized for specific topics without training, while the latter is designed for custom classification.

254
MCQmedium

A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?

A.Deploy the challenger to a separate endpoint and use a proxy to split traffic.
B.Use Cloud Load Balancing with weighted backend services.
C.Deploy both models to the same endpoint and use traffic splitting.
D.Use Vertex AI Experiments to manage model versions.
AnswerC

Vertex AI Endpoints support deploying multiple model versions to one endpoint with traffic splitting, letting you route a percentage to the challenger and gradually shift it. This satisfies the A/B testing and champion-to-challenger migration requirement.

Why this answer

Vertex AI Endpoints natively support traffic splitting, allowing you to deploy multiple model versions (e.g., champion and challenger) to the same endpoint and assign a percentage of traffic to each. This enables gradual A/B testing without additional infrastructure, as the endpoint automatically routes requests based on the configured split. Option C is correct because it leverages this built-in feature, which is designed specifically for this use case.

Exam trap

The PMLE exam often tests the misconception that traffic splitting requires external load balancers or proxies, when in fact Vertex AI Endpoints provide this capability natively, and candidates may overlook the built-in feature in favor of more complex architectures.

How to eliminate wrong answers

Option A is wrong because deploying the challenger to a separate endpoint and using a proxy adds unnecessary complexity, latency, and management overhead; Vertex AI Endpoints already provide traffic splitting without external proxies. Option B is wrong because Cloud Load Balancing operates at the network layer (HTTP(S) or TCP/UDP) and is designed for distributing traffic across regional backends, not for splitting traffic between model versions on the same Vertex AI Endpoint; it would require separate endpoints and does not integrate with Vertex AI's model versioning. Option D is wrong because Vertex AI Experiments is a tool for tracking and comparing model training runs and hyperparameters, not for serving or routing live traffic; it has no mechanism to split traffic between deployed models.

255
MCQhard

A company uses Vertex AI Prediction with a custom container for a TensorFlow model. They notice that after deploying a new model version, requests still go to the old version. What is the most likely cause?

A.The custom container is not compatible with Vertex AI
B.The model is cached and needs cache invalidation
C.Traffic is not split to the new model version
D.The new model version was not deployed to the same endpoint
AnswerC

Vertex AI Prediction routes requests according to each version's traffic split percentage. Deploying a new version does not automatically shift traffic; until the split is adjusted, the old version continues serving all requests, so requests never reach the new model.

Why this answer

In Vertex AI Prediction, when you deploy a new model version to an existing endpoint, you must explicitly allocate traffic to it. By default, the new version receives 0% traffic, so all requests continue to be served by the old version. The correct fix is to update the endpoint's traffic split, for example via the console or the `gcloud ai endpoints update` command with the `--traffic-split` flag.

Exam trap

Google Cloud often tests the misconception that deploying a new model version automatically replaces the old one, when in fact Vertex AI requires an explicit traffic split update to shift requests to the new version.

How to eliminate wrong answers

Option A is wrong because Vertex AI supports custom containers for TensorFlow models as long as they implement the required HTTP health check and prediction endpoints; incompatibility would cause deployment failure, not silent routing to an old version. Option B is wrong because Vertex AI does not cache model predictions at the endpoint level; caching is not a factor in traffic routing between model versions. Option D is wrong because deploying to the same endpoint is exactly what the user did; the issue is that the new version was deployed but not given any traffic share, not that it was deployed to a different endpoint.

256
Multi-Selecthard

A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)

Select 2 answers
A.Use Vertex AI Experiments to track the traffic split percentages.
B.Enable Cloud Memorystore to cache identical requests for both models.
C.Deploy the challenger model to the same endpoint as the champion with a separate deployed model.
D.Configure traffic split in the endpoint's traffic_split field (e.g., champion:90, challenger:10).
E.Use Cloud Monitoring to track custom metrics like conversion rate per model version.
AnswersC, D

Vertex AI Endpoints support multiple deployed models behind one endpoint, each with its own deployed model ID. Deploying the challenger alongside the champion on the same endpoint is the prerequisite that lets traffic be split between them.

Why this answer

Option C is correct because Vertex AI Endpoints support multiple deployed models behind a single endpoint, so the challenger must be deployed as a separate DeployedModel (with its own model ID and deployed_model_id) on the same endpoint as the champion to enable side-by-side serving. Option D is correct because gradual traffic shifting is implemented by setting the endpoint's traffic_split map, which assigns percentage weights keyed by deployed_model_id (e.g., champion deployed model: 90, challenger deployed model: 10), allowing controlled canary rollout. Option A is not required because Vertex AI Experiments tracks training/experiment runs and metrics, not live endpoint traffic splitting.

Option B is not required because Memorystore is a caching layer and does not perform traffic splitting or model comparison. Option E, while useful for observing conversion rate, is not a required step to shift traffic and is not marked correct; the question asks for the steps needed to perform the gradual shift itself.

Exam trap

This question tests the distinction between monitoring (which is optional after deployment) and the actual configuration steps required to shift traffic; candidates mistakenly select Cloud Monitoring (Option E) as a required step, but the question specifically asks for steps to 'gradually shift traffic,' which is accomplished by deploying to the same endpoint and setting the traffic split, not by monitoring after the fact.

257
MCQmedium

A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?

A.Delta Lake on Dataproc
B.BigQuery table snapshots
C.DVC (Data Version Control)
D.Vertex AI Feature Store
AnswerA

Delta Lake provides ACID transactions and scalable metadata handling over Parquet files in Cloud Storage, letting concurrent Dataproc jobs read and write multi-terabyte datasets consistently. Its transaction log delivers the snapshot isolation and versioning the stem demands.

Why this answer

Delta Lake on Dataproc provides ACID transactions on cloud storage data lakes, enabling concurrent reads/writes with consistency.

258
Multi-Selectmedium

A retail company wants to build a low-code ML solution to predict customer lifetime value (CLV) using historical transaction data stored in BigQuery. They have limited ML expertise and want to use BigQuery ML. Which two steps are necessary to train and evaluate a model using BigQuery ML? (Choose two.)

Select 2 answers
A.Use ML.PREDICT to generate predictions on new data.
B.Manually split the data into training and test sets using a CREATE TABLE statement.
C.Use ML.EVALUATE to assess the model's performance on a test dataset.
D.Export the model to a TensorFlow SavedModel for deployment.
E.Create a model using the CREATE MODEL statement with the appropriate model type.
AnswersC, E

ML.EVALUATE is a function that computes evaluation metrics for a trained model, such as RMSE for regression. It is essential to assess model performance and ensure it meets business needs. You run it against a dataset not used in training. This step is necessary for model validation and is performed via SQL, maintaining the low-code approach.

Why this answer

Training a model in BigQuery ML requires the CREATE MODEL statement, which defines and trains the model using SQL. Evaluation is performed with ML.EVALUATE to compute metrics like RMSE or accuracy. These two steps are essential for building and assessing a model.

Other steps like exporting or predicting are optional or occur after evaluation, and manual data splitting is unnecessary due to automatic splitting.

Exam trap

The trap here is thinking that data must be manually split or that model export is required for training; BigQuery ML handles splitting automatically and export is only for external deployment.

259
MCQmedium

You are monitoring a machine learning pipeline that runs on Vertex AI Pipelines. The pipeline occasionally fails with a 'ResourceExhausted' error when attempting to read data from BigQuery. Which action should you take to resolve this issue?

A.Switch from BigQuery to Cloud Storage for data source
B.Increase the memory allocated to the pipeline step
C.Reduce the complexity of the BigQuery query or increase the reservation size
D.Reduce the batch size of the data being read
AnswerC

ResourceExhausted signals BigQuery quota or slot limits being hit, not a pipeline bug. Simplifying the query reduces scanned bytes and slot demand, while enlarging the reservation raises available capacity, both directly relieving the exhausted resource.

Why this answer

The 'ResourceExhausted' error when reading from BigQuery indicates that the query is consuming more resources than the BigQuery reservation allows. Option C is correct because reducing query complexity (e.g., using fewer JOINs, aggregations, or partitions) or increasing the reservation size directly addresses the root cause by either lowering resource demand or allocating more capacity. Other options like switching to Cloud Storage or adjusting pipeline memory do not fix the BigQuery-specific quota or slot exhaustion.

Exam trap

Google Cloud often tests the misconception that memory or batch size adjustments in the pipeline environment can fix backend service quota errors, when in fact the error is specific to BigQuery's resource management (slots/queries) and requires query optimization or reservation changes.

How to eliminate wrong answers

Option A is wrong because switching to Cloud Storage does not resolve the BigQuery resource exhaustion; it changes the data source but introduces new latency and format compatibility issues without addressing the query's resource consumption. Option B is wrong because increasing memory allocated to the pipeline step only affects the compute environment (e.g., the container running the pipeline), not the BigQuery service's slot or query quota limits. Option D is wrong because reducing the batch size of data being read may reduce memory pressure on the pipeline but does not affect the BigQuery query's resource usage; the error originates from BigQuery's backend, not from the client-side read volume.

260
Multi-Selecthard

An ML engineer is building a monitoring dashboard for a Vertex AI pipeline that includes training, evaluation, and batch prediction. Which THREE components should be included to provide comprehensive observability? (Select THREE.)

Select 3 answers
A.Pipeline execution status, duration, and failure rates for each component.
B.Compute engine CPU and memory logs for each pipeline step.
C.Model evaluation metrics (e.g., accuracy, AUC) after training and validation.
D.Data validation reports showing anomaly counts and feature statistics.
E.Online prediction latency and request count from the deployed model endpoint.
AnswersA, C, D

Pipeline execution status, duration, and failure rates expose orchestration-level telemetry for each training, evaluation, and batch prediction step, satisfying the requirement for comprehensive observability across the whole Vertex AI pipeline. These metrics reveal where runs stall or fail, which resource-level metrics alone cannot attribute to a specific component.

Why this answer

Option A is correct because tracking pipeline execution status, duration, and failure rates per component gives the operational observability needed to detect where a training, evaluation, or batch prediction step fails or slows down. Option C is correct because model evaluation metrics such as accuracy and AUC after training and validation are essential to monitor model quality and detect degradation before deployment. Option D is correct because data validation reports with anomaly counts and feature statistics surface input data drift, skew, and schema issues that directly affect pipeline and model reliability.

Option B is not appropriate because raw Compute Engine CPU and memory logs for each step are low-level infrastructure telemetry, not the pipeline-level observability components the scenario requires. Option E is not appropriate because online prediction latency and request counts apply to a deployed endpoint, whereas this pipeline uses batch prediction and has no online serving endpoint.

Exam trap

The trap here is that candidates often confuse infrastructure monitoring (CPU/memory logs) or serving-layer metrics (online prediction latency) with pipeline-specific observability, leading them to select options that are relevant to different stages of the ML lifecycle rather than the pipeline itself.

261
Multi-Selectmedium

A data science team collaborates using Vertex AI Workbench user-managed notebooks. They want to version control their notebook code and share it with team members. Which TWO tools should they use? (Choose 2)

Select 2 answers
A.Git integration in Vertex AI Workbench
B.Cloud Functions
C.Vertex AI Model Registry
D.Vertex AI Experiments
E.Cloud Source Repositories
AnswersA, E

Git integration in Vertex AI Workbench lets users commit notebook code to a remote repository, track revisions and share branches with teammates. This satisfies the version-control and collaboration requirement without leaving the managed notebook environment.

Why this answer

Option A, Git integration in Vertex AI Workbench, is correct because user-managed notebooks include built-in Git support, allowing data scientists to clone repositories, commit, push, and pull notebook code directly from the JupyterLab interface for version control and collaboration. Option E, Cloud Source Repositories, is correct because it is a fully managed private Git repository service on Google Cloud that can host the team's notebook code and integrate with Workbench's Git tooling for sharing and versioning. Option B, Cloud Functions, is incorrect because it is a serverless compute service for event-driven functions, not a version control or code-sharing tool.

Option C, Vertex AI Model Registry, is incorrect because it manages trained model versions and their metadata, not notebook source code. Option D, Vertex AI Experiments, is incorrect because it tracks experiment runs, parameters, and metrics, not Git-based notebook version control.

Exam trap

The trap is selecting other Vertex AI components like Model Registry or Experiments, which sound related to ML workflows but are not for source code version control — candidates must distinguish between code versioning and model/experiment tracking.

262
MCQhard

You have a very large language model that does not fit on a single GPU. You need to train it efficiently across multiple GPUs on a single machine. Which approach should you use?

A.Data parallelism with MirroredStrategy
B.Data parallelism with MultiWorkerMirroredStrategy
C.Use TPU training as TPUs have more memory
D.Model parallelism using pipeline parallelism
AnswerD

Pipeline parallelism splits the model's layers across GPUs so each device holds only a subset of parameters, with activations passed between stages. This fits a model too large for one GPU while training efficiently within a single machine.

Why this answer

When a model does not fit on a single GPU, model parallelism is required because it partitions the model itself across devices. Pipeline parallelism is a specific model parallelism technique that splits the model into stages across GPUs and pipelines micro-batches to maintain utilization, making it the appropriate approach for training a very large model on multiple GPUs in one machine.

Exam trap

PMLE often tests the misconception that data parallelism can handle models too large for one GPU, when in fact data parallelism replicates the model and only model parallelism (including pipeline parallelism) partitions it across devices.

How to eliminate wrong answers

Option A is wrong because data parallelism with MirroredStrategy replicates the entire model on each GPU, which is impossible if the model does not fit on one GPU. Option B is wrong because MultiWorkerMirroredStrategy is also data parallelism, just across multiple machines, and still requires the full model to fit on each worker. Option C is wrong because TPUs are not a given in a single-machine multi-GPU scenario, and switching hardware does not address the architectural need for model partitioning; also, TPUs have their own memory limits.

263
MCQeasy

What is the primary purpose of Vertex AI Edge Manager?

A.To run batch predictions on edge devices
B.To deploy and manage ML models on edge devices at scale
C.To convert models to TensorFlow Lite automatically
D.To train models on edge devices using federated learning
AnswerB

Vertex AI Edge Manager deploys and manages ML models across fleets of edge devices at scale, handling distribution, versioning and monitoring. It targets edge hardware rather than cloud endpoints, which is precisely the capability the question asks about.

Why this answer

Vertex AI Edge Manager is specifically designed to deploy, monitor, and manage ML models on edge devices at scale. It handles model packaging, over-the-air updates, and health monitoring across fleets of edge devices, which is distinct from simply running batch predictions or converting model formats.

Exam trap

Google Cloud often tests the distinction between 'managing models at scale' (deployment, updates, monitoring) and 'running inference' or 'converting formats' — candidates confuse the operational management role with the execution or preprocessing steps.

How to eliminate wrong answers

Option A is wrong because batch predictions on edge devices are a use case, not the primary purpose; Vertex AI Edge Manager focuses on lifecycle management (deployment, updates, monitoring) rather than just executing predictions. Option C is wrong because model conversion to TensorFlow Lite is handled by tools like the TensorFlow Lite Converter or Vertex AI's model optimization services, not by Edge Manager itself. Option D is wrong because training on edge devices using federated learning is a separate paradigm (e.g., TensorFlow Federated) and is not a core function of Vertex AI Edge Manager, which manages already-trained models.

264
MCQeasy

To enable collaboration on notebook-based experiments across teams, what is the recommended approach in Google Cloud?

A.Use Colab Enterprise notebooks with shared runtimes and IAM permissions
B.Share Docker images containing the notebook environment
C.Each team member works on their own local Jupyter notebook and shares screenshots
D.Store notebooks in a Cloud Storage bucket and open them with Vertex AI Workbench
AnswerA

Colab Enterprise enables collaborative editing and shared compute resources.

Why this answer

Colab Enterprise notebooks with shared runtimes and IAM permissions is the recommended approach because it provides a fully managed, collaborative environment where multiple users can work on the same notebook simultaneously, with fine-grained access control via IAM and consistent runtime configurations. This eliminates version conflicts and environment drift, which are common in distributed notebook workflows.

Exam trap

Google Cloud often tests the misconception that shared storage (like Cloud Storage) alone is sufficient for collaboration, but the key requirement is shared runtimes and concurrent editing, which only Colab Enterprise provides among the options.

How to eliminate wrong answers

Option B is wrong because sharing Docker images containing the notebook environment addresses environment reproducibility but does not enable real-time collaboration or shared runtime execution; each user would still need to launch their own instance and manually sync changes. Option C is wrong because each team member working on their own local Jupyter notebook and sharing screenshots is a manual, non-scalable approach that lacks version control, concurrent editing, and centralized data access, making it unsuitable for team collaboration. Option D is wrong because storing notebooks in a Cloud Storage bucket and opening them with Vertex AI Workbench provides shared storage but does not inherently support shared runtimes or concurrent editing; Vertex AI Workbench instances are typically single-user, and multiple users would need to coordinate access to avoid conflicts.

265
MCQhard

You are deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that involves scaling numerical features and one-hot encoding categorical features. You have packaged the preprocessing and model into a single Python script that uses a custom prediction routine. You need to ensure that the online prediction service can handle varying input formats and provide low-latency responses. What should you do?

A.Use Vertex AI's custom prediction routine with a preprocessor that accepts raw input and transforms it, and deploy the model as a model artifact with the custom routine.
B.Deploy the model as a custom container on Vertex AI, and implement the preprocessing logic inside the container's HTTP server.
C.Use Vertex AI's feature store to perform preprocessing, and then feed the features directly to the model.
D.Deploy the model using Vertex AI's built-in scikit-learn container, and perform preprocessing on the client side before sending requests.
AnswerA

Vertex AI supports custom prediction routines where you can define a preprocessor that handles raw input and transforms it before passing to the model. This allows you to encapsulate preprocessing and model logic in a single Python package. It leverages Vertex AI's managed infrastructure for scaling and low-latency serving, and is the recommended approach for scikit-learn models with custom preprocessing.

Why this answer

Using a custom prediction routine with a preprocessor allows you to encapsulate preprocessing and model logic in a single package, ensuring consistency between training and serving. Vertex AI manages the infrastructure, providing low-latency and scalability. This is the most efficient and maintainable approach for scikit-learn models with custom preprocessing.

Exam trap

The trap here is assuming that a custom container is necessary for custom preprocessing, when Vertex AI's custom prediction routines already provide this capability without the overhead of managing a container.

266
MCQhard

A team is deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that scales numerical features using statistics computed from the training set. The preprocessing must be identical between training and serving. The team wants to minimize latency and ensure consistency. What should they do?

A.Wrap the preprocessing and model into a single scikit-learn Pipeline object, and deploy that pipeline as a custom model on Vertex AI.
B.Use a Vertex AI pipeline to preprocess the data before training, and then deploy the model without preprocessing, assuming the input data at serving time will already be scaled.
C.Save the scikit-learn model and the scaler as separate artifacts, and in the prediction script, load both and apply the scaler before calling model.predict.
D.Use Vertex AI Feature Store to store the scaled features and serve them online, bypassing the need for preprocessing in the model.
AnswerA

A scikit-learn Pipeline encapsulates all preprocessing and the final estimator, ensuring that the exact same transformations are applied during training and serving. Deploying the pipeline as a custom model on Vertex AI guarantees consistency and reduces the risk of training-serving skew, while keeping the serving logic simple.

Why this answer

Wrapping preprocessing and the model into a single scikit-learn Pipeline ensures that the exact same transformations are applied during both training and serving. This eliminates training-serving skew and simplifies deployment. Vertex AI supports deploying custom models with custom prediction routines, but using a native scikit-learn Pipeline is more straightforward and less error-prone, as it encapsulates all steps and can be serialized as one artifact.

Exam trap

The trap here is assuming that preprocessing can be handled separately or that the client will send preprocessed data, overlooking the need for a single artifact that guarantees consistency.

267
MCQmedium

Your team trains a scikit-learn model locally and uploads it to Vertex AI Model Registry. A colleague needs to deploy it to a Vertex AI Endpoint for online prediction with a prebuilt container. The model artifacts are stored in a Cloud Storage bucket. Which deployment approach should they use?

A.Create a custom prediction routine in a Python source distribution, upload it as a model artifact, and deploy it without specifying a serving container.
B.Export the model to a SavedModel format, import it with a TensorFlow prebuilt container, and deploy it to an Endpoint.
C.Import the model with the prebuilt scikit-learn container image and deploy it to an Endpoint using the Model Registry UI or the gcloud ai endpoints deploy-model command.
D.Upload the model artifact to Vertex AI Model Registry as a BigQuery ML model and deploy it using the BigQuery ML serving container.
AnswerC

This is correct because Vertex AI Model Registry supports importing scikit-learn models with a prebuilt serving container, and the imported model can be deployed directly to an Endpoint. The prebuilt scikit-learn container handles prediction requests without requiring custom inference code, and deployment can be done through the console or gcloud CLI. This matches the standard workflow for deploying a locally trained scikit-learn model.

Why this answer

The scikit-learn model was trained locally and stored in Cloud Storage, so it should be imported into Vertex AI Model Registry using the prebuilt scikit-learn container. This allows direct deployment to a Vertex AI Endpoint for online prediction. The prebuilt container removes the need for custom inference code and supports the standard deployment workflow via console or gcloud CLI.

Exam trap

The trap here is assuming that a locally trained scikit-learn model must be converted to TensorFlow SavedModel or packaged with custom code before it can be deployed to a Vertex AI Endpoint.

268
MCQhard

A large organization uses a multi-project setup with a central data lake. Different teams manage their own models. To enable cross-team sharing of features, they want to use Vertex AI Feature Store. What is the best practice to manage access?

A.Create a single Feature Store in a central project and grant fine-grained IAM roles
B.Export features to Cloud Storage
C.Create separate Feature Stores per team project
D.Use BigQuery authorized views
AnswerA

A single central Feature Store lets teams share features across projects, while fine-grained IAM roles restrict each team to only the features and operations it needs. This satisfies the cross-team sharing requirement without granting broad project-level access, which would violate least privilege.

Why this answer

Creating a single Feature Store in a central project with fine-grained IAM roles is the best practice because it centralizes feature management while allowing cross-team access control at the feature group or feature level. Vertex AI Feature Store supports IAM roles like `aiplatform.featureStoreAdmin` and `aiplatform.featureStoreDataViewer` to grant granular permissions, enabling teams to share features without duplicating data or exposing sensitive information. This approach avoids data silos and ensures consistent governance across the organization.

Exam trap

Google Cloud often tests the misconception that separate Feature Stores per team are needed for isolation, but the correct approach is to use a single Feature Store with fine-grained IAM to enable sharing while maintaining security.

How to eliminate wrong answers

Option B is wrong because exporting features to Cloud Storage introduces data duplication, latency, and manual synchronization overhead, defeating the purpose of a centralized feature store for real-time serving. Option C is wrong because creating separate Feature Stores per team project creates data silos, preventing cross-team sharing and requiring complex cross-project networking or data replication. Option D is wrong because BigQuery authorized views are designed for table-level access control in BigQuery, not for managing access to Vertex AI Feature Store entities like feature groups or online/offline stores, and they lack the low-latency serving capabilities of Feature Store.

269
MCQeasy

Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?

A.Vertex AI Model Optimization
B.Vertex AI Model Monitoring
C.Vertex AI Matching Engine
D.Vertex AI Continuous Training
AnswerA

Vertex AI Model Optimization applies quantisation and pruning to shrink trained models, reducing size and inference latency on constrained edge hardware. It meets the stem's requirement of faster edge inference without significant accuracy loss, unlike retraining or distillation approaches.

Why this answer

Vertex AI Model Optimization is the correct feature because it provides model quantization, pruning, and distillation techniques specifically designed to reduce model size and improve inference latency on edge devices. This service applies post-training quantization (e.g., FP32 to INT8) and structured weight pruning to shrink the model footprint while maintaining accuracy within acceptable thresholds, directly addressing the need for efficient deployment on resource-constrained hardware.

Exam trap

Google PMLE exams often test the distinction between 'optimization' (size/speed improvements) and 'monitoring' (observability), leading candidates to confuse Model Monitoring with performance tuning because both involve 'model performance' terminology.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Monitoring is used for detecting prediction drift, data skew, and feature attribution changes in deployed models, not for reducing model size or optimizing inference speed. Option C is wrong because Vertex AI Matching Engine is a vector similarity search service for large-scale embedding-based retrieval (e.g., recommendation systems), not a model compression or optimization tool. Option D is wrong because Vertex AI Continuous Training automates retraining pipelines based on new data or schedules, but it does not perform model size reduction or inference optimization.

270
MCQmedium

A retail company wants to build a product recommendation system using BigQuery ML for their e-commerce platform. The data includes customer purchase history, product metadata, and clickstream logs. The ML engineer needs to minimize manual feature engineering and leverage pre-built solutions. Which approach should the engineer take?

A.Use a pre-built recommendation model from Vertex AI Model Garden and deploy it to an endpoint.
B.Write a custom TensorFlow model using the Vertex AI Training service and deploy it via Vertex AI Prediction.
C.Export the data to CSV and use AutoML Tables to train a recommendation model.
D.Use BigQuery ML's matrix factorization model (CREATE MODEL with model_type='matrix_factorization') to train directly on historical interaction data.
AnswerD

BigQuery ML's matrix_factorization model learns latent user and item embeddings directly from historical interaction data, so purchase history and clickstream logs feed training without hand-crafted features. This satisfies the stem's constraint of minimising manual feature engineering while using a pre-built solution, unlike approaches requiring custom pipelines or external frameworks.

Why this answer

BigQuery ML's matrix factorization model (model_type='matrix_factorization') is purpose-built for recommendation systems using implicit or explicit feedback data. It trains directly on historical interaction data (e.g., user-item purchases) without requiring manual feature engineering, aligning with the goal of minimizing low-code ML effort. This approach leverages BigQuery's native SQL interface and scales automatically, making it ideal for the described e-commerce scenario.

Exam trap

The trap here is that candidates may assume Vertex AI Model Garden (Option A) is the go-to for pre-built ML, but it does not offer a pre-trained recommendation model that can be directly deployed without custom training on the company's data.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden provides pre-built models for tasks like vision or NLP, not a ready-to-use recommendation model that can be directly deployed without training on the company's specific interaction data. Option B is wrong because writing a custom TensorFlow model and training it via Vertex AI Training contradicts the requirement to minimize manual feature engineering and leverage pre-built solutions. Option C is wrong because exporting data to CSV and using AutoML Tables would require additional data preparation and does not natively handle the user-item interaction structure as efficiently as BigQuery ML's matrix factorization, which operates directly on the data in place.

271
MCQmedium

A team uses Vertex AI Pipelines with a custom training component that reads data from a BigQuery table. They need to ensure that a new pipeline run uses a specific snapshot of the training data for reproducibility. Which approach should they take?

A.Create a BigQuery table snapshot of the training data and pass the snapshot name to the training component as an input artifact.
B.Pass the BigQuery table name and a WHERE clause that filters on a timestamp column to the training component, and record the query in the pipeline parameters.
C.Export the BigQuery table to a Cloud Storage bucket in CSV format and pass the Cloud Storage URI to the training component.
D.Use BigQuery's time travel feature by adding FOR SYSTEM_TIME AS OF a fixed timestamp to the query in the training component.
AnswerA

A BigQuery table snapshot is an immutable, point-in-time copy of the table that persists even if the source table changes. By creating a snapshot and passing its name to the training component, you ensure the pipeline run always reads the same data. This provides reproducibility and aligns with Vertex AI Pipelines artifact tracking, as the snapshot can be recorded as an input artifact.

Why this answer

To ensure reproducibility, the training data must be immutable. A BigQuery table snapshot provides a point-in-time copy that does not change even if the source table is updated. Passing the snapshot name to the training component as an input artifact also integrates with Vertex AI Pipelines lineage tracking, making the data reference explicit and stable.

Exam trap

The trap here is relying on time travel or filtered queries for reproducibility, which are not durable or immutable beyond their retention limits.

272
MCQmedium

A machine learning engineer notices that the online prediction latency for a custom TensorFlow model deployed on Vertex AI has increased significantly over the past week. Cloud Monitoring shows that the CPU utilization of the endpoints remains below 40%, but the number of concurrent requests has doubled. What is the most likely cause of the latency increase?

A.Data skew causing longer inference time
B.Memory leak in the serving container
C.Insufficient number of replicas for autoscaling
D.Model overfitting
AnswerC

Doubling concurrent requests without CPU saturation points to queueing at the replica layer, not compute exhaustion. Vertex AI autoscaling adds replicas based on utilisation thresholds, so with CPU below 40% the autoscaler never scales out, leaving too few replicas to serve the higher concurrency, which inflates prediction latency.

Why this answer

The CPU utilization remains below 40% while concurrent requests have doubled, indicating that the existing replicas are not saturated on CPU but are bottlenecked by request queuing or thread contention. Vertex AI autoscaling scales based on CPU utilization by default; if the threshold is not crossed, new replicas are not provisioned, causing requests to queue and latency to spike. The engineer should verify the autoscaling configuration and consider scaling on request count or reducing the CPU utilization target.

Exam trap

Google Cloud often tests the misconception that low CPU utilization always means there is spare capacity, when in reality the bottleneck can be request queuing or thread pool exhaustion that does not raise CPU usage.

How to eliminate wrong answers

Option A is wrong because data skew would cause a persistent increase in per-request inference time, but the observation shows CPU utilization is low and latency increased only after request volume doubled, not due to a change in data distribution. Option B is wrong because a memory leak would manifest as increasing memory usage over time, potentially causing OOM kills or garbage collection pauses, but the described symptom is low CPU and doubled concurrency, not memory pressure. Option D is wrong because model overfitting affects prediction accuracy, not inference latency; overfitting does not change the computational cost of a forward pass.

273
MCQmedium

You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?

A.Configure the pipeline's training component to query the BigQuery table using a parameter that specifies the current date, and log that date in the pipeline run metadata.
B.Before each pipeline run, create a BigQuery table snapshot of the source table and configure the pipeline to read from that snapshot.
C.Set up a Cloud Scheduler job that copies the BigQuery table to a Cloud Storage bucket as CSV files, and have the pipeline read from that bucket.
D.Use BigQuery's streaming inserts to push data into a new table that the pipeline reads, ensuring that the data is frozen at the time of insertion.
AnswerB

BigQuery table snapshots are immutable and preserve the table's data at the time of creation. By reading from a snapshot, each pipeline run uses a consistent, reproducible dataset, and the snapshot can be referenced later for auditing. This directly addresses both consistency and reproducibility without altering the source table.

Why this answer

Using BigQuery table snapshots ensures that each pipeline run reads an immutable, point-in-time copy of the source data. This guarantees consistency across runs and enables exact reproducibility for audits, because the snapshot preserves the data as it existed when created. Other methods either do not freeze the data or introduce complexity without the same guarantees.

Exam trap

The trap here is assuming that querying a table with a date filter or exporting data provides a consistent snapshot, when only a native BigQuery snapshot guarantees immutability.

274
MCQeasy

You have a TensorFlow training script that runs on a single machine. To speed up training on Vertex AI with 8 GPUs on a single machine, which strategy should you use?

A.tf.distribute.ParameterServerStrategy
B.tf.distribute.MirroredStrategy
C.tf.distribute.TPUStrategy
D.tf.distribute.MultiWorkerMirroredStrategy
AnswerB

MirroredStrategy performs synchronous, all-reduce data-parallel training across multiple GPUs within one machine, replicating the model on each device and aggregating gradients. This directly satisfies the stem's constraint of 8 GPUs on a single machine, where MultiWorkerMirroredStrategy would add unnecessary cross-machine networking overhead.

Why this answer

tf.distribute.MirroredStrategy is designed for synchronous, data-parallel training across multiple GPUs on a single machine. It replicates the model on each GPU, splits each batch across replicas, and uses all-reduce (via NCCL) to aggregate gradients, which is exactly the scenario described: 8 GPUs on one machine. This is the canonical strategy for single-node multi-GPU TensorFlow training.

Exam trap

The trap is confusing single-machine multi-GPU (MirroredStrategy) with multi-machine multi-GPU (MultiWorkerMirroredStrategy) — candidates who skim the question miss the 'single machine' qualifier and pick the multi-worker variant.

How to eliminate wrong answers

Option A is wrong because ParameterServerStrategy is designed for asynchronous or synchronous training across many machines with dedicated parameter servers, which is overkill and inappropriate for a single machine with 8 GPUs. Option C is wrong because TPUStrategy targets Google Cloud TPUs, not GPUs — using it on a GPU machine will fail or fall back incorrectly. Option D is wrong because MultiWorkerMirroredStrategy is for multi-node, multi-GPU training across multiple machines; on a single machine it adds unnecessary coordination overhead and is not the intended strategy.

275
MCQhard

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a custom component for data validation. The component takes a dataset URI and outputs a validation report. The engineer wants to fail the pipeline immediately if the validation report indicates that the data is invalid, without running subsequent steps. How should the engineer implement this?

A.Use a conditional step after validation that checks the report and only proceeds if valid, otherwise skips subsequent steps.
B.Have the data validation component raise an exception if the data is invalid, causing the task to fail and the pipeline to stop.
C.Write the validation result to a Cloud Storage bucket and use a Cloud Function to monitor the bucket and cancel the pipeline if invalid.
D.Configure the pipeline to use a 'fail_fast' execution option that stops the pipeline if any component outputs a failure status.
AnswerB

If the data validation component raises an exception when data is invalid, the task fails, and Vertex AI Pipelines will not execute downstream steps that depend on its output. This is the simplest and most direct way to halt the pipeline upon validation failure. The pipeline status will reflect the failure, allowing for alerting and debugging.

Why this answer

The most effective way to fail the pipeline immediately upon invalid data is to have the data validation component raise an exception. This causes the task to fail, and Vertex AI Pipelines will not run any downstream tasks that depend on its output. The pipeline will be marked as failed, which can trigger alerts.

This approach is simple, reliable, and leverages the native failure handling of the pipeline orchestrator.

Exam trap

The trap here is thinking that a conditional step or an external monitor is needed to stop the pipeline, when simply failing the component achieves the goal.

276
MCQmedium

An organisation wants to monitor fairness of their loan approval model across demographic subgroups. They have predictions stored in BigQuery along with ground truth. Which GCP service can evaluate model performance for each subgroup and identify disparities?

A.Cloud Data Loss Prevention (DLP)
B.Vertex AI Explainable AI
C.Vertex AI Model Evaluation
D.Vertex AI Model Monitoring
AnswerC

Vertex AI Model Evaluation computes performance metrics and slices results by configured demographic columns, exposing fairness disparities across subgroups. This directly satisfies the stem's need to evaluate each subgroup using BigQuery predictions and ground truth labels.

Why this answer

Vertex AI Model Evaluation is the correct answer because it provides built-in functionality to compute performance metrics (e.g., accuracy, precision, recall, AUC) for specific slices of data, such as demographic subgroups. It allows you to evaluate fairness by comparing metrics across these subgroups to identify disparities. This service directly supports the requirement to monitor fairness by analyzing model performance on different segments.

Exam trap

The trap here is confusing model evaluation with model monitoring or explainability; candidates might think Model Monitoring handles fairness because it monitors models, but it only tracks drift, not performance disparities.

How to eliminate wrong answers

Option A is wrong because Cloud Data Loss Prevention (DLP) is designed to discover, classify, and protect sensitive data like PII, not to evaluate model performance or fairness. Option B is wrong because Vertex AI Explainable AI provides feature attributions (e.g., SHAP values) to explain individual predictions, but it does not compute subgroup performance metrics or identify disparities. Option D is wrong because Vertex AI Model Monitoring detects drift and anomalies in production data, but it does not evaluate model performance against ground truth for fairness analysis.

277
MCQmedium

A retail company wants to forecast daily sales for each of its 500 stores for the next 90 days. They have three years of historical daily sales data stored in BigQuery, including promotions, holidays, and store attributes. The data science team has minimal ML expertise and wants to use SQL to build and deploy the model with minimal coding. Which approach should they use?

A.Build a custom LSTM model in Vertex AI using TensorFlow, training on the historical sales data.
B.Use AutoML Forecasting in Vertex AI, exporting the data from BigQuery to Cloud Storage.
C.Train a linear regression model in BigQuery ML with store ID as a feature and date as a numeric feature.
D.Create an ARIMA_PLUS model in BigQuery ML using the sales time series and specify the store ID as the time series identifier.
AnswerD

ARIMA_PLUS is designed for univariate time series forecasting and can handle multiple time series by specifying a time series ID column. It automatically handles seasonality, holidays, and trends, and can incorporate additional regressors like promotions. This requires only SQL, matching the team's low-code requirement and the need to forecast per store.

Why this answer

BigQuery ML's ARIMA_PLUS is purpose-built for time series forecasting and can handle multiple related time series by specifying an identifier column. It automatically models seasonality, holidays, and trends, and supports additional regressors. This requires only SQL, making it ideal for teams with limited ML expertise who need per-store forecasts.

Exam trap

The trap here is assuming that any regression model can handle time series forecasting, but linear regression fails to capture temporal patterns like seasonality and trends.

278
MCQeasy

What is the purpose of the 'importer' component in Vertex AI Pipelines?

A.To import data from external sources into the pipeline.
B.To import existing ML artifacts (e.g., models, datasets) into a pipeline as inputs.
C.To import Python libraries into the pipeline environment.
D.To import pipeline definitions from other projects.
AnswerB

The importer component registers an existing artefact, such as a model or dataset already in Cloud Storage or Artifact Registry, as a pipeline input without rerunning upstream steps. This lets pipelines consume pre-existing ML artefacts as declared inputs.

Why this answer

The Importer component in Vertex AI Pipelines (a Kubeflow Pipelines component) brings existing ML artifacts — such as pre-trained models, datasets, or other resources already registered in Vertex AI — into a pipeline as inputs. It creates an artifact reference so downstream components can consume it without regenerating it. This is essential for reusing models or datasets across pipeline runs.

Exam trap

PMLE often tests whether candidates confuse the Importer (existing Vertex AI artifacts) with data ingestion components (external raw data) — the word 'import' is deliberately ambiguous.

How to eliminate wrong answers

Option A is wrong because importing raw data from external sources is done by data ingestion components (e.g., custom Python components or BigQuery loaders), not the Importer — the Importer works with existing Vertex AI artifacts, not arbitrary external data. Option C is wrong because Python library imports are handled by container images and requirements files, not by a pipeline component. Option D is wrong because pipeline definitions are compiled and submitted via the Vertex AI SDK or templates; the Importer does not import pipeline definitions from other projects.

279
MCQeasy

A company wants to classify customer support emails into categories like 'billing', 'technical', or 'account'. They have labeled email text data. Which AutoML solution should they use?

A.AutoML Tables
B.AutoML Natural Language
C.AutoML Video
D.AutoML Vision
AnswerB

AutoML Natural Language performs text classification on labelled data, training a model that assigns support emails to categories such as billing, technical or account, matching the labelled email text and multi-class requirement in the stem.

Why this answer

AutoML Natural Language is designed for text classification tasks, including sentiment analysis, entity extraction, and content categorization. Since the company has labeled email text data and wants to classify emails into categories like 'billing', 'technical', or 'account', AutoML Natural Language is the correct choice. It handles text data natively and provides a simple interface to train custom models without requiring deep ML expertise.

Exam trap

PMLE often tests the distinction between AutoML services based on data modality; candidates may confuse AutoML Natural Language with AutoML Tables when the data is text but stored in a tabular format.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for structured tabular data, not unstructured text. Option C is wrong because AutoML Video is for video classification, object tracking, and action recognition, not text. Option D is wrong because AutoML Vision is for image classification, object detection, and image segmentation, not text.

280
MCQmedium

You are training a scikit-learn random forest on a dataset that fits in memory using Vertex AI custom training. The prototype notebook took 15 minutes, but the Vertex AI job takes over an hour and occasionally fails with a resource error. You want the production job to complete reliably without changing the model or preprocessing. What should you do?

A.Switch the training job to use a preemptible VM to reduce contention on shared resources.
B.Select a machine type with more memory and vCPUs, and set the appropriate boot disk size for the training job.
C.Configure the training job to run on a single n1-standard-4 machine with no accelerator.
D.Increase the number of worker replicas and use a distributed reduction strategy for the random forest.
AnswerB

Vertex AI custom training lets you choose a machine type and boot disk for the worker pool. A scikit-learn job that fits in memory but fails on a small default machine benefits from more RAM and vCPUs. Increasing available memory prevents out-of-memory failures, and a larger boot disk ensures the container and any temporary files have sufficient space, matching the reliability goal without changing the model.

Why this answer

The job fits in memory but fails on the default Vertex AI worker configuration, which points to insufficient RAM or vCPUs on the training machine. Choosing a larger machine type and an adequate boot disk gives the scikit-learn process the resources it needs. Distributed training is unnecessary and would require changing the algorithm, while preemptible VMs reduce reliability rather than improve it.

Exam trap

The trap here is assuming that any training job that fails must be fixed with distributed training, when a single-machine resource increase is the correct and simpler remedy.

281
MCQhard

In a Vertex AI Pipeline, a component produces a Metrics artifact that includes an evaluation metric. The engineer wants to use this metric value as a condition to decide whether to deploy the model. However, the metric value is stored in the artifact's metadata and not directly as a pipeline parameter. How can the engineer pass the metric value to a downstream conditional task?

A.Configure the component that produces the Metrics artifact to also output the metric as a pipeline parameter.
B.Use the importer component to convert the artifact into a parameter.
C.Add a component that reads the artifact's metadata and outputs the metric as a parameter, then use that parameter in the condition.
D.Use the artifact directly in the dsl.If condition, as artifacts are comparable.
AnswerC

Metrics stored in artifact metadata are not pipeline parameters, so dsl.If cannot consume them directly. A component that reads the artifact's metadata and emits the metric as an output parameter makes the value usable in the downstream condition, satisfying the stem's constraint.

Why this answer

Vertex AI Pipeline conditions require pipeline parameters (typed values) to evaluate expressions like `dsl.If`. A Metrics artifact's metadata is stored as an artifact property, not a pipeline parameter, so a custom component must read that metadata and output the metric as a parameter. This parameter can then be used in the `dsl.If` condition to control downstream deployment.

Exam trap

A common trap in this Google exam is the misconception that artifact metadata can be directly used in pipeline conditions, but conditions require typed parameters, not artifact objects or their metadata fields.

How to eliminate wrong answers

Option A is wrong because modifying the upstream component to output the metric as a pipeline parameter would require changing the component's implementation, which may not be feasible if the component is from a shared library or third-party. Option B is wrong because the importer component is designed to bring external artifacts into the pipeline, not to extract metadata from an existing artifact and convert it into a parameter. Option D is wrong because artifacts are not directly comparable in `dsl.If` conditions; conditions only work with pipeline parameters (e.g., integers, strings), not with artifact objects or their metadata.

282
Multi-Selectmedium

You are deploying a prototype model to Vertex AI for online prediction. The model was trained with a custom preprocessing step that must run on raw JSON input before inference. You need the endpoint to return predictions with minimal latency and to support rolling updates of new model versions without downtime. (Choose two.)

Select 2 answers
A.Deploy the model to an endpoint with a traffic split between the current version and a new version, gradually shifting traffic to the new version.
B.Deploy the preprocessing step as a separate Vertex AI endpoint and call it from the client before sending data to the model endpoint.
C.Use a Vertex AI batch prediction job for the initial rollout and switch to online prediction after the model is validated.
D.Store the preprocessing logic in a Cloud Function and invoke it from the model's prediction container at request time.
E.Package the preprocessing and the model into a single custom container that implements the Vertex AI HTTP prediction contract.
AnswersA, E

Vertex AI endpoints support deploying multiple models and assigning a traffic split percentage to each deployed model. This enables canary or rolling updates where a small share of traffic goes to the new version first, and the split is adjusted as confidence grows. It provides zero-downtime updates because both versions serve simultaneously during the transition.

Why this answer

Bundling preprocessing and model inference into one custom container that follows the Vertex AI HTTP prediction contract keeps the request path short and self-contained, which minimizes latency. Deploying model versions to the same endpoint with a traffic split enables canary or rolling updates, so new versions can be validated with a small share of live traffic before full promotion, achieving zero-downtime releases.

Exam trap

The trap here is separating preprocessing into its own service or using batch prediction for rollout, when the latency and zero-downtime goals are best met by a single container plus endpoint traffic splitting.

283
MCQeasy

An ML engineer is monitoring a Vertex AI Feature Store used for online serving. Which metrics are most important to track for ensuring low-latency online serving?

A.Number of feature stores and feature values.
B.Storage utilization and write throughput to the feature store.
C.Batch export duration and number of exported features.
D.Feature value retrieval latency (p99) and error rate.
AnswerD

Online serving latency is dominated by feature retrieval, so p99 retrieval latency and error rate directly expose slow or failing lookups that breach the low-latency constraint; aggregate CPU or storage metrics cannot reveal per-request serving delays.

Why this answer

For online serving, the primary concern is the latency and reliability of feature value retrieval at inference time. The p99 retrieval latency directly measures the worst-case delay experienced by users, while the error rate captures failures that could cause serving disruptions. Other metrics like storage utilization or batch export duration are relevant for offline or batch pipelines, not real-time serving.

Exam trap

The trap here is that candidates confuse metrics for offline batch operations (like export duration) with those for online serving, or assume that storage-level metrics (like utilization) are sufficient for performance monitoring, when in fact only retrieval latency and error rate directly reflect the serving quality.

How to eliminate wrong answers

Option A is wrong because the number of feature stores and feature values does not directly impact serving latency; it is a capacity planning metric, not a performance indicator. Option B is wrong because storage utilization and write throughput are important for data ingestion and maintenance, but they do not measure the online retrieval performance that affects inference latency. Option C is wrong because batch export duration and number of exported features pertain to offline batch serving or data export jobs, not the low-latency online serving path.

284
MCQmedium

A hospital's radiology department wants to build a model that flags possible pneumonia on chest X-rays. They have 8,000 labeled DICOM studies in a Cloud Storage bucket and no in-house data science staff. They need a managed service that can ingest the images, train a classifier, and provide an endpoint for their viewing software. What should they do?

A.Use Vertex AI Vision to build a pipeline that counts objects in the X-ray streams
B.Use the Cloud Vision API label detection feature on each X-ray
C.Create a Vertex AI dataset from the images and train an AutoML image classification model
D.Deploy a pre-trained TensorFlow Hub model to Vertex AI Endpoints without retraining
AnswerC

AutoML image classification accepts labeled images referenced from Cloud Storage, trains a managed classifier, and deploys an endpoint callable by the viewing software. It requires no model code, fitting a radiology team without data scientists who need a fast, managed path to a pneumonia flagger.

Why this answer

AutoML image classification is the managed, low-code route for a labeled image set: it reads images from Cloud Storage, trains and tunes a classifier, and serves predictions through an endpoint the viewing software can call. The alternatives either use fixed-label APIs, skip the essential fine-tuning, or apply a streaming analytics product instead of a diagnostic classifier.

Exam trap

The trap here is reaching for a general Vision API because it 'sees images,' when it cannot be trained on the department's pneumonia labels.

285
MCQeasy

A data science team uses Vertex AI Workbench and wants to share notebooks with version history. Which service should they use?

A.Artifact Registry
B.Cloud Storage
C.Data Catalog
D.Cloud Source Repositories
AnswerD

Cloud Source Repositories provides Git-based version control, so notebooks committed from Vertex AI Workbench retain full commit history and can be shared and cloned by teammates. Workbench itself stores notebooks locally without versioning, and Cloud Storage offers object storage rather than revision tracking.

Why this answer

Cloud Source Repositories (CSR) is the correct choice because it provides Git-based version control for notebooks, enabling teams to track changes, collaborate, and maintain a full version history. Vertex AI Workbench integrates natively with CSR, allowing users to clone, commit, and push notebook files directly from the JupyterLab interface, which is essential for collaborative development with revision tracking.

Exam trap

Google Cloud often tests the distinction between storage services (Cloud Storage) and version control services (Cloud Source Repositories), leading candidates to choose Cloud Storage because it has object versioning, but it lacks the collaborative Git workflow required for notebook version history.

How to eliminate wrong answers

Option A is wrong because Artifact Registry is designed for storing and managing container images and ML artifacts (e.g., models, packages), not for version-controlling notebook files or providing a Git-based history. Option B is wrong because Cloud Storage is an object store for unstructured data; it supports object versioning but lacks the branching, merging, and collaborative workflow features of a Git repository, making it unsuitable for notebook version history. Option C is wrong because Data Catalog is a metadata management service for discovering and tagging assets (e.g., datasets, models), not a version control system for code or notebooks.

286
Multi-Selecthard

An organization is deploying a mission-critical model on Vertex AI Endpoints. They need to ensure high availability and meet a strict SLO of 99.9% uptime. Which THREE steps should they take? (Choose 3)

Select 3 answers
A.Use Cloud CDN to cache responses.
B.Set minReplicas to at least 2 to ensure redundancy within a region.
C.Use a single large instance instead of multiple small ones.
D.Deploy the endpoint in multiple regions.
E.Configure health checks to detect and replace unhealthy instances.
AnswersB, D, E

Setting minReplicas to at least two keeps additional replicas running within the region, so a single node failure does not take the endpoint offline. This removes the single point of failure underpinning the 99.9% uptime SLO.

Why this answer

Option B is correct because setting minReplicas to at least 2 on a Vertex AI Endpoint ensures that multiple model server replicas run within the region, so a single replica failure or rolling update does not cause an outage and the 99.9% SLO can be maintained. Option D is correct because deploying the endpoint in multiple regions provides geographic redundancy, so a regional outage of Vertex AI or its underlying infrastructure does not take the mission-critical model offline. Option E is correct because configuring health checks (Vertex AI endpoint health checks/readiness probes) lets the service detect unhealthy replicas and route traffic away from or replace them, preserving availability.

Option A is not appropriate because Cloud CDN caches HTTP responses at the edge and does not address the availability of the Vertex AI prediction backend or its SLO. Option C is not appropriate because using a single large instance creates a single point of failure and provides no redundancy, which works against high availability.

Exam trap

PMLE often tests the misconception that a single large instance provides better availability than multiple small ones — candidates confuse vertical scaling (more power) with horizontal scaling (more redundancy), but only the latter provides fault tolerance.

287
MCQeasy

Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?

A.Use global variables to share state between components.
B.Pass data through Cloud Storage URIs rather than in-memory.
C.Write component outputs to a database with timestamps.
D.Use the same output name for all runs to avoid duplication.
AnswerB

Cloud Storage URIs persist data outside component memory, so a retried or rerun component reads the same immutable input and writes the same output, satisfying idempotency. In-memory passing loses state between container executions and cannot guarantee reproducible reruns.

Why this answer

Passing data through Cloud Storage URIs ensures that component outputs are stored persistently and can be retrieved by downstream components, even if the original component instance is terminated or scaled down. This aligns with the principle of idempotency because the same input will always produce the same output stored at the same URI, and re-running the component will not cause side effects or data loss. In contrast, in-memory data is ephemeral and tied to a specific runtime instance, breaking idempotency across retries or parallel executions.

Exam trap

The Google PMLE exam often tests the misconception that idempotency is about avoiding duplication of output names or using timestamps for uniqueness, when in fact idempotency requires that repeated executions produce the same result without side effects, which is achieved by using immutable, deterministic storage like Cloud Storage URIs rather than mutable state or time-dependent writes.

How to eliminate wrong answers

Option A is wrong because using global variables to share state between components introduces mutable shared state that can cause non-deterministic behavior across retries or parallel runs, violating idempotency. Option C is wrong because writing component outputs to a database with timestamps introduces a side effect that changes with each run (different timestamps), making the component non-idempotent; idempotent components should produce the same output regardless of how many times they are executed. Option D is wrong because using the same output name for all runs does not guarantee idempotency; it can lead to overwriting or collision of outputs, and idempotency requires that repeated executions produce the same result without unintended side effects, not just the same output name.

288
MCQeasy

A data scientist has trained a model using Vertex AI Training and wants to deploy it to a Vertex AI Endpoint for online predictions. Which orchestration service should be used to automate the deployment step after training completes?

A.Vertex AI Pipelines
B.App Engine
C.Cloud Functions
D.Cloud Build
AnswerA

Vertex AI Pipelines orchestrates the full ML workflow as a directed acyclic graph, so a deployment component can be triggered automatically once the training component succeeds. This satisfies the stem's requirement to automate deployment to a Vertex AI Endpoint after training completes, without manual intervention.

Why this answer

Vertex AI Pipelines is the correct orchestration service because it is purpose-built for automating and managing end-to-end ML workflows on Google Cloud. It allows you to define a pipeline that includes both the training step (using Vertex AI Training) and the subsequent deployment step (creating or updating a Vertex AI Endpoint) as a single, repeatable, and monitored workflow. This ensures that after training completes, the model is automatically deployed without manual intervention, leveraging the pipeline's ability to pass artifacts and trigger conditional logic.

Exam trap

Google Cloud often tests the distinction between general-purpose compute services (Cloud Functions, App Engine) and ML-specific orchestration tools (Vertex AI Pipelines), trapping candidates who think any serverless or CI/CD tool can handle the unique requirements of ML workflow automation.

How to eliminate wrong answers

Option B (App Engine) is wrong because it is a platform-as-a-service (PaaS) for building and hosting web applications, not an ML pipeline orchestrator; it lacks native integration with Vertex AI Training and Endpoint APIs for automated model deployment. Option C (Cloud Functions) is wrong because it is a serverless compute service for event-driven, single-purpose functions, not designed for orchestrating multi-step ML workflows with dependencies and artifact tracking. Option D (Cloud Build) is wrong because it is a CI/CD service primarily for building, testing, and deploying software artifacts (e.g., container images), not for orchestrating ML pipelines that involve training jobs and endpoint deployments with state management.

289
Multi-Selectmedium

Your team is preparing to hand a trained model to a separate operations team that will deploy it to a Vertex AI endpoint. The operations team needs to understand the model's input schema, the training run that produced it, and which alias currently points to production. Which two Vertex AI resources should you share with them to provide this information? (Choose two.)

Select 2 answers
A.The Vertex AI Pipeline job resource that was used to schedule nightly retraining.
B.The Cloud Storage bucket containing the raw training data.
C.The registered model resource in Vertex AI Model Registry, including its version and alias.
D.The Vertex AI Feature Store online store resource used during training.
E.The Vertex AI Experiments run that logged the training parameters, metrics, and input schema.
AnswersC, E

The Model Registry resource holds the model version, its metadata, and aliases such as 'production'. Sharing this resource name lets the operations team see exactly which version is aliased for production and deploy it, satisfying the alias and versioning part of the handoff.

Why this answer

A complete handoff to operations requires the registered model resource for version and alias information, and the training run record for parameters, metrics, and input schema. Feature stores, pipeline jobs, and raw data buckets serve different purposes and do not provide the deployment team with the model-specific governance details they need.

Exam trap

The trap here is equating data access with model handoff, assuming the operations team needs the raw training data rather than the model's schema and version metadata.

290
MCQeasy

A machine learning team uses Vertex AI Pipelines to orchestrate training workflows. They want to share pipeline runs and artifacts with stakeholders who do not have Google Cloud accounts. What should they do?

A.Generate a pipeline run report using the Vertex AI Pipelines SDK and export it as a static HTML file to share.
B.Use Vertex AI Pipelines' built-in sharing feature to generate a public URL for the pipeline run.
C.Export the pipeline run's metadata and artifacts to a Cloud Storage bucket and grant public access.
D.Create a custom dashboard in Looker Studio that reads from Vertex ML Metadata and share it with the stakeholders.
AnswerA

The Vertex AI Pipelines SDK allows generating a detailed report of a pipeline run, including artifacts and metrics, which can be exported as a static HTML file. This file can be shared with anyone, regardless of Google Cloud account, providing a snapshot of the run. It is a secure and straightforward way to share information with external stakeholders.

Why this answer

Exporting a pipeline run report as a static HTML file using the Vertex AI Pipelines SDK enables sharing with stakeholders who lack Google Cloud accounts. This method provides a comprehensive view of the run's artifacts and metrics without requiring access to the cloud environment, ensuring both security and accessibility.

Exam trap

The trap here is assuming that Vertex AI Pipelines has a built-in public sharing URL or that public bucket access is acceptable, when the correct approach is to generate a static report for external sharing.

291
MCQeasy

A marketing team wants to build a model that predicts customer lifetime value (CLV) using historical transaction data. They are comfortable with spreadsheets but have no coding experience. They need a low-code solution that automatically handles feature engineering and model selection. Which Google Cloud service should they use?

A.AI Platform Training
B.Vertex AI AutoML
C.BigQuery ML
D.Vertex AI Workbench
AnswerB

Vertex AI AutoML provides a fully managed, no-code environment where users can upload data and automatically train models with feature engineering and hyperparameter tuning handled by the service. It supports tabular data for regression tasks like predicting CLV. The marketing team can use the UI without writing code, making it the ideal low-code solution.

Why this answer

Vertex AI AutoML is designed for users with limited ML expertise to build models without writing code. It automates feature engineering, model selection, and hyperparameter tuning. For predicting CLV from historical data, the marketing team can simply upload a dataset and let AutoML train a regression model.

Other options require SQL or Python coding, which the team lacks.

Exam trap

The trap here is assuming that BigQuery ML is a no-code solution because it uses SQL, but it still requires query writing and manual model configuration.

292
MCQeasy

A marketing team wants to use a pre-built natural language processing (NLP) model from Vertex AI Model Garden to analyze customer feedback. They need to extract sentiment from text data stored in Cloud Storage. The team has no experience with model serving infrastructure. Which deployment option minimizes operational overhead?

A.Deploy the model as a Cloud Function invoked by Cloud Storage events.
B.Deploy the model as a Cloud Run service using a custom Docker container.
C.Deploy the model on App Engine flexible environment.
D.Deploy the model to a Vertex AI Endpoint directly from Model Garden.
AnswerD

Deploying directly from Model Garden to a Vertex AI Endpoint gives a fully managed serving endpoint with autoscaling and no infrastructure to maintain. This satisfies the constraint of minimising operational overhead for a team with no model-serving experience.

Why this answer

Deploying directly to a Vertex AI Endpoint from Model Garden eliminates all infrastructure management. Vertex AI handles model serving, scaling, and monitoring automatically, which is ideal for a team with no experience in model serving infrastructure. This is a fully managed, serverless deployment that requires no containerization or server configuration.

Exam trap

The trap here is that candidates often assume Cloud Functions or Cloud Run are simpler because they are 'serverless,' but they fail to recognize that deploying a large NLP model requires specialized infrastructure (GPUs, model serving frameworks) that these services do not natively provide without significant custom work.

How to eliminate wrong answers

Option A is wrong because Cloud Functions are designed for lightweight, stateless event-driven code, not for hosting large NLP models with significant memory and GPU requirements; they also lack built-in model serving capabilities like autoscaling for inference. Option B is wrong because deploying as a Cloud Run service with a custom Docker container requires the team to containerize the model, manage dependencies, and configure scaling, which introduces significant operational overhead for a team with no serving experience. Option C is wrong because App Engine flexible environment still requires the team to build a custom runtime, manage instances, and handle model dependencies, and it is not optimized for ML inference workloads like Vertex AI endpoints.

293
MCQeasy

An organization wants to implement continuous training for a model that serves predictions via Vertex AI Endpoints. Which approach best automates the retrain-deploy cycle?

A.Schedule a Vertex AI Pipeline to retrain and conditionally deploy
B.Use Vertex AI Model Registry to auto-deploy on new model upload
C.Manually retrain and deploy monthly
D.Use Cloud Composer to schedule retraining only
E.Use a Cloud Function to retrain the model and update the endpoint
AnswerA

A scheduled Vertex AI Pipeline can retrain on a defined cadence and use a condition to deploy only when evaluation metrics pass, directly automating the retrain-deploy cycle for models served through Vertex AI Endpoints without manual intervention.

Why this answer

Vertex AI Pipelines can be scheduled to run a retraining workflow and include a conditional step that deploys the new model to the endpoint only if it passes validation (e.g., evaluation metrics meet a threshold). This fully automates the retrain-deploy cycle without manual intervention, leveraging the pipeline's orchestration capabilities.

Exam trap

Google Cloud often tests the distinction between partial automation (e.g., only retraining or only deploying) and full end-to-end automation; the trap here is that candidates may choose an option that automates only one part of the cycle (like retraining with Cloud Composer or auto-deployment with Model Registry) and miss that the question requires both retraining and deployment to be automated in a single, orchestrated workflow.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Registry auto-deploys a model to an endpoint only if the endpoint is configured for automatic deployment, but it does not trigger retraining; it merely deploys an already uploaded model, so it does not automate the retrain step. Option C is wrong because manual retraining and deployment monthly is not automated and defeats the purpose of continuous training. Option D is wrong because Cloud Composer (Airflow) can schedule retraining, but it does not automatically deploy the model to the endpoint; deployment requires an additional step, so it does not fully automate the cycle.

Option E is wrong because a Cloud Function can trigger retraining and update an endpoint, but it lacks built-in orchestration for complex workflows like conditional deployment based on model evaluation, and it is less robust for managing dependencies and state compared to a pipeline.

294
Multi-Selecteasy

Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)

Select 2 answers
A.Use a larger machine type with more memory
B.Optimize the model using quantization or pruning
C.Deploy the model in the same region as the clients
D.Use batch prediction instead of online prediction
E.Enable model caching at the endpoint
AnswersB, C

Quantization reduces weight precision and pruning removes redundant parameters, shrinking the model so each inference requires fewer compute cycles and less memory bandwidth. This directly lowers per-request serving latency on Vertex AI Endpoints, satisfying the stem's latency-reduction constraint without changing the endpoint's infrastructure.

Why this answer

Option B is correct because quantization (e.g., reducing weights from FP32 to INT8) and pruning (removing redundant parameters) shrink the model size and reduce the compute required per inference, directly lowering prediction latency on Vertex AI Endpoints. Option C is correct because deploying the model in the same region as the clients minimizes network round-trip time, which is a significant component of end-to-end serving latency for online predictions. Option A is not a best practice for latency specifically: a larger machine with more memory may help with throughput or memory-bound models, but it does not inherently reduce per-request latency and increases cost.

Option D is wrong because batch prediction is an asynchronous, offline mode that does not serve real-time requests and is not a latency optimization for online endpoints. Option E is not a supported Vertex AI Endpoint feature; there is no endpoint-level 'model caching' toggle that reduces serving latency.

Exam trap

PMLE often tests whether candidates confuse throughput optimizations (bigger machines, batching) with latency optimizations — the trap is picking 'larger machine type' when the question specifically asks about latency.

295
MCQeasy

A user receives the error "Deployment failed due to insufficient memory. Please use a machine type with higher memory." when deploying an AutoML model. What should they do?

A.Change the region to us-west1
B.Use machine type n1-highmem-2
C.Increase the min-replica-count to 2
D.Remove the traffic-split flag
AnswerB

The error indicates the chosen machine type has insufficient RAM for the AutoML model. n1-highmem-2 supplies 13 GB of memory with a high memory-to-vCPU ratio, directly resolving the constraint, whereas standard or compute-optimised types offer less memory per core.

Why this answer

When deploying an AutoML model, memory constraints are a common cause of deployment failures. Using a machine type with higher memory, such as n1-highmem-2, helps ensure the model can be loaded and served without out-of-memory errors. The other options do not address memory requirements: changing the region does not affect compute resources, increasing the min replica count does not increase per-instance memory, and removing traffic-split flags does not resolve memory issues.

Exam trap

The trap here is that candidates often confuse scaling (increasing replicas) with resource allocation (increasing memory per replica), leading them to choose Option C instead of addressing the per-instance memory bottleneck.

How to eliminate wrong answers

Option A is wrong because changing the region to `us-west1` does not affect the memory capacity of the machine type; the error is due to insufficient memory, not regional availability or latency. Option C is wrong because increasing `min-replica-count` to 2 only adds more replicas for scaling, but each replica still uses the same underpowered machine type, so the OOM error persists. Option D is wrong because removing the `traffic-split` flag would disrupt traffic routing but does not address the root cause of insufficient memory for model loading.

296
MCQmedium

You are deploying a custom model to Vertex AI for online prediction. The model requires a preprocessing step that normalizes input features. You want to ensure that the same preprocessing is applied during both training and serving to avoid training-serving skew. What should you do?

A.Use Vertex AI Feature Store to serve features and apply preprocessing at training time only.
B.Perform preprocessing on the client side before sending data to the model for prediction.
C.Export the preprocessing as a TensorFlow Transform (tf.Transform) function and include it in both training and serving graphs.
D.Include the preprocessing logic in the training script and replicate it in the serving container.
AnswerC

tf.Transform allows you to define preprocessing as a function that can be applied consistently during training and serving. By exporting the transform function and including it in both the training and serving graphs, you ensure identical preprocessing. This is a best practice for avoiding training-serving skew, as the same code is used in both phases.

Why this answer

Using TensorFlow Transform (tf.Transform) to define preprocessing ensures that the same transformations are applied during both training and serving. The transform function is included in the model graph, so the serving container automatically applies the same logic. This eliminates manual replication and reduces the risk of training-serving skew.

Exam trap

The trap here is thinking that duplicating preprocessing code in both training and serving is sufficient, but it often leads to skew due to maintenance issues; using a shared transform function is more reliable.

297
MCQeasy

A data scientist is defining a Vertex AI pipeline and needs to include a step that imports a pre-existing model from Cloud Storage into the pipeline as an artifact. Which Kubeflow Pipelines SDK v2 component should they use?

A.dsl.Collected
B.dsl.importer
C.dsl.Importer
D.dsl.Artifact
AnswerB

dsl.importer registers an existing Cloud Storage artefact, such as a trained model, as a pipeline artefact without executing a container to recreate it. This satisfies the requirement to bring a pre-existing model into the pipeline graph as a usable input.

Why this answer

The `dsl.importer` component in Kubeflow Pipelines SDK v2 is specifically designed to import existing artifacts (such as models, datasets, or metrics) from external storage (e.g., Cloud Storage) into a pipeline as a pipeline artifact. It allows you to reference a pre-existing model without retraining or re-uploading, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse the Python class naming convention (capitalized `Importer`) with the actual SDK v2 function name (lowercase `importer`), or mistakenly think `dsl.Artifact` can import artifacts when it only defines the artifact schema.

How to eliminate wrong answers

Option A is wrong because `dsl.Collected` is not a valid Kubeflow Pipelines SDK v2 component; it does not exist in the API. Option C is wrong because `dsl.Importer` (capital 'I') is not a valid class or function in the SDK v2; the correct name is all lowercase `dsl.importer`. Option D is wrong because `dsl.Artifact` is a base class for defining custom artifact types, not a component for importing artifacts into a pipeline.

298
MCQmedium

A data science team uses TFX to train and deploy a model on Vertex AI. They want automated monitoring for pipeline health. Which set of metrics should they monitor to quickly detect issues in the training pipeline?

A.Prediction request count, latency, and error rate on the serving endpoint.
B.Pipeline execution status (success/failure), component completion times, and data validation anomalies.
C.Number of pipeline runs, average CPU utilization, and memory usage.
D.Model accuracy, precision, and recall on the evaluation dataset.
AnswerB

TFX pipeline health hinges on orchestration-level signals, not model accuracy. Execution status catches failed runs, component completion times expose bottlenecks or stalls in individual steps, and data validation anomalies flag schema or distribution problems before training proceeds, directly satisfying the requirement to detect training pipeline issues quickly.

Why this answer

The question specifically asks about monitoring the training pipeline's health, not the serving infrastructure. Pipeline execution status directly indicates whether the pipeline ran successfully, component completion times help identify bottlenecks or failures, and data validation anomalies catch data quality issues early in the pipeline — all of which are essential for detecting issues in the training pipeline itself.

Exam trap

The trap here is that candidates confuse serving endpoint metrics (like latency and error rate) with pipeline health metrics, because both are part of an ML system, but the question explicitly asks about the training pipeline, not the serving infrastructure.

How to eliminate wrong answers

Option A is wrong because prediction request count, latency, and error rate are metrics for monitoring the serving endpoint (model serving), not the training pipeline. Option C is wrong because number of pipeline runs, average CPU utilization, and memory usage are infrastructure-level metrics that do not directly indicate pipeline health or data quality issues. Option D is wrong because model accuracy, precision, and recall are evaluation metrics for model performance, not for detecting issues in the training pipeline's execution or data validation.

299
MCQhard

You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?

A.Model parallelism using pipeline parallelism
B.Data parallelism with MultiWorkerMirroredStrategy
C.Mixed precision training (FP16)
D.Data parallelism with tf.distribute.MirroredStrategy
AnswerA

Pipeline parallelism splits the model's layers across GPUs, with each device holding a subset and passing activations onward. This addresses the constraint that seven billion parameters exceed single-GPU memory, unlike data parallelism which replicates the full model per device.

Why this answer

When a model is too large to fit on a single GPU, model parallelism is required to split the model's layers across multiple devices. Pipeline parallelism is a specific form of model parallelism that partitions layers into stages and pipelines micro-batches across devices, enabling training of models like a 7B-parameter LLM across multiple GPUs. This is the correct approach when memory, not throughput, is the binding constraint.

Exam trap

PMLE often tests the misconception that data parallelism solves memory constraints, when in fact it replicates the model and only model/pipeline parallelism addresses models too large for one GPU.

How to eliminate wrong answers

Option B is wrong because data parallelism with MultiWorkerMirroredStrategy replicates the full model on each worker — if the model does not fit on one GPU, it will not fit on any, so this approach fails. Option C is wrong because mixed precision (FP16) reduces memory footprint but does not solve the fundamental problem of a model that exceeds single-GPU memory; it is an optimization, not a distribution strategy. Option D is wrong because MirroredStrategy is single-machine data parallelism that also replicates the full model on each GPU, which is impossible when the model exceeds single-GPU memory.

300
Multi-Selectmedium

Which THREE actions are best practices for managing ML models in production on Google Cloud? (Choose 3)

Select 3 answers
A.Manually tune hyperparameters for each retraining run.
B.Monitor model performance and data drift continuously.
C.Use a central model registry for model governance.
D.Version all model artifacts and training datasets.
E.Store all raw training data indefinitely for auditability.
AnswersB, C, D

Continuous monitoring of model performance and data drift detects degradation after deployment, satisfying the production oversight requirement. Unlike static validation, it tracks live inference distributions against training baselines, triggering retraining when drift exceeds thresholds. This directly addresses the stem's demand for ongoing operational management of deployed ML models.

Why this answer

Option B is correct because continuous monitoring of model performance and data drift is essential in production to detect degradation and trigger retraining before business impact occurs. Option C is correct because a central model registry (such as Vertex AI Model Registry) provides governance, lineage, and controlled promotion of models across environments. Option D is correct because versioning model artifacts and training datasets ensures reproducibility, traceability, and rollback capability for every deployed model.

Option A is not a best practice because manual hyperparameter tuning does not scale and should be automated with tools like Vertex AI Vizier. Option E is not a best practice because retaining all raw training data indefinitely increases cost and compliance risk; retention should follow defined policies and lifecycle rules.

Exam trap

Google Cloud often tests the misconception that manual hyperparameter tuning is acceptable for production, when in fact automation (e.g., Vertex AI Vizier) is the recommended practice to ensure reproducibility and efficiency.

Page 3

Page 4 of 11

Page 5

All pages