Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 301–375

775 questions total · 11pages · All types, answers revealed

Page 4

Page 5 of 11

Page 6
301
MCQhard

You are training a TensorFlow model on Vertex AI using a custom container with a single Tesla T4 GPU. You notice that training is slower than expected, and GPU utilization is consistently below 20%. Profiling shows that the input pipeline is the bottleneck. Which change should you make to improve GPU utilization?

A.Switch to a larger GPU instance with more memory to reduce data loading overhead.
B.Move the dataset to a local SSD on the training VM to reduce I/O latency.
C.Increase the batch size to fully utilize GPU memory.
D.Use tf.data with prefetching and parallel data extraction to overlap data loading with GPU computation.
AnswerD

The tf.data API with prefetching and parallel extraction allows data preprocessing to occur on CPU while the GPU computes on the previous batch, effectively overlapping I/O and compute. This directly addresses the input pipeline bottleneck, increasing GPU utilization and reducing training time. It is the recommended approach for optimizing input pipelines in TensorFlow.

Why this answer

When the input pipeline is the bottleneck, the GPU waits for data. Using tf.data with prefetching and parallel extraction overlaps CPU preprocessing with GPU computation, ensuring the GPU is continuously fed. This is the standard TensorFlow optimization for such scenarios and directly improves GPU utilization.

Exam trap

The trap here is assuming that a faster GPU or more memory will fix low GPU utilization, when the issue is actually data starvation from an inefficient input pipeline.

302
Multi-Selectmedium

Which TWO metrics should you monitor to detect data drift in a batch prediction pipeline?

Select 2 answers
A.Model accuracy on recent labeled data
B.Model prediction latency
C.Feature distribution drift (e.g., KS test)
D.Prediction distribution drift
E.Training data size
AnswersC, D

Feature distribution drift compares each input feature's distribution between training and serving data using tests such as Kolmogorov-Smirnov, directly detecting changes in the input data itself, which is the definition of data drift in a batch prediction pipeline.

Why this answer

Feature distribution drift (C) is correct because data drift is detected by statistically comparing the distribution of each input feature in recent batches against the training/reference distribution, typically using tests like the Kolmogorov-Smirnov (KS) test or Population Stability Index (PSI). Prediction distribution drift (D) is correct because shifts in the distribution of model outputs (e.g., changing class proportions or score histograms) are a direct, label-free signal that the input data relationship has changed in a batch pipeline. Model accuracy on recent labeled data (A) measures performance degradation, not drift itself, and labels are usually unavailable or delayed in batch prediction.

Model prediction latency (B) is an operational performance metric unrelated to distributional change, and training data size (E) is a static dataset property, not a monitoring signal for drift.

Exam trap

Google Cloud often tests the distinction between monitoring for data drift (input distribution changes) versus monitoring for model performance degradation (accuracy), leading candidates to incorrectly select accuracy as a drift metric when it is actually a downstream effect.

303
MCQhard

A machine learning team is deploying a PyTorch model on Vertex AI Prediction for real-time inference. The model was trained with preprocessing that includes tokenization and normalization. They want to embed the preprocessing logic in the model to reduce prediction latency and avoid additional service calls. Which approach should they take?

A.Deploy the preprocessing logic as a Cloud Function and invoke it before calling the prediction endpoint
B.Wrap the preprocessing logic in a Flask application and deploy it as a separate microservice in front of the prediction endpoint
C.Use TorchScript to trace the preprocessing steps and export the entire pipeline as a single scripted model
D.Use TensorFlow Transform to convert preprocessing into a SavedModel and call it from the PyTorch model
AnswerC

TorchScript tracing captures the tokenisation and normalisation operations as graph nodes, fusing them with the PyTorch model into one serialised artefact. Vertex AI Prediction then serves this single scripted model, eliminating the separate preprocessing service call and satisfying the stem's latency-reduction constraint.

Why this answer

TorchScript allows you to trace or script a PyTorch model, including preprocessing operations like tokenization and normalization, into a single serialized artifact. By embedding preprocessing in the TorchScript model, the entire pipeline runs in one forward pass on the Vertex AI endpoint, eliminating extra service calls and reducing latency. This is the standard approach for consolidating preprocessing with a PyTorch model for real-time inference.

Exam trap

PMLE often tests the misconception that preprocessing must be a separate service — candidates pick Cloud Functions or Flask microservices because they are familiar patterns, but the question explicitly asks to embed preprocessing in the model, and TorchScript is the PyTorch-native way to do that.

How to eliminate wrong answers

Option A is wrong because a Cloud Function adds a network hop and cold-start latency, increasing prediction latency rather than reducing it — the goal is to avoid additional service calls, not add one. Option B is wrong because a separate Flask microservice in front of the endpoint introduces another network call and operational overhead, directly contradicting the requirement to reduce latency and avoid additional service calls. Option D is wrong because TensorFlow Transform produces a SavedModel for TensorFlow, not PyTorch — it cannot be directly called from a PyTorch model without a cross-framework bridge, which adds complexity and latency.

304
MCQmedium

A company uses Vertex AI AutoML to train a vision model, but the model has low accuracy. What should they do first?

A.Add more labeled images to the dataset
B.Switch to a custom model
C.Increase the training budget
D.Reduce image size to speed up training
AnswerA

Low accuracy usually stems from insufficient or unrepresentative training data, so adding more labelled images addresses the root cause first. This satisfies the scenario's need to improve the dataset before tuning hyperparameters or changing model architecture.

Why this answer

Adding more labeled images directly addresses the most common cause of low accuracy in AutoML vision models: insufficient or unrepresentative training data. Vertex AI AutoML relies on transfer learning from pre-trained models, and its performance is heavily dependent on the quality and quantity of labeled examples. Before adjusting hyperparameters or infrastructure, the first step should always be to improve the dataset, as AutoML is designed to handle model architecture and training budget automatically.

Exam trap

Google Cloud often tests the misconception that AutoML models are 'black boxes' where tuning budgets or switching to custom models is the first fix, when in reality the platform is optimized to handle those aspects automatically, and the primary lever is data quality.

How to eliminate wrong answers

Option B is wrong because switching to a custom model would require manual architecture design and hyperparameter tuning, which contradicts the low-code premise of AutoML and is not the first troubleshooting step. Option C is wrong because increasing the training budget (e.g., node hours) only helps if the model has not converged; with low accuracy, the root cause is typically data quality, not insufficient training time. Option D is wrong because reducing image size may speed up training but can discard critical features, further degrading accuracy; AutoML already handles resizing internally.

305
MCQhard

An ML engineer is using Vertex AI distributed training for a TensorFlow model that uses the MirroredStrategy. They notice that the training throughput drops significantly when moving from a single GPU to multiple GPUs on the same machine. What is the most likely cause?

A.The GPUs are not properly configured in TF_CONFIG.
B.The batch size is too small, causing each GPU to complete its forward pass quickly, but the sync wait dominates.
C.The learning rate is too high, causing instability.
D.The model uses TensorFlow 1.x instead of 2.x.
AnswerB

MirroredStrategy performs an all-reduce gradient sync across GPUs after every step. With a small batch size, each GPU's forward and backward pass finishes quickly, so communication overhead dominates and throughput falls rather than scaling with added GPUs.

Why this answer

In MirroredStrategy, all GPUs must finish their forward/backward pass before the all-reduce gradient synchronization can occur, so throughput is bounded by the slowest replica plus the sync overhead. With a small batch size, each GPU's compute time is very short, so the fixed cost of the NCCL all-reduce (and the sync wait) dominates the step time, making multi-GPU training slower than single-GPU. Increasing the per-replica batch size amortizes the synchronization cost over more compute, restoring scaling efficiency.

Exam trap

The trap here is assuming any multi-GPU slowdown must be a configuration error (TF_CONFIG, framework version) rather than recognizing that synchronization overhead in data-parallel training is amortized by batch size — a classic distributed-training performance pitfall.

How to eliminate wrong answers

Option A is wrong because TF_CONFIG is used for multi-worker distributed training (MultiWorkerMirroredStrategy, ParameterServerStrategy) to specify cluster/worker/task info, not for single-machine MirroredStrategy, which auto-discovers local GPUs via NCCL. Option C is wrong because a high learning rate causes divergence or NaN loss, not a throughput drop — it affects convergence quality, not step time. Option D is wrong because TensorFlow 1.x vs 2.x affects API style and graph/eager execution, not the fundamental synchronization overhead that causes the observed throughput regression.

306
MCQmedium

You are using DVC for data versioning in an ML project on Google Cloud. Your training data is stored in Cloud Storage. You want to track a new version of the dataset after preprocessing. Which DVC command should you use to register the changes?

A.dvc add data/processed
B.dvc push
C.dvc run -n preprocess
D.dvc commit
AnswerA

dvc add computes the hash of the processed directory, writes a .dvc file and updates .gitignore, registering the new dataset version for tracking. dvc push only uploads already-tracked data to remote storage; it does not register changes.

Why this answer

The `dvc add` command registers a file or directory with DVC, creating a .dvc metafile that captures the hash and metadata of the data, and adds the actual data to the DVC cache. After preprocessing produces a new dataset in data/processed, running `dvc add data/processed` creates the versioned pointer that DVC tracks in Git, which is exactly what 'registering the changes' means.

Exam trap

PMLE often tests the confusion between `dvc add` (register a new dataset/file) and `dvc commit` (update cache for existing stage outputs) — candidates pick `dvc commit` when the question asks about registering a brand-new dataset.

How to eliminate wrong answers

Option B is wrong because `dvc push` uploads cached data to remote storage (e.g., Cloud Storage) — it transfers data but does not create or update the .dvc metafile that registers a new version. Option C is wrong because `dvc run -n preprocess` (or `dvc stage add`) defines a pipeline stage that executes a command and tracks its outputs; it is used to create the preprocessing step, not to register an already-produced dataset. Option D is wrong because `dvc commit` writes the current state of tracked outputs to the DVC cache without re-running the pipeline — it is used after manual edits to outputs of an existing stage, not for adding a brand-new dataset to DVC tracking.

307
MCQhard

An ML engineer is running a Vertex AI pipeline that includes a data validation component and a training component. The engineer wants the pipeline to stop before training if data validation fails, but wants the validation component to record its result as an output artifact for later inspection. Which combination of pipeline features should the engineer use?

A.Configure the validation component with a retry policy of 3 so transient validation errors are retried, then let training proceed.
B.Use a dsl.Condition to skip training when validation fails, and have the validation component always succeed while writing the report.
C.Have the validation component write the report artifact and return a success status, then check the artifact contents in the training component before training.
D.Have the validation component raise an exception on failure and write a validation report artifact before raising.
AnswerD

Raising an exception causes the pipeline task to fail, which stops downstream training tasks because they depend on the validation task. Writing the validation report artifact before raising ensures the artifact is persisted and available for inspection even though the task failed. This combination satisfies both the stop-before-training and record-for-inspection requirements.

Why this answer

In Vertex AI Pipelines, a task that raises an exception fails, and dependent tasks are not executed. Writing the validation report artifact before raising ensures the artifact is captured for later inspection. The other options either let training proceed, rely on retries that do not address deterministic data failures, or move the validation check into the training component, none of which stop the pipeline before training while preserving the report.

Exam trap

The trap here is assuming that a failed component cannot produce artifacts, when artifacts written before the exception are still persisted and available for inspection.

308
Multi-Selectmedium

An ML team is optimizing an inference model for deployment on edge devices. They need to reduce the model size and improve latency while maintaining accuracy as much as possible. Which two techniques should they use? (Choose TWO.)

Select 2 answers
A.Use a larger pre-trained model as a starting point.
B.Post-training quantization to INT8.
C.Use half-precision (FP16) instead of INT8.
D.Apply weight pruning to remove small weights.
E.Increase the number of layers in the model.
AnswersB, D

Reduces size and latency with minimal accuracy loss.

Why this answer

Post-training quantization to INT8 reduces model size by converting 32-bit floating-point weights and activations to 8-bit integers, which also speeds up inference on edge devices with integer-optimized hardware. This technique typically maintains accuracy within 1-2% of the original model while significantly lowering memory footprint and latency.

Exam trap

Candidates often think that FP16 is always better than INT8 for edge devices, but INT8 offers greater size reduction and is more widely supported on edge hardware, including Google's Edge TPU.

309
MCQmedium

An ML team is scaling a prototype to production. The data pipeline currently reads from Cloud Storage and transforms data with a custom Python script. They need to handle higher throughput and add monitoring. Which approach should they take?

A.Deploy the Python script on a large Compute Engine instance with a cron job
B.Migrate the pipeline to Apache Beam on Dataflow with Cloud Monitoring
C.Rewrite the pipeline to use Pub/Sub and Cloud Functions for processing
D.Use Cloud Composer to orchestrate the Python script at scale
AnswerB

Apache Beam on Dataflow provides autoscaling runners that distribute the Python transforms across workers, satisfying the higher-throughput constraint that a single script cannot meet. Dataflow's native integration with Cloud Monitoring supplies the required pipeline metrics and alerting without custom instrumentation.

Why this answer

Apache Beam on Dataflow provides a unified programming model for batch and streaming data processing, enabling automatic scaling to handle higher throughput. Cloud Monitoring integrates natively with Dataflow to track pipeline metrics, latency, and error rates, addressing the monitoring requirement. This approach is purpose-built for production-grade data pipelines, unlike ad-hoc solutions.

Exam trap

Google Cloud often tests the distinction between orchestration (Cloud Composer) and execution (Dataflow), leading candidates to choose an orchestrator when a dedicated processing engine is required for scaling and monitoring.

How to eliminate wrong answers

Option A is wrong because deploying a Python script on a single large Compute Engine instance with a cron job does not provide horizontal scaling, fault tolerance, or built-in monitoring; it creates a single point of failure and cannot handle throughput spikes. Option C is wrong because rewriting the pipeline to use Pub/Sub and Cloud Functions is suitable for event-driven, lightweight processing but not for complex data transformations or high-throughput batch workloads; Cloud Functions have timeouts (up to 9 minutes for HTTP functions) and lack stateful processing capabilities. Option D is wrong because Cloud Composer (managed Apache Airflow) is an orchestration tool, not a data processing engine; it would still rely on the Python script's execution, inheriting its scaling and monitoring limitations without addressing the core transformation throughput.

310
Multi-Selectmedium

An ML engineer is setting up Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. They want to receive alerts when either feature skew or prediction drift exceeds a threshold. Which two configurations are required to enable these alerts? (Choose two.)

Select 2 answers
A.Enable explainable AI on the Endpoint to generate feature attributions.
B.Specify a training dataset for feature skew detection.
C.Configure an email notification channel in Cloud Monitoring.
D.Set up a Cloud Scheduler job to trigger the monitoring pipeline.
E.Define a monitoring frequency and window for drift and skew analysis.
AnswersB, E

Feature skew detection compares live input features to the training data distribution. Without specifying a training dataset, Vertex AI cannot compute the skew metric. Therefore, providing a training dataset is a required configuration for feature skew alerts. This is a fundamental step in setting up skew monitoring.

Why this answer

To enable feature skew and prediction drift alerts in Vertex AI Model Monitoring, you must specify a training dataset for skew detection and define a monitoring frequency and window. The training dataset provides the baseline for input feature distributions, while the frequency and window control how often and over what period the analysis runs. Cloud Scheduler, email notifications, and explainable AI are not required for the core monitoring functionality.

Exam trap

The trap here is assuming that external scheduling or notification setup is necessary for monitoring alerts, when Vertex AI handles scheduling internally and notifications are optional.

311
Multi-Selectmedium

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced and that artifacts are tracked. Which two of the following practices should they follow? (Choose two.)

Select 2 answers
A.Version all pipeline components and store them in a version control system.
B.Use the same container image tags for all pipeline steps to ensure consistency.
C.Manually copy all artifacts to a Cloud Storage bucket after each run.
D.Use Vertex AI Metadata to record parameters, metrics, and artifacts for each pipeline run.
E.Disable caching for all pipeline steps to ensure fresh executions.
AnswersA, D

Versioning pipeline components ensures that changes are tracked and that specific versions can be reproduced. Storing them in version control allows the team to revert to previous versions and understand the history of changes. This is essential for reproducibility and artifact tracking.

Why this answer

Versioning pipeline components and using Vertex AI Metadata are key practices for reproducibility and artifact tracking. Versioning ensures that the exact code and dependencies are captured, while Metadata automatically records parameters, metrics, and artifacts for each run, providing lineage and enabling reproducibility.

Exam trap

The trap here is thinking that manual artifact management or disabling caching is necessary for reproducibility, but automated metadata tracking and component versioning are the core practices.

312
MCQmedium

You have deployed a regression model that predicts house prices. Over the past month, the model's predictions have been consistently too high. You suspect data drift in the input features. Which monitoring metric should you prioritize to confirm this?

A.Monitor prediction drift (prediction distribution)
B.Monitor feature distribution drift using a divergence metric like Jensen-Shannon divergence
C.Monitor feature attribution drift using SHAP values
D.Monitor residual distribution drift
AnswerB

Jensen-Shannon divergence quantifies how far each input feature's live distribution has drifted from its training baseline, directly confirming whether covariate shift explains the biased predictions. Aggregate prediction metrics alone cannot isolate which features moved, so distribution-level comparison is the appropriate diagnostic.

Why this answer

The question describes a scenario where predictions are consistently too high, which is a symptom of data drift—a change in the distribution of input features. Monitoring feature distribution drift using a divergence metric like Jensen-Shannon divergence directly measures whether the input data has shifted from the training distribution, which would cause the model to make biased predictions. This is the most direct way to confirm data drift in the input features.

Exam trap

Google Cloud often tests the distinction between monitoring prediction drift (output) and feature drift (input), trapping candidates who assume that a change in predictions automatically implies data drift without verifying the input distributions.

How to eliminate wrong answers

Option A is wrong because monitoring prediction drift (prediction distribution) only tells you that the outputs have changed, not why; it does not isolate whether the cause is data drift in features or other issues like concept drift. Option C is wrong because monitoring feature attribution drift using SHAP values measures changes in feature importance, not changes in the feature distributions themselves; it can indicate which features are driving predictions differently but does not directly confirm data drift. Option D is wrong because monitoring residual distribution drift focuses on the errors (residuals) between predictions and actual values, which can be influenced by both data drift and concept drift; it does not specifically confirm data drift in input features.

313
Matchingmedium

Match each Google Cloud AI/ML service to its primary purpose.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

End-to-end ML platform for building, deploying, and managing models

Train high-quality custom ML models with minimal effort

Managed service for distributed training of ML models

Custom ASIC for accelerating ML training workloads

Create and execute ML models using SQL queries

Why these pairings

In this matching question, the correct pairings are: Option A (Vertex AI: End-to-end ML platform) is correct because Vertex AI unifies the ML workflow. Option C (Cloud AutoML: Train custom ML models with minimal coding) is correct because AutoML provides no-code training. Option E (Vision API: Analyze images for objects, faces, text) is correct because Vision API specializes in image analysis.

Option B is wrong because Vertex AI is not primarily for graphical no-code training—that is AutoML's strength. Option D is wrong because Cloud AutoML is not an end-to-end platform—that's Vertex AI. Option F is wrong because Vision API does not analyze text; that is the Natural Language API's purpose.

Exam trap

Candidates often confuse Vertex AI with Cloud AutoML, thinking Vertex AI is only for no-code training or that AutoML is the end-to-end platform. Remember: Vertex AI is the unified platform; AutoML is a component for no-code model training.

314
Multi-Selectmedium

A data science team uses Vertex AI Experiments to compare multiple model training runs. They want to capture and compare hyperparameters, metrics, and code versions for each run. Which TWO steps should they take?

Select 2 answers
A.Use Cloud Logging to capture all training outputs
B.Store code versions in Cloud Storage and link them to experiments manually
C.Log hyperparameters and metrics using the Vertex AI SDK's experiment logging functions
D.Export experiment data to BigQuery for comparison
E.Integrate the training code with Git and use the commit hash as a run parameter
AnswersC, E

The Vertex AI SDK's experiment logging functions (aiplatform.log_params and log_metrics) attach hyperparameters and metrics to a named experiment run, which is exactly what enables side-by-side comparison of runs in the Vertex AI Experiments console. Code versions are captured separately via Git.

Why this answer

Option C is correct because the Vertex AI SDK provides dedicated experiment logging functions (e.g., aiplatform.log_params() and aiplatform.log_metrics()) that record hyperparameters and metrics directly into a Vertex AI Experiment run, which is exactly what the team needs to capture and compare across runs. Option E is correct because integrating training code with Git and passing the commit hash as a run parameter ties each experiment run to a specific, reproducible code version, satisfying the requirement to capture code versions alongside hyperparameters and metrics. Option A is not appropriate because Cloud Logging captures log output, not structured experiment parameters or metrics for comparison in Vertex AI Experiments.

Option B is unnecessary and less precise than using Git commit hashes, since manually linking Cloud Storage artifacts does not automatically associate code versions with runs. Option D is not required because Vertex AI Experiments already provides comparison capabilities natively, so exporting to BigQuery is an extra step not needed for the stated goal.

Exam trap

PMLE often tests the misconception that Cloud Logging or BigQuery export are part of the experiment tracking workflow, when in fact Vertex AI Experiments requires explicit SDK logging and manual code version parameterization.

315
Multi-Selectmedium

A team has trained a sentiment analysis model using PyTorch on Vertex AI Training. They now want to deploy it for online predictions with low latency. Which TWO actions should they take? (Choose 2)

Select 2 answers
A.Create multiple model versions for A/B testing.
B.Use a machine type with a GPU for faster inference.
C.Enable batch prediction instead of online prediction.
D.Convert the model to TensorFlow SavedModel format.
E.Package the model in a custom container with a web server (e.g., FastAPI).
AnswersB, E

A GPU machine type accelerates the matrix operations underpinning PyTorch inference, directly satisfying the low-latency requirement for online predictions. Vertex AI supports attaching GPUs to prediction nodes, so tensor computations complete faster than CPU-only serving, reducing per-request response times for the sentiment model.

Why this answer

Option B is correct because deploying a PyTorch model for low-latency online predictions on Vertex AI benefits from GPU-backed machine types, which accelerate the matrix/tensor computations of deep neural network inference and reduce per-request latency. Option E is correct because Vertex AI custom containers let you serve a PyTorch model with your own web server (e.g., FastAPI) implementing the required predict/health routes, which is the standard way to deploy non-TensorFlow frameworks for online prediction. Option A is not required for low latency; model versions and A/B testing address traffic splitting and evaluation, not inference speed.

Option C is wrong because batch prediction is for asynchronous, bulk scoring and does not provide the low-latency online endpoint requested. Option D is wrong because converting a PyTorch model to TensorFlow SavedModel is unnecessary and would require a framework conversion rather than serving the trained PyTorch artifact.

Exam trap

Google Cloud often tests the misconception that converting to TensorFlow SavedModel is required for Vertex AI, but the platform supports PyTorch natively via custom containers, making conversion an unnecessary and potentially error-prone step.

316
MCQhard

A large e-commerce company deploys a recommendation model on Vertex AI with autoscaling enabled. During Black Friday, traffic spikes rapidly. The autoscaler adds new instances, but new instances take several minutes to become ready (cold start). As a result, many requests time out. What should they do to mitigate this issue?

A.Use a larger machine type to reduce the number of instances needed.
B.Configure the autoscaler to use CPU utilization metric instead of request count.
C.Increase the health check grace period for new instances.
D.Set a higher minimum number of instances to handle the expected peak.
AnswerD

Cold starts mean newly added instances cannot serve traffic for several minutes, so autoscaling alone cannot absorb Black Friday spikes. Raising the minimum instance count keeps enough warm capacity already running to handle the expected peak without waiting for scale-out.

Why this answer

Setting a higher minimum number of instances ensures that a baseline capacity is always running and ready to serve traffic. This pre-warms instances, eliminating the cold-start latency during rapid traffic spikes, such as Black Friday, because new instances do not need to initialize from scratch.

Exam trap

The trap here is that candidates confuse scaling metrics or instance readiness with the fundamental need for pre-provisioned capacity, leading them to choose options that adjust autoscaling behavior without eliminating the cold-start latency.

How to eliminate wrong answers

Option A is wrong because using a larger machine type reduces the number of instances needed but does not address the cold-start delay; each new instance still takes minutes to become ready. Option B is wrong because switching to CPU utilization metric does not solve the cold-start problem; the autoscaler still adds instances that take time to initialize, and CPU utilization may not react as quickly to a sudden traffic surge as request count. Option C is wrong because increasing the health check grace period only delays when the load balancer considers an instance healthy, but the instance still takes the same time to become ready; requests will still time out during the cold-start window.

317
Multi-Selectmedium

A data scientist needs to scale a prototype deep learning model to train on a massive dataset using multiple GPUs. Which three strategies are essential for efficient distributed training? (Select THREE)

Select 3 answers
A.Use a single large batch size across all workers.
B.Implement data parallelism.
C.Ensure that the input pipeline is not a bottleneck by using tf.data.Dataset with prefetching and parallel reads.
D.Use synchronous gradient updates.
E.Use asynchronous gradient updates to reduce communication overhead.
AnswersB, C, D

Data parallelism replicates the model across GPUs and splits each batch between them, satisfying the massive-dataset constraint by scaling throughput linearly with device count. Gradients are then aggregated, letting all GPUs train the same model concurrently.

Why this answer

Option B is correct because data parallelism is the foundational strategy for multi-GPU training: the model is replicated on each GPU and each replica processes a different shard of the massive dataset, with gradients aggregated across workers to scale throughput. Option C is correct because at multi-GPU scale the input pipeline can easily starve the accelerators; using tf.data.Dataset with parallel reads (num_parallel_reads / interleave) and prefetching (prefetch(AUTOTUNE)) overlaps data loading and preprocessing with GPU computation so the GPUs are not idle. Option D is correct because synchronous gradient updates (e.g., all-reduce via NCCL, as in tf.distribute.MirroredStrategy) aggregate gradients from all workers before applying the optimizer step, which keeps replicas consistent and yields stable, reproducible convergence for deep learning training.

Option A is not correct because a single large batch size is not itself an essential distributed-training strategy; batch size is a tuning choice, and naively enlarging it can hurt convergence and may not fit in memory. Option E is not correct because asynchronous gradient updates are an alternative to, not a requirement for, efficient distributed training, and they can introduce stale gradients that degrade model accuracy and reproducibility.

Exam trap

PMLE often tests the trade-off between synchronous and asynchronous updates, and candidates may incorrectly choose asynchronous as 'more efficient' when synchronous is preferred for convergence.

318
MCQeasy

A machine learning engineer wants to use Vertex AI Vizier to tune three hyperparameters: learning rate (log scale), number of layers (integer), and optimizer (categorical). They have 50 parallel trials available. Which parameter specification types should they define?

A.learning_rate: CATEGORICAL, layers: INTEGER, optimizer: CATEGORICAL
B.learning_rate: DOUBLE (unit_log_scale), layers: INTEGER (unit_linear_scale), optimizer: CATEGORICAL
C.learning_rate: DOUBLE (unit_log_scale), layers: DOUBLE (unit_linear_scale), optimizer: DISCRETE
D.learning_rate: DOUBLE (unit_linear_scale), layers: INTEGER (unit_linear_scale), optimizer: CATEGORICAL
AnswerB

Correct types and scales for the parameters.

Why this answer

Vertex AI Vizier requires parameter specifications that match the nature of each hyperparameter. Learning rate is best explored on a logarithmic scale, so it should be a DOUBLE parameter with unit_log_scale. Number of layers is a discrete integer count, so it should be an INTEGER parameter with unit_linear_scale.

Optimizer is a categorical choice among named algorithms, so it should be CATEGORICAL. This combination correctly reflects the mathematical and structural properties of each hyperparameter.

Exam trap

PMLE often tests the confusion between DISCRETE and CATEGORICAL parameter types, and between linear and log scales, causing candidates to pick specifications that do not match the hyperparameter's mathematical nature.

How to eliminate wrong answers

Option A is wrong because treating learning rate as CATEGORICAL discards the continuous, ordered nature of the value and prevents Vizier from exploring intermediate values on a log scale. Option C is wrong because number of layers should be an INTEGER, not a DOUBLE, since layers are discrete counts, and optimizer should be CATEGORICAL, not DISCRETE, because DISCRETE is for ordered numeric values rather than named categories. Option D is wrong because learning rate on a unit_linear_scale is inefficient; log scale is standard for learning rates that span orders of magnitude.

319
MCQmedium

An ML engineer needs to deploy a model to an endpoint and gradually shift traffic from the previous version (champion) to a new version (challenger) for A/B testing. How should they configure the endpoint?

A.Use a canary deployment with Cloud Run
B.Manually update the endpoint to point to the challenger after testing
C.Create a new endpoint for the challenger and route traffic via load balancer
D.Deploy both versions to the same endpoint and set traffic splitting
AnswerD

Deploying both versions to one endpoint with traffic splitting lets the engineer route a defined percentage to the challenger while the champion serves the remainder, enabling gradual A/B comparison. Separate endpoints would not permit proportional traffic distribution between versions.

Why this answer

To gradually shift traffic between a champion and challenger model for A/B testing, the engineer should deploy both versions to the same endpoint and configure traffic splitting. This allows the endpoint to route a percentage of requests to each version, enabling controlled experimentation and rollback. This is the standard pattern for A/B testing on managed ML platforms.

Exam trap

PMLE often tests deployment strategies, and candidates may confuse infrastructure-level canary deployments (e.g., Cloud Run) with model-level traffic splitting, or they may choose manual switching, missing the need for gradual, controlled experimentation.

How to eliminate wrong answers

Option A is wrong because Cloud Run is a container platform, not an ML endpoint service, and canary deployment there does not provide model-level traffic splitting. Option B is wrong because manually updating the endpoint to point to the challenger after testing does not allow gradual traffic shift or A/B testing; it is an all-or-nothing switch. Option C is wrong because creating a new endpoint and routing via load balancer adds complexity and does not use the native traffic splitting feature of the ML platform.

320
Multi-Selecteasy

A company is deploying a machine learning model for real-time inference on Vertex AI. Which TWO practices improve serving performance and reliability?

Select 2 answers
A.Use batch prediction for all requests.
B.Enable autoscaling to handle traffic variations.
C.Use manual scaling with a fixed number of replicas.
D.Deploy all models on the same machine type for consistency.
E.Set up model monitoring for prediction drift and data quality.
AnswersB, E

Autoscaling adjusts resources dynamically.

Why this answer

Vertex AI's autoscaling dynamically adjusts the number of replicas based on incoming request traffic, ensuring low latency during spikes and cost savings during lulls. This is critical for real-time inference, where consistent response times are required and manual scaling would either over-provision or under-provision resources. Autoscaling uses metrics like CPU utilization or request count to scale up or down, directly improving serving performance and reliability.

Exam trap

Google Cloud often tests the distinction between batch and real-time serving, trapping candidates who think batch prediction can be used for low-latency inference, or who assume that manual scaling is more reliable than autoscaling for variable workloads.

321
MCQeasy

You have trained a scikit-learn model and saved it as a joblib file in Cloud Storage. You need to deploy this model to Vertex AI for online predictions with minimal effort. What should you do?

A.Upload the joblib file to Vertex AI Model Registry and deploy it directly without specifying a container.
B.Convert the scikit-learn model to TensorFlow SavedModel format and use the pre-built TensorFlow container.
C.Write a custom container that loads the joblib file and serves predictions using Flask, then deploy it to Vertex AI.
D.Use the pre-built scikit-learn container provided by Vertex AI and specify the model artifact path in Cloud Storage.
AnswerD

Vertex AI provides pre-built containers for popular frameworks like scikit-learn. You can deploy the model by specifying the pre-built container image and the path to the joblib file in Cloud Storage. This requires no custom code and is the fastest way to deploy a scikit-learn model for online predictions.

Why this answer

Vertex AI offers pre-built containers for scikit-learn that can load joblib files directly. By using the pre-built container and specifying the model artifact path, you can deploy the model with minimal effort. Writing a custom container or converting the model format adds unnecessary work, and Model Registry does not support containerless deployment.

Exam trap

The trap here is assuming that you need to write a custom container or convert the model format to deploy a scikit-learn model on Vertex AI.

322
Multi-Selecteasy

Which TWO are best practices for deploying models to Vertex AI Prediction? (Choose 2.)

Select 2 answers
A.Monitor prediction latency and error rates with Cloud Monitoring alerts.
B.Log all raw prediction inputs and outputs for every request for auditing.
C.Use a dedicated service account with minimal permissions for the endpoint.
D.Always deploy the model in the same environment as training to avoid incompatibility.
E.Use the default model version alias 'default' for all deployments to simplify updates.
AnswersA, C

Cloud Monitoring alerts on prediction latency and error rates surface degradation before users are broadly affected, enabling timely rollback or scaling. This observability practise is essential for maintaining reliable Vertex AI Prediction deployments in production.

Why this answer

Option A is correct because Vertex AI Prediction exposes built-in metrics such as prediction latency, request count, and error rates to Cloud Monitoring, and configuring alerting policies on these metrics is a recommended operational best practice to detect model degradation or endpoint failures early. Option C is correct because the endpoint and model deployment should run under a dedicated service account scoped with least-privilege IAM roles (e.g., only the permissions needed to read model artifacts from Cloud Storage and write logs), which limits blast radius if the endpoint is compromised. Option B is not a best practice because logging every raw prediction input and output can expose sensitive data (PII), incur high logging costs, and is not required for auditing; sampling or redaction is preferred.

Option D is incorrect because training and serving environments are typically decoupled, and Vertex AI supports deploying models across environments as long as the serving container and dependencies are compatible; forcing identical environments is not a deployment best practice. Option E is incorrect because relying solely on the 'default' alias for all deployments removes version control and makes rollback or A/B testing harder; explicit, immutable model versions or dedicated aliases should be used.

Exam trap

PMLE often tests whether candidates over-apply 'log everything' as a best practice — the trap is picking option B, which sounds like good auditing but actually violates privacy and cost best practices.

323
MCQhard

An ML team is using Population Stability Index (PSI) to monitor feature drift on a Vertex AI Endpoint. The PSI value for a feature is 0.25, which exceeds the alert threshold of 0.2. The feature has high SHAP importance. The team wants to automatically retrain the model. What is the correct end-to-end setup?

A.Vertex AI Model Monitoring alert → Cloud Scheduler → Vertex AI Training
B.Cloud Functions → Pub/Sub → Cloud Monitoring → Vertex AI Pipeline
C.Cloud Monitoring alert → Pub/Sub → Cloud Functions → Vertex AI Pipeline (with training and deployment steps)
D.Cloud Monitoring alert → Cloud Logging → Cloud Functions → Vertex AI Training
AnswerC

Correct: This is the recommended architecture for automated retraining.

Why this answer

The correct end-to-end flow is: Vertex AI Model Monitoring detects the PSI threshold breach and emits a metric to Cloud Monitoring; a Cloud Monitoring alerting policy fires and publishes to a Pub/Sub topic; a Cloud Function subscribes and triggers a Vertex AI Pipeline that retrains and redeploys the model. This chain uses the native monitoring-to-automation path on GCP.

Exam trap

PMLE often tests the correct GCP service chain for event-driven ML automation — candidates confuse Cloud Logging with Cloud Monitoring or invert the Pub/Sub/Cloud Functions order, missing that Monitoring is the alerting source and Logging is not an event trigger.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Monitoring does not directly trigger Cloud Scheduler, and Cloud Scheduler is a cron service — it cannot react to an alert event, only to time. Option B is wrong because the order is inverted: Cloud Functions should not be the entry point, and Cloud Monitoring should be the alerting layer, not a downstream step after Pub/Sub. Option D is wrong because Cloud Logging is for log storage/querying, not for triggering automation — routing through Logging adds latency and is not the designed alerting path.

324
MCQhard

A company uses Vertex AI Feature Store for feature engineering. They need to ensure point-in-time correctness to avoid data leakage during training. Which feature retrieval method should they use?

A.Use the `get_features` API without specifying a timestamp.
B.Use BigQuery to manually join features with a sliding window.
C.Use the offline store with point-in-time join using the `feature_view` with a timestamp column.
D.Use the online store to retrieve the latest feature values.
AnswerC

The offline store performs point-in-time joins, matching each training example's timestamp against feature values valid at that moment. This guarantees the model only sees historically accurate features, eliminating the temporal leakage the scenario requires avoiding during training.

Why this answer

Point-in-time correctness requires retrieving feature values as they existed at the timestamp of each training example, which is exactly what the offline store's point-in-time join does when a feature_view is configured with an event/timestamp column. Vertex AI Feature Store uses this timestamp column to perform an as-of join, preventing future data from leaking into the training row. The online store and timestamp-less get_features calls return only the latest values, which is the classic source of label leakage.

Exam trap

PMLE often tests the distinction between online (latest-value, low-latency) and offline (historical, point-in-time) feature retrieval, tricking candidates into choosing the online store because it sounds more 'real-time' and therefore more accurate.

How to eliminate wrong answers

Option A is wrong because calling get_features without a timestamp returns the most recent feature values, which leaks future information into historical training rows. Option B is wrong because manually joining in BigQuery with a sliding window is error-prone, not the Feature Store-native mechanism, and does not leverage the feature_view's point-in-time semantics. Option D is wrong because the online store is optimized for low-latency serving of the latest feature values, not for historical as-of retrieval during training.

325
MCQmedium

A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?

A.Create a new endpoint for the new version and update the client to call both endpoints.
B.Deploy the new version and set the minimum replicas to 0, then gradually increase.
C.Use Cloud Load Balancing to distribute traffic between two endpoints.
D.Deploy the new version as a separate model on the same endpoint and use the `traffic_split` parameter in the deployment request.
AnswerD

Deploying the new version as a separate model on the same Vertex AI endpoint and setting the `traffic_split` parameter to route 5% of requests to it directly satisfies the constraint of testing with a controlled fraction of live traffic before a full rollout. This mechanism uses the endpoint’s built-in traffic routing to allocate a precise percentage of inference requests to the new model version without requiring a separate endpoint or external load balancer.

Why this answer

Vertex AI endpoints support traffic splitting between multiple deployed models. By deploying the new model version to the same endpoint and setting `traffic_split` to 5% for the new version and 95% for the existing version, the endpoint automatically routes a corresponding proportion of inference requests to each model without any client-side changes.

Exam trap

The trap here is that candidates may confuse traffic splitting with scaling or load balancing, assuming that adjusting replicas or using an external load balancer is required, when Vertex AI's native `traffic_split` is the simplest and correct method for canary deployments.

How to eliminate wrong answers

Option A is wrong because creating a new endpoint and updating clients to call both endpoints introduces unnecessary complexity, latency, and risk of client misconfiguration; Vertex AI endpoints natively support traffic splitting, making this approach redundant. Option B is wrong because setting minimum replicas to 0 does not control traffic distribution; it only affects autoscaling behavior, and gradually increasing replicas does not route a specific percentage of traffic to the new version. Option C is wrong because Cloud Load Balancing operates at the network layer and cannot intelligently split traffic between two Vertex AI endpoints based on model version; it would require additional proxy logic and defeats the purpose of Vertex AI's built-in traffic management.

326
Multi-Selectmedium

An ML engineer is building a continuous training pipeline that retrains a model when new data arrives. The pipeline should also detect skew between training and serving data. Which TWO Google Cloud services should they use? (Choose two.)

Select 2 answers
A.Cloud Logging
B.Vertex AI Model Monitoring
C.Cloud Functions
D.Vertex AI Pipelines
E.Cloud Monitoring
AnswersB, D

Vertex AI Model Monitoring detects training-serving skew by comparing live prediction traffic against a baseline dataset, directly satisfying the pipeline's skew-detection requirement. It computes distribution drift metrics on features and predictions, alerting when serving data diverges from training data, so retraining triggers fire on genuine data drift rather than a fixed schedule.

Why this answer

Vertex AI Pipelines (D) is correct because it is the managed service for orchestrating and automating ML workflows, allowing the engineer to define a pipeline that triggers retraining when new data arrives. Vertex AI Model Monitoring (B) is correct because it is specifically designed to detect training-serving skew and drift by comparing the statistical distribution of incoming prediction requests against the training data baseline. Together, these two services directly satisfy both requirements: continuous training orchestration and skew detection.

Cloud Logging (A) only stores and queries log entries and does not orchestrate pipelines or compute skew. Cloud Functions (C) is a lightweight event-driven compute service that could trigger code but lacks native ML pipeline orchestration and skew detection. Cloud Monitoring (E) tracks operational metrics and alerts but does not perform training-serving skew analysis for ML models.

Exam trap

The Google PMLE exam often tests the distinction between monitoring for infrastructure health (Cloud Monitoring) versus monitoring for ML-specific data skew (Vertex AI Model Monitoring), leading candidates to confuse general observability with ML-specific drift detection.

327
MCQeasy

You have a prototype ML model that you want to scale to production on Vertex AI. The model is a Python function that performs simple data preprocessing and then calls a pre-trained scikit-learn model. You need to deploy this as a batch prediction job that runs weekly on a large dataset stored in BigQuery. What is the most efficient way to accomplish this?

A.Deploy the model to a Vertex AI Endpoint and write a script that sends all BigQuery rows as individual prediction requests, then store the results in BigQuery.
B.Create a custom container that includes the preprocessing code and the scikit-learn model, push it to Artifact Registry, and use it to run a Vertex AI Batch Prediction job with BigQuery as the input source.
C.Use a Vertex AI Pipelines job that runs a Dataflow transform to preprocess the BigQuery data, then calls a pre-built scikit-learn model for prediction and writes results back to BigQuery.
D.Export the BigQuery data to Cloud Storage as CSV files, then run a Vertex AI Batch Prediction job using a pre-built scikit-learn container that reads from Cloud Storage.
AnswerB

Vertex AI Batch Prediction supports custom containers, allowing you to package preprocessing logic and the model together. Using BigQuery as the input source is natively supported, and the job can scale to handle large datasets. This approach is efficient because it leverages Vertex AI's managed batch prediction infrastructure without needing to export data manually.

Why this answer

For batch prediction with custom preprocessing, a custom container is required because pre-built containers do not include your code. Vertex AI Batch Prediction natively supports BigQuery as both input and output, so you can avoid data movement. This method is efficient, scalable, and integrates with Vertex AI's managed infrastructure.

Exam trap

The trap here is assuming that a pre-built container can handle custom preprocessing, or that data must be exported from BigQuery, when Vertex AI Batch Prediction supports custom containers and direct BigQuery input.

328
Multi-Selecteasy

An ML team wants to monitor their recommendation model for fairness. Which TWO metrics should they track to detect potential bias? (Select TWO.)

Select 2 answers
A.Pair-wise fairness metrics such as equal opportunity difference.
B.Recall for the minority group only.
C.Overall accuracy on the test set.
D.Average prediction confidence per request.
E.Prediction distribution (e.g., top-K recommendations) across different sensitive attribute groups.
AnswersA, E

Pair-wise fairness metrics such as equal opportunity difference compare true positive rates between sensitive attribute groups, directly quantifying disparate error rates. This satisfies the requirement to detect bias by revealing whether one group receives systematically worse outcomes than another.

Why this answer

Option A is correct because pair-wise fairness metrics such as equal opportunity difference directly compare model performance (e.g., true positive rates) between groups defined by a sensitive attribute, which is the standard way to quantify disparate impact and bias. Option E is correct because examining the prediction distribution (e.g., top-K recommendations) across sensitive attribute groups reveals whether the model systematically under- or over-represents certain groups in its outputs, a key bias signal for recommender systems. Option B is not sufficient because tracking recall for only the minority group provides no comparison baseline, so it cannot by itself detect bias relative to other groups.

Option C is not appropriate because overall accuracy on the test set can mask large disparities between groups and is not a fairness metric. Option D is not appropriate because average prediction confidence per request measures model certainty, not equitable treatment or outcomes across sensitive groups.

Exam trap

Google Cloud often tests the misconception that overall accuracy or group-specific recall alone is sufficient for fairness monitoring, when in fact comparative metrics across groups are required to detect bias.

329
MCQmedium

A company deploys a classification model on Vertex AI for loan approval. After a month, they notice the precision has dropped significantly. What should they do first?

A.Retrain the model with more data
B.Increase the number of prediction nodes
C.Check for data drift using Vertex AI Model Monitoring
D.Revert to the previous model version
AnswerC

Precision degradation after deployment typically stems from the input distribution shifting away from training data. Vertex AI Model Monitoring computes drift against the training baseline, so checking it first identifies whether feature distributions have moved, directly addressing the drop in precision before retraining or tuning.

Why this answer

A sudden drop in precision indicates that the model's predictions are no longer aligning with the ground truth, which is a classic symptom of data drift. Vertex AI Model Monitoring can automatically detect drift in feature distributions or prediction output compared to a baseline, allowing you to identify the root cause before taking corrective action. Retraining or reverting without first diagnosing the drift could waste resources or mask the underlying issue.

Exam trap

Google Cloud often tests the misconception that any performance degradation should be immediately fixed by retraining or rolling back, rather than first diagnosing the cause through monitoring tools like Vertex AI Model Monitoring.

How to eliminate wrong answers

Option A is wrong because retraining with more data does not address the root cause if the data distribution has shifted; it may even reinforce the drift if the new data is also drifted. Option B is wrong because increasing prediction nodes only improves throughput and latency, not prediction quality or precision. Option D is wrong because reverting to a previous model version is a reactive rollback that does not diagnose why precision dropped; the old model may also suffer from drift if the environment has changed.

330
MCQeasy

A company needs to serve a model with strict latency requirements (<100ms). They are using Vertex AI Prediction with CPU. During testing, latency is 150ms. What should they do?

A.Enable batching to improve throughput
B.Use a smaller machine type with more replicas
C.Export the model to TensorFlow Lite
D.Switch to a GPU machine type
AnswerD

GPUs accelerate the matrix multiplications dominating neural network inference, cutting compute time well below the 100ms threshold that CPU inference currently exceeds at 150ms. For latency-bound serving, GPU machine types provide the parallel throughput needed, satisfying the stem's strict sub-100ms constraint where CPU cannot.

Why this answer

The model's latency of 150ms exceeds the 100ms requirement. Switching to a GPU machine type (Option D) is correct because GPUs are optimized for parallel computation, significantly reducing inference latency for many ML models, especially deep learning models, compared to CPUs. Vertex AI Prediction supports GPU machine types, and this change directly addresses the latency bottleneck without altering the model or its serving configuration.

Exam trap

The trap here is that candidates confuse throughput optimization (batching or scaling replicas) with latency reduction, failing to recognize that GPUs directly address compute-bound latency while CPU-based solutions cannot meet strict sub-100ms requirements for complex models.

How to eliminate wrong answers

Option A is wrong because batching improves throughput (requests per second) by grouping multiple inference requests, but it typically increases per-request latency due to queuing and processing delays, making it unsuitable for a strict sub-100ms latency requirement. Option B is wrong because using a smaller machine type with more replicas can improve throughput and availability but does not reduce per-request inference latency; smaller machines often have less compute power, potentially increasing latency. Option C is wrong because exporting the model to TensorFlow Lite is designed for edge or mobile deployment with limited resources, not for optimizing latency in a cloud-based Vertex AI Prediction serving environment; it would require significant model conversion and may not be compatible with all model architectures.

331
MCQmedium

A data scientist needs to forecast daily sales for the next 30 days using historical sales data stored in BigQuery. They want to use BigQuery ML. Which model type should they choose?

A.LINEAR_REG
B.BOOSTED_TREE_REGRESSOR
C.K_MEANS
D.ARIMA_PLUS
AnswerD

ARIMA_PLUS handles time-series forecasting natively in BigQuery ML, modelling trend, seasonality and holidays directly from historical data. It satisfies the requirement to forecast the next 30 days of daily sales without exporting data, and supports the forecast horizon through its horizon parameter.

Why this answer

ARIMA_PLUS is the correct choice because it is specifically designed for time-series forecasting, such as predicting daily sales over a future horizon. BigQuery ML's ARIMA_PLUS model automatically handles seasonality, trend, and holiday effects, making it ideal for 30-day sales forecasts from historical data.

Exam trap

The trap here is that candidates often confuse regression models (like LINEAR_REG or BOOSTED_TREE_REGRESSOR) with time-series forecasting, not realizing that standard regression assumes independent observations and cannot inherently model temporal dependencies or extrapolate beyond the training period.

How to eliminate wrong answers

Option A is wrong because LINEAR_REG is a linear regression model for predicting a continuous target from input features, but it does not inherently model time-series dependencies like autocorrelation or seasonality, making it unsuitable for forecasting sequential daily sales. Option B is wrong because BOOSTED_TREE_REGRESSOR is an ensemble tree-based model for regression tasks, but it treats each row independently and cannot capture temporal patterns or extrapolate into the future without explicit feature engineering of time lags. Option C is wrong because K_MEANS is an unsupervised clustering algorithm used to partition data into groups, not for forecasting numerical values over time.

332
MCQmedium

A team is monitoring a model and observes that the error rate (prediction failures) has increased. They have enabled request/response logging on the Vertex AI Endpoint. How can they set up a metric and alert for prediction error rate?

A.Configure Cloud Monitoring to pull error rate from Cloud Endpoints
B.Use Vertex AI Model Monitoring to monitor error rate directly
C.Create a log-based metric in Cloud Logging for error logs and set up an alert in Cloud Monitoring
D.Enable Vertex AI Pipelines to track errors
AnswerC

Request/response logging writes prediction failures to Cloud Logging, so a log-based metric counts matching error entries. Cloud Monitoring then alerts on that metric, satisfying the need to detect a rising prediction error rate from the endpoint's existing logs.

Why this answer

To monitor prediction error rate, you can create a log-based metric in Cloud Logging that counts error logs from the Vertex AI Endpoint, and then set up an alert in Cloud Monitoring based on that metric. This leverages the request/response logging already enabled.

Exam trap

PMLE often tests the integration of Cloud Logging and Cloud Monitoring for custom metrics, and candidates might incorrectly assume Vertex AI Model Monitoring can track error rates.

How to eliminate wrong answers

Option A is wrong because Cloud Endpoints is a different service for API management, not for pulling error rates from Vertex AI. Option B is wrong because Vertex AI Model Monitoring is designed for drift and skew detection, not for monitoring prediction errors directly. Option D is wrong because Vertex AI Pipelines is for orchestrating ML workflows, not for tracking errors.

333
MCQeasy

An ML engineer has a prototype that trains a TensorFlow model on a single CPU machine using Vertex AI custom training. The job now needs to train on a larger dataset and must use multiple GPUs on one machine. The training script already uses tf.distribute.MirroredStrategy. What change is required to scale the job?

A.Add a second worker pool with one replica and set the distribution strategy to MultiWorkerMirroredStrategy.
B.Enable TPU training by specifying a TPU machine type and changing the script to use TPUStrategy.
C.Set the worker pool machine type to a GPU machine and specify the number of accelerators and accelerator type.
D.Increase the boot disk size and set the training container to use the GPU-enabled base image.
AnswerC

Vertex AI custom training uses the worker pool specification to select the machine type, accelerator type, and accelerator count. Because the script already uses MirroredStrategy, it will automatically detect the local GPUs and replicate training across them. Providing a GPU machine type with the desired accelerator count is the only configuration change needed to scale this single-machine job.

Why this answer

Scaling a single-machine TensorFlow job that already uses MirroredStrategy to multiple GPUs is primarily a resource configuration task. You set the machine type to a GPU-capable machine and specify the accelerator type and count in the worker pool spec. The strategy code detects the local GPUs automatically.

Multi-worker strategies, TPU strategies, and disk or image changes are not required for this scenario.

Exam trap

The trap here is assuming that multi-GPU training always requires a multi-worker strategy, when a single machine with multiple accelerators works with MirroredStrategy.

334
MCQeasy

A company wants to monitor the cost of their Vertex AI prediction endpoint. They are charged per hour per replica and per request for GPU instances. Which approach should they use to track these costs?

A.Set up Cloud Billing budget alerts and export billing data to BigQuery for analysis
B.Use Vertex AI Model Monitoring to track cost metrics
C.Enable Cloud Monitoring dashboards for cost metrics
D.Use Vertex AI Pipelines to track cost per job
AnswerA

Cloud Billing export to BigQuery captures per-replica hourly GPU charges and per-request costs as granular line items, letting you query and attribute spend by endpoint. Budget alerts alone only notify on thresholds; BigQuery analysis satisfies the requirement to track both billing dimensions.

Why this answer

Cloud Billing budget alerts plus BigQuery billing export is the standard GCP approach for tracking and analyzing Vertex AI endpoint costs. Budget alerts notify when spend crosses thresholds, and the detailed billing export to BigQuery enables granular analysis by SKU, label, and resource — including per-replica and per-request GPU charges.

Exam trap

PMLE often tests the confusion between operational monitoring (Cloud Monitoring) and cost monitoring (Cloud Billing) — candidates pick Cloud Monitoring dashboards assuming cost metrics appear there, but billing data requires the Billing export.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Monitoring tracks data drift and model quality, not cost metrics — it has no billing visibility. Option C is wrong because Cloud Monitoring dashboards surface operational metrics (latency, CPU, request count), not billing line items; cost data lives in Cloud Billing, not Cloud Monitoring. Option D is wrong because Vertex AI Pipelines orchestrates ML workflows and tracks pipeline runs, not the ongoing hourly and per-request costs of a deployed endpoint.

335
MCQhard

A company uses Vertex AI Pipelines to train and deploy models. The pipeline has a step that runs a custom container. The step fails intermittently with a timeout error. Which approach should be taken to robustly handle this?

A.Switch to Kubeflow Pipelines
B.Set up a Cloud Composer DAG to monitor and rerun the pipeline
C.Reduce the size of the training data
D.Increase the timeout for the step in the pipeline definition
E.Use Cloud Functions to retry the step
AnswerD

Raising the step timeout directly addresses the intermittent timeout by allowing the custom container more execution time before Vertex AI kills it. This satisfies the robustness requirement when the container legitimately needs longer than the default, without altering pipeline logic.

Why this answer

Vertex AI Pipelines (built on Kubeflow Pipelines) allows you to define a `timeout` parameter for each pipeline step. Increasing this timeout directly addresses the intermittent timeout error by giving the custom container more time to complete its work, without changing the pipeline architecture or introducing external monitoring components. This is the most robust and minimal-change solution for a step that occasionally exceeds its current time limit.

Exam trap

The trap here is that candidates may over-engineer the solution by choosing external retry mechanisms (Cloud Functions, Cloud Composer) or changing the pipeline framework, when the simplest and most correct fix is to adjust the step's timeout configuration within the pipeline definition itself.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipelines is already built on Kubeflow Pipelines; switching does not solve a timeout issue and would require re-architecting the pipeline. Option B is wrong because Cloud Composer (Apache Airflow) is an external orchestrator; adding it to monitor and rerun the pipeline adds complexity and latency, and does not fix the root cause of the step timing out. Option C is wrong because reducing training data size may degrade model quality and does not address the timeout—the step might still fail if the container itself is slow for other reasons.

Option E is wrong because Cloud Functions are stateless and event-driven; they cannot directly retry a step within a Vertex AI Pipeline—retries should be configured natively in the pipeline definition using the `retry_count` or `timeout` parameters.

336
MCQeasy

A company wants to predict customer churn using a dataset with 10,000 rows and 20 features. They have no ML expertise. Which low-code solution should they use?

A.Kubeflow Pipelines
B.Custom TensorFlow model
C.BigQuery ML
D.Vertex AI AutoML Tables
AnswerD

Vertex AI AutoML Tables suits the no-expertise constraint: it automatically selects algorithms, engineers features and tunes hyperparameters on tabular data, so the team needs no ML knowledge. It handles 10,000 rows and 20 features directly, satisfying the low-code requirement without manual model development.

Why this answer

Vertex AI AutoML Tables is the correct low-code solution because it allows users with no ML expertise to train high-quality tabular models on structured data (10,000 rows, 20 features) without writing any code. It automates feature engineering, model selection, and hyperparameter tuning, and provides a simple UI to upload data and get predictions. This directly matches the requirement of a low-code, no-expertise solution for a tabular churn prediction problem.

Exam trap

Google Cloud often tests the distinction between low-code/no-code solutions (like AutoML Tables) and platforms that still require coding or infrastructure expertise (like Kubeflow or custom TensorFlow), leading candidates to pick a technically capable but overly complex option.

How to eliminate wrong answers

Option A is wrong because Kubeflow Pipelines is a platform for building and deploying ML pipelines that requires significant coding and Kubernetes expertise, making it unsuitable for users with no ML expertise. Option B is wrong because a custom TensorFlow model requires writing Python code, defining neural network architectures, and tuning hyperparameters, which demands ML expertise. Option C is wrong because BigQuery ML is a low-code option for SQL-based ML, but it requires knowledge of SQL and ML concepts (e.g., creating models with CREATE MODEL statements), and it is less automated than AutoML Tables for users with zero ML background.

337
MCQmedium

A team is using Vertex AI Experiments to compare different hyperparameters. They want to automatically record the hyperparameters. What is the correct way?

A.Manually log to console
B.Use the `aiplatform.start_run()` context manager
C.Write to a CSV file
D.Use BigQuery
AnswerB

This context manager automatically logs hyperparameters and metrics to Vertex AI Experiments.

Why this answer

Vertex AI Experiments provides a native `aiplatform.start_run()` context manager that automatically captures hyperparameters passed as key-value arguments, logging them to the experiment run metadata without manual intervention. This integrates directly with the Vertex AI SDK, ensuring consistency and traceability across runs.

Exam trap

Google Cloud often tests the misconception that any logging method (console, CSV, BigQuery) is equivalent to native SDK integration, but the key requirement is automatic, structured recording tied to the experiment run, which only the SDK's context manager provides.

How to eliminate wrong answers

Option A is wrong because manually logging to console only outputs data to stdout, which is not persisted in Vertex AI Experiments and cannot be queried or compared programmatically. Option C is wrong because writing to a CSV file requires custom I/O code, lacks integration with Vertex AI's experiment tracking, and does not associate the hyperparameters with a specific experiment run. Option D is wrong because BigQuery is a data warehouse for analytics, not a mechanism for automatically recording hyperparameters during model training; it would require additional infrastructure to capture and store the parameters.

338
MCQhard

A team monitors features in Vertex AI Feature Store for drift. They want to set up automated alerts when a feature's distribution deviates significantly from the baseline. Which feature monitoring configuration should they use?

A.Enable feature monitoring on the feature group with drift threshold and notification channel.
B.Use Cloud Monitoring custom metrics and log-based alerts manually.
C.Use Vertex AI Experiments to compare distributions.
D.Export features to BigQuery and set up scheduled queries with alerts.
AnswerA

Feature group monitoring computes drift statistics against a baseline distribution and triggers alerts via a configured notification channel when the drift threshold is exceeded. Enabling it at the feature group level with both parameters satisfies the automated-alert requirement without custom pipeline code.

Why this answer

Feature monitoring in Vertex AI Feature Store allows defining drift thresholds and alerting via Cloud Monitoring.

339
MCQeasy

A data scientist wants to use a pre-trained ResNet model from Keras Applications and fine-tune it on a small custom dataset. Which approach should they take to avoid overfitting?

A.Freeze the first few layers and train the rest.
B.Add more convolutional layers to the model.
C.Use a larger learning rate to speed up training.
D.Train the entire model from scratch on the custom dataset.
AnswerA

Freezing early convolutional layers preserves the generic features ResNet learned on ImageNet, so only the later, task-specific layers train. Restricting updates to fewer parameters on a small dataset reduces overfitting while still adapting the model.

Why this answer

Freezing the earlier layers (which capture general features) and only training the later layers is a common transfer learning approach for small datasets, reducing overfitting.

340
MCQmedium

Your team owns a Vertex AI Model Registry entry that other teams depend on for production serving. A new retrained model shows better offline metrics, and you need to roll it out gradually to a small percentage of live traffic while keeping the ability to revert instantly if quality degrades. What should you do?

A.Create a second endpoint for the new model and update the client application to call both endpoints.
B.Register the new model under a new model resource and delete the previous version to avoid ambiguity.
C.Assign the new model version the default alias in Model Registry and redeploy the endpoint from that alias.
D.Deploy the new model version to the existing endpoint with a traffic split, then shift the split percentage as confidence grows.
AnswerD

Vertex AI Endpoints support deploying multiple model versions and splitting prediction traffic by percentage. Deploying the new version alongside the current one and starting with a small split enables a controlled canary rollout, and because the old version remains deployed, reverting is simply resetting the split back to the previous version.

Why this answer

Gradual rollout with instant rollback on Vertex AI means deploying both model versions to the same endpoint and controlling the percentage of traffic each receives. Starting with a small split limits blast radius, and because the prior version stays deployed, reverting is a split change rather than a redeployment. Alias promotion, duplicate endpoints, and deleting the old version all fail to provide controlled, reversible traffic shifting.

Exam trap

The trap here is confusing version promotion via aliases with traffic management, when only an endpoint traffic split gives gradual exposure and instant rollback.

341
Multi-Selecthard

Which THREE should be considered when setting up an automated retraining pipeline using Vertex AI Pipelines and Cloud Composer? (Choose THREE.)

Select 3 answers
A.Setting performance thresholds for new models to decide deployment
B.Including hyperparameter tuning in every retraining run
C.Optimizing resource allocation to control costs
D.Frequency of code commits to the repository
E.Monitoring for data drift to trigger retraining
AnswersA, C, E

Automated retraining produces candidate models that must be gated before deployment; defining performance thresholds lets the pipeline compare each new model against the incumbent and promote it only when metrics justify replacement, preventing silent quality regressions in production.

Why this answer

Option A is correct because an automated retraining pipeline needs defined performance thresholds (e.g., accuracy, AUC, or RMSE gates) so that a newly trained model is only promoted to deployment when it demonstrably beats or meets the required metric, preventing regression in production. Option C is correct because Vertex AI Pipelines runs training, tuning, and evaluation steps on billable compute, so resource allocation choices such as machine type, accelerator (GPU/TPU) usage, and pipeline caching directly control cost and must be planned for a sustainable retraining cadence. Option E is correct because the trigger for retraining is typically data drift or concept drift detected via Vertex AI Model Monitoring, which emits alerts that Cloud Composer DAGs can consume to launch the pipeline, making drift monitoring a core design consideration.

Option B is not required: hyperparameter tuning need not run on every retraining cycle, since it is expensive and often only needed when the model or data distribution changes materially. Option D is not relevant: the frequency of code commits reflects developer workflow, not the operational triggers or resource decisions of an automated retraining pipeline.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning must be part of every retraining run, but in practice it is a separate, infrequent optimization step to avoid excessive compute costs and pipeline latency.

342
MCQeasy

You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

A.Use Vertex AI Experiments to run an A/B test between the two versions.
B.Deploy the new version to a separate endpoint and update your client to use the new endpoint for a percentage of requests.
C.Update the endpoint to split traffic between the two model versions using the traffic split configuration.
D.Delete the old version and redeploy the new version with a different endpoint name, then update DNS.
AnswerC

Vertex AI endpoints support traffic split configuration across deployed model versions, letting you assign percentage weights to each. Adjusting these weights gradually shifts requests from old to new, achieving the controlled 24-hour rollout without redeployment.

Why this answer

Vertex AI endpoints support a built-in traffic split configuration that allows you to gradually shift traffic between model versions deployed to the same endpoint. By updating the endpoint's traffic split percentages (e.g., from 100% old / 0% new to 0% old / 100% new over 24 hours), you can achieve a smooth, controlled rollout without changing client code or managing multiple endpoints.

Exam trap

Google often tests the misconception that traffic splitting requires separate endpoints or client-side logic, when in fact Vertex AI provides a native traffic split configuration on a single endpoint.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is designed for tracking and comparing model training runs, not for managing production traffic splits or A/B testing at the serving layer. Option B is wrong because deploying to a separate endpoint and updating the client to split requests manually introduces unnecessary complexity, client-side changes, and potential inconsistency; Vertex AI's traffic split feature handles this natively at the server side. Option D is wrong because deleting the old version and redeploying with a different endpoint name, then updating DNS, would cause a complete traffic cutover (not gradual) and disrupt service during the DNS propagation period, which can take minutes to hours.

343
MCQeasy

A company deploys a model on Vertex AI Endpoints for real-time inference. They notice latency spikes during peak hours. Which action is most effective to reduce latency without sacrificing accuracy?

A.Enable autoscaling based on CPU utilization
B.Use a larger machine type
C.Reduce model size by pruning
D.Implement client-side caching
AnswerA

Autoscaling adds replica capacity when CPU utilisation rises, spreading inference requests across more nodes during peak load. This reduces per-request queueing latency while the same model and precision are served, so accuracy is unchanged. It directly addresses the peak-hour latency spikes.

Why this answer

Latency spikes during peak hours indicate the endpoint is under-provisioned for concurrent load. Enabling autoscaling based on CPU utilization (or a custom metric like request count per replica) lets Vertex AI add replicas as demand rises, absorbing the spike without changing the model or sacrificing accuracy. This is the most direct and effective action for peak-hour latency.

Exam trap

PMLE often tests whether candidates pick a static fix (bigger machine, pruning) when the scenario describes a dynamic load problem — the trap is missing that 'peak hours' implies autoscaling.

How to eliminate wrong answers

Option B is wrong because a larger machine type increases per-replica capacity but does not scale with load — during off-peak you overpay, and during extreme peaks a single larger replica can still saturate. Option C is wrong because pruning reduces model size and may slightly reduce latency, but it can degrade accuracy (contradicting 'without sacrificing accuracy') and does not address the load-driven nature of the spikes. Option D is wrong because client-side caching only helps for repeated identical requests; real-time inference with diverse inputs will not benefit, and it does not address server-side saturation.

344
MCQmedium

A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?

A.Brute-force (flat) index
B.Approximate nearest neighbor (ANN) index with Scann
C.Tree-based index
D.Hashing-based index
AnswerB

Scann's approximate nearest neighbour index trades exact search for graph-based traversal, delivering far lower query latency at high recall on large vector sets. For 10 million 512-dimensional vectors, this satisfies the lowest-latency-at-high-recall constraint better than brute-force exact search.

Why this answer

For a dataset of 10 million 512-dimensional vectors, a brute-force (flat) index would be far too slow for low-latency queries. Approximate Nearest Neighbor (ANN) with ScaNN (Scalable Nearest Neighbors) is specifically designed by Google for high-dimensional vector search, offering sub-linear query time while maintaining high recall through techniques like anisotropic quantization and tree-based partitioning. This makes it the optimal choice for balancing latency and recall at this scale.

Exam trap

The trap here is that candidates often assume brute-force is the only way to guarantee high recall, but the question explicitly asks for lowest latency at high recall, which is the exact trade-off that ANN indexes like ScaNN are designed to optimize.

How to eliminate wrong answers

Option A is wrong because a brute-force (flat) index computes exact distances to every vector, resulting in O(N) complexity per query, which is prohibitively slow for 10 million vectors and cannot achieve low latency. Option C is wrong because tree-based indexes (e.g., KD-trees, R-trees) suffer from the curse of dimensionality in high-dimensional spaces (512-D), where their performance degrades to near brute-force due to the sparsity of data. Option D is wrong because hashing-based indexes (e.g., LSH) typically require multiple hash tables to achieve high recall, leading to high memory usage and often lower recall compared to optimized ANN methods like ScaNN, especially for 512-dimensional vectors.

345
MCQhard

A company runs a Vertex AI Pipeline that includes a custom component for hyperparameter tuning. The component uses a large search space and runs many trials. The pipeline is taking too long to complete, and the team wants to reduce the execution time without sacrificing model quality. They have already optimized the training code. Which Vertex AI Pipelines feature should they use to speed up the tuning component?

A.Enable pipeline caching for the tuning component so that repeated runs with identical parameters are skipped.
B.Use the ParallelFor loop in Vertex AI Pipelines to run multiple trials concurrently as separate tasks.
C.Reduce the number of trials by using a random search instead of a grid search.
D.Increase the machine type of the tuning component to a higher CPU or GPU configuration.
AnswerB

The ParallelFor loop allows you to execute multiple tasks in parallel, which is ideal for hyperparameter tuning where trials are independent. By running trials concurrently, the overall tuning time is reduced. This approach leverages Vertex AI Pipelines' orchestration to manage parallel execution and resource allocation, and it integrates with the pipeline's tracking and caching mechanisms.

Why this answer

ParallelFor in Vertex AI Pipelines enables concurrent execution of independent tasks, which is well-suited for hyperparameter tuning trials. By running multiple trials in parallel, the overall tuning time is significantly reduced while maintaining the same search space and model quality. This leverages the pipeline's orchestration capabilities and is more effective than caching or increasing machine type.

Exam trap

The trap here is assuming that caching or larger machines solve the runtime problem, when the real bottleneck is the sequential execution of independent trials.

346
Multi-Selectmedium

A team is using Vertex AI Model Registry to manage models. They need to ensure that when a new model version is registered, it is automatically evaluated for fairness and bias before being deployed. Which two Google Cloud services should they integrate to achieve this? (Choose two.)

Select 2 answers
A.Vertex AI Pipelines
B.Vertex AI Feature Store
C.Vertex AI Model Evaluation
D.Cloud Build
E.Vertex AI Model Monitoring
AnswersA, C

Vertex AI Pipelines can orchestrate a workflow that triggers upon model registration, running evaluation steps such as fairness and bias checks. It allows integration with other services and custom code, enabling automated pre-deployment assessment. By using pipelines, the team can enforce that no model is deployed without passing the fairness evaluation.

Why this answer

Vertex AI Pipelines can orchestrate an automated workflow triggered by model registration, and Vertex AI Model Evaluation provides the fairness and bias metrics. Together, they enable pre-deployment evaluation. Other services like Model Monitoring and Feature Store focus on different aspects and cannot perform fairness assessment at registration time.

Exam trap

The trap here is confusing model monitoring with model evaluation; monitoring detects drift in deployed models, while evaluation computes fairness metrics for a model version.

347
MCQmedium

An engineer is using TensorFlow Transform (tf.Transform) to preprocess training data. They want to ensure that the same preprocessing logic is applied during inference without code duplication. Which approach should they take?

A.Use tf.Transform at prediction time by running a separate Beam pipeline
B.Use Dataflow to preprocess data for both training and serving
C.Use tf.Transform to generate a transform_fn and save it as a SavedModel; then use tf.saved_model.load to apply it in the serving pipeline
D.Write separate preprocessing code for training and serving in Python
AnswerC

Saving the transform_fn as a SavedModel exports the full preprocessing graph, including constants such as vocabulary and mean values computed during training. Loading it with tf.saved_model.load in the serving pipeline applies identical logic at inference, satisfying the no-duplication constraint without reimplementing feature engineering.

Why this answer

TensorFlow Transform outputs a SavedModel that contains the preprocessing graph. This can be exported as a transform_fn and embedded in the serving model, ensuring consistency between training and serving.

348
MCQeasy

A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?

A.Store the dataset as a Dataset artifact and pass the artifact between components.
B.Write the dataset to a temporary BigQuery table and pass the table name.
C.Serialize the dataset to a string and pass it as a pipeline parameter.
D.Use a Cloud Storage bucket and pass the bucket name as a parameter.
AnswerA

Dataset artifacts are references to Cloud Storage or BigQuery URIs, so components exchange a lightweight metadata pointer rather than the bytes themselves. This satisfies the stem's large-dataset constraint, since passing raw data through component inputs would exceed the metadata limits.

Why this answer

In Vertex AI Pipelines, the recommended way to pass large datasets between components is to use a `Dataset` artifact. Artifacts are metadata references that point to the underlying data stored in Cloud Storage, enabling efficient, scalable, and type-safe data passing without serialization overhead or size limits. This approach leverages the Kubeflow Pipelines SDK's artifact tracking, which automatically handles lineage and versioning.

Exam trap

The trap here is that candidates often assume passing a Cloud Storage bucket name (Option D) is sufficient, but they miss that artifacts provide automatic metadata tracking, type safety, and integration with Vertex AI's lineage system, which is required for production ML pipelines.

How to eliminate wrong answers

Option B is wrong because writing a large dataset to a temporary BigQuery table introduces unnecessary latency, cost, and complexity; BigQuery is designed for analytical queries, not as an intermediate data transfer mechanism for pipeline components. Option C is wrong because serializing a large dataset to a string and passing it as a pipeline parameter violates the parameter size limit (typically 64KB in Kubeflow Pipelines) and would cause out-of-memory errors or pipeline failures. Option D is wrong because passing only the bucket name as a parameter lacks the structured metadata and type safety that artifacts provide; it forces components to independently resolve file paths and does not automatically track lineage or versioning.

349
MCQmedium

An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?

A.Use a canary deployment by deploying to a separate endpoint and using a load balancer with weighted routing.
B.Deploy the new model to a new endpoint, then update DNS to point to the new endpoint.
C.Delete the old model and deploy the new one with the same endpoint.
D.Deploy the new model to the same endpoint with 0% traffic initially, then gradually increase traffic while monitoring.
AnswerD

Deploying the new model to the same endpoint at 0% traffic creates a second deployed model behind the shared endpoint, letting you shift traffic incrementally via the endpoint's traffic split. This satisfies the no-downtime constraint: the existing version keeps serving while you monitor error rates before promoting the new one.

Why this answer

Vertex AI supports deploying multiple models to the same endpoint and splitting traffic via a deployed model's traffic percentage. The correct canary procedure is to deploy the new model version to the existing endpoint with 0% traffic, then gradually increase its traffic share while monitoring error rates and latency. This avoids downtime and enables safe rollback by shifting traffic back to the old version.

Exam trap

The trap here is confusing Vertex AI's native traffic-splitting capability with generic load-balancer-based canary patterns — candidates who don't know Vertex AI endpoints support multiple models with weighted traffic pick the external load balancer answer.

How to eliminate wrong answers

Option A is wrong because Vertex AI endpoints natively support traffic splitting — introducing an external load balancer and separate endpoint adds unnecessary complexity and is not the standard procedure. Option B is wrong because DNS-based cutover is a blue/green pattern with no gradual traffic control and DNS propagation delays cause unpredictable behavior. Option C is wrong because deleting the old model before the new one is validated risks downtime and eliminates rollback capability.

350
MCQmedium

A team wants to implement CI/CD for their ML models using Cloud Build. They have a pipeline that trains a model and deploys it. What is the best practice for triggering the pipeline when a new commit is pushed to the source repository?

A.Set up a Cloud Scheduler job to poll the repository periodically
B.Deploy a custom web service on App Engine to call Cloud Build API
C.Use Pub/Sub to notify Cloud Build of new commits
D.Configure a Cloud Build trigger on the source repository (e.g., Cloud Source Repositories, GitHub)
AnswerD

A Cloud Build trigger bound to the repository fires automatically on push events, invoking the build that trains and deploys the model. This satisfies the new-commit constraint by replacing manual invocation with event-driven execution tied directly to source changes.

Why this answer

Cloud Build natively supports triggers that automatically start a pipeline when a new commit is pushed to a connected source repository (e.g., Cloud Source Repositories, GitHub, Bitbucket). This is the simplest, most event-driven approach, requiring no polling, custom services, or additional messaging infrastructure. It directly maps the git push event to a build invocation, ensuring near-instantaneous pipeline execution.

Exam trap

The trap here is that candidates may overthink the solution and choose Pub/Sub (Option C) because they know Pub/Sub is used for event-driven architectures, but they miss that Cloud Build triggers already abstract this complexity away, making direct trigger configuration the best practice.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler polling is inefficient, introduces latency (minimum 1-minute intervals), and is not event-driven; it would waste resources and delay pipeline starts. Option B is wrong because deploying a custom web service on App Engine to call the Cloud Build API adds unnecessary complexity, cost, and maintenance overhead, and is not a best practice when native triggers exist. Option C is wrong because while Pub/Sub can be used to trigger builds, it requires an intermediary (e.g., a Cloud Function) to receive the commit notification and call the Cloud Build API, adding latency and complexity compared to a direct Cloud Build trigger.

351
MCQhard

A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?

A.Create a custom service account with required permissions and assign it to the Dataflow job.
B.Grant the 'roles/storage.objectViewer' role to 'allUsers' on the Cloud Storage bucket.
C.Use the Composer environment's service account for all pipeline components.
D.Move the Dataflow job to run after the pipeline so that data is already processed.
AnswerA

Dataflow workers execute under a service account, so granting the required Cloud Storage permissions to a dedicated custom service account and assigning it to the job fixes the access failures directly. This is more efficient than broadly widening project-level roles.

Why this answer

The most efficient way to resolve insufficient permissions for Dataflow workers accessing Cloud Storage is to create a custom service account with the required roles (e.g., roles/storage.objectViewer) and assign it to the Dataflow job via the --serviceAccount option. This follows the principle of least privilege and ensures that only the Dataflow workers have the necessary permissions, without affecting other pipeline components or exposing the bucket publicly.

Exam trap

Google Cloud often tests the misconception that using a single service account for all components (like the Composer environment's service account) is simpler and sufficient, but this ignores the principle of least privilege and can cause security vulnerabilities or permission conflicts in distributed pipelines.

How to eliminate wrong answers

Option B is wrong because granting roles/storage.objectViewer to 'allUsers' makes the Cloud Storage bucket publicly readable, which is a severe security risk and violates least privilege principles. Option C is wrong because the Composer environment's service account typically has broader permissions than needed for Dataflow workers, and using it for all components can lead to over-privileging and potential security issues; moreover, Dataflow workers require a separate identity to access resources independently. Option D is wrong because moving the Dataflow job to run after the pipeline does not address the root cause of insufficient permissions; the Dataflow job will still fail when it tries to access Cloud Storage data, regardless of when it runs.

352
MCQmedium

You need to serve a large embedding model for similarity search with low latency. The model was trained to generate 256-dimensional embeddings. You plan to use Vertex AI Vector Search. Which index type should you choose to balance accuracy and performance for a dataset with 10 million vectors?

A.Tree-based index
B.Approximate nearest neighbor (ANN) index using ScaNN
C.Brute-force index
D.Hash-based index
AnswerB

ScaNN builds an approximate nearest neighbour index, trading exact recall for substantially lower query latency and memory at 10 million vectors. This satisfies the stem's balance of accuracy and performance, where exact search would be too slow at that scale.

Why this answer

Vertex AI Vector Search uses ScaNN (Scalable Nearest Neighbors) as its underlying ANN algorithm, which is specifically designed for high-dimensional embeddings (like 256-d) and large-scale datasets (10M vectors). ScaNN balances accuracy and performance by employing anisotropic quantization and tree-based partitioning, making it the optimal choice for low-latency similarity search without requiring exhaustive comparison.

Exam trap

Candidates often mistakenly choose Tree-based index (Option A) because ScaNN uses tree-based partitioning internally, but a standalone tree index fails in high dimensions. Vertex AI Vector Search’s ScaNN combines tree partitioning with quantization to overcome the curse of dimensionality and balance accuracy and performance for 10 million 256-d vectors.

How to eliminate wrong answers

Option A is wrong because a pure tree-based index (e.g., KD-tree) suffers from the 'curse of dimensionality' at 256 dimensions, where performance degrades to near brute-force levels. Option C is wrong because a brute-force index computes exact distances for all 10M vectors, resulting in O(n) latency that is unacceptable for real-time serving. Option D is wrong because hash-based indexes (e.g., LSH) are typically used for approximate nearest neighbor search in lower dimensions or for specific distance metrics, but they are not natively supported as a primary index type in Vertex AI Vector Search, and they often require extensive tuning to match ScaNN's accuracy-latency trade-off.

353
MCQmedium

What is the most likely cause of the error?

A.The data split column contains only NULL values, so no rows are assigned to the training set
B.The model type 'linear_reg' is incompatible with the column 'price' because of missing values
C.The model creation does not have permission to read the dataset in BigQuery
D.The model creation did not specify a training budget, so default is insufficient
AnswerA

AutoML splits data using the designated split column; if every value is NULL, no row is assigned to the training, validation, or test sets, so training cannot proceed and the job fails with this error.

Why this answer

When the data split column contains only NULL values, BigQuery ML cannot assign any rows to the training set. The `DATA_SPLIT_METHOD` using a custom column requires non-NULL values in that column to partition data into training and evaluation sets; if all values are NULL, the training set receives zero rows, causing the model creation to fail with an error about insufficient training data.

Exam trap

Google Cloud often tests the subtle distinction between missing values in the label column (which are handled gracefully) versus missing values in the data split column (which can cause a complete failure), leading candidates to incorrectly blame missing values in the target column.

How to eliminate wrong answers

Option B is wrong because the `linear_reg` model type is fully compatible with the `price` column even if it has missing values; BigQuery ML handles NULLs in the label column by excluding those rows during training, but the error here is about no training rows, not missing values. Option C is wrong because if the user lacked permission to read the dataset, the error would be a permissions-related message (e.g., 'Access Denied'), not a training set size error. Option D is wrong because BigQuery ML does not require a training budget for linear regression models; the default settings are sufficient, and the error is not budget-related.

354
MCQeasy

You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?

A.Increase the machine type to a larger instance with more memory and vCPUs.
B.Configure autoscaling on the endpoint with a minimum replica count of one and a maximum replica count that covers peak load.
C.Deploy the model to a batch prediction job and schedule it to run every hour.
D.Enable request logging and set up an alert when latency exceeds a threshold.
AnswerB

Vertex AI endpoint autoscaling adjusts replicas based on utilization or request concurrency. Setting a low minimum keeps cost down during normal traffic, while a maximum that covers peak load allows the endpoint to scale out during bursts. This matches the requirement to handle tenfold spikes automatically without over-provisioning all day, because replicas are added only when needed and removed when traffic subsides.

Why this answer

Vertex AI endpoints support autoscaling based on metrics such as CPU utilization or request concurrency. By setting a low minimum replica count, you keep baseline cost low during normal periods. A maximum replica count sized for peak load allows the endpoint to scale out automatically when traffic increases tenfold.

This elastic behavior handles short bursts without manual intervention and avoids paying for peak capacity all day, which is exactly what the scenario requires.

Exam trap

The trap here is thinking that a larger machine type provides elasticity, when autoscaling adds replicas horizontally and is what actually handles burst traffic.

355
MCQmedium

You are deploying a scikit-learn model for online predictions. The model size is 200 MB. You want to minimize latency and cost. Which serving option should you choose?

A.Deploy to Vertex AI online prediction using a prebuilt container for scikit-learn.
B.Use Cloud Run with a custom container.
C.Create a Kubernetes cluster on GKE and deploy the model there.
D.Export the model as a Cloud Function.
AnswerA

A prebuilt scikit-learn container on Vertex AI online prediction loads the 200 MB model and serves low-latency requests without custom infrastructure, minimising both latency and cost. Batch prediction adds delay, and larger custom deployments raise cost, so this option satisfies both constraints.

Why this answer

Vertex AI online prediction with a prebuilt scikit-learn container is purpose-built for serving ML models with managed autoscaling, low-latency endpoints, and no infrastructure management. It handles model loading, versioning, traffic splitting, and monitoring out of the box, so a 200 MB scikit-learn model can be deployed with minimal effort and cost. This directly satisfies the requirement to minimize both latency and operational cost.

Exam trap

The trap here is assuming that any container-based serverless option (Cloud Run) is automatically cheaper and lower-latency than a managed ML service, when in fact Vertex AI's prebuilt containers remove the custom-serving-code overhead and provide ML-specific features like model versioning and traffic splitting.

How to eliminate wrong answers

Option B is wrong because Cloud Run requires you to build and maintain a custom container with a web server (e.g., Flask/FastAPI) wrapping the model, adding cold-start latency and engineering overhead not needed for a standard scikit-learn model. Option C is wrong because standing up a GKE cluster for a single 200 MB model introduces significant cost and operational burden (node pools, autoscaling, upgrades) that Vertex AI manages for you. Option D is wrong because Cloud Functions has deployment package size limits (well under 200 MB uncompressed in most configurations) and is not designed for hosting large ML model artifacts or GPU-backed inference.

356
MCQeasy

A team uses Vertex AI Feature Store for storing features. They want to share feature definitions with other teams in a collaborative manner. What is the best way to collaborate on feature definitions?

A.Use a shared repository with feature definition files and CI/CD to update the feature store.
B.Grant all teams write access to the same feature store so they can modify definitions directly.
C.Export the feature definitions as CSV and email them to the other teams.
D.Use a wiki page to document feature definitions and update it manually.
AnswerA

Feature definitions stored as version-controlled files in a shared repository let multiple teams review, reuse and propose changes collaboratively. CI/CD then applies those definitions to Vertex AI Feature Store consistently, satisfying the requirement for collaborative sharing rather than ad-hoc manual updates.

Why this answer

Using a shared repository with feature definition files and CI/CD pipelines enables version control, peer review, and automated deployment to Vertex AI Feature Store. This approach ensures consistency, traceability, and collaboration without risking direct, uncoordinated changes to the production feature store.

Exam trap

The trap here is that candidates may assume direct write access (Option B) is efficient for collaboration, but the exam tests understanding that feature stores require controlled, versioned updates to maintain data integrity and avoid breaking downstream models.

How to eliminate wrong answers

Option B is wrong because granting all teams write access to the same feature store allows uncoordinated, direct modifications to feature definitions, which can lead to conflicts, data corruption, and lack of version control. Option C is wrong because exporting feature definitions as CSV and emailing them is error-prone, lacks versioning, and does not provide a single source of truth for collaboration. Option D is wrong because using a wiki page for manual documentation is static, easily outdated, and does not integrate with the feature store's actual schema or deployment process.

357
Multi-Selectmedium

You are deploying a model to a Vertex AI Endpoint and need to reduce inference latency for a real-time application. Which two actions should you take? (Choose two.)

Select 2 answers
A.Enable request-response logging to monitor latency metrics.
B.Optimize the model for inference using techniques such as quantization or pruning.
C.Choose a machine type with a GPU accelerator that matches the model's compute profile.
D.Enable autoscaling with a high maximum replica count to handle traffic spikes.
E.Set the endpoint's minReplicaCount to zero to save cost during idle periods.
AnswersB, C

Quantization reduces numerical precision, shrinking model size and memory bandwidth needs, which often lowers latency. Pruning removes redundant weights, reducing computation. Both techniques can speed up inference on suitable hardware with minimal accuracy loss, making them effective for latency-sensitive real-time serving when applied carefully.

Why this answer

Latency reduction comes from making each prediction faster, not from adding capacity. Using a GPU that fits the model's compute profile accelerates the math, and inference-specific optimizations like quantization or pruning reduce the work per prediction. Autoscaling, zero minimum replicas, and logging affect scalability, availability, or observability rather than per-request latency.

Exam trap

The trap here is equating scalability features like autoscaling with latency improvements, when they primarily affect throughput and availability.

358
MCQmedium

You are training a scikit-learn random forest model on a large tabular dataset using a Vertex AI custom training job. The training script reads a CSV file from a Cloud Storage bucket and writes the trained model artifact to the same bucket. You need to ensure the training job can access the Cloud Storage bucket without embedding long-lived credentials in the container. What should you do?

A.Grant the Vertex AI Service Agent role to the default Compute Engine service account and rely on the default credentials provided by the training environment to access the bucket.
B.Configure the training job to use a user-managed service account that has the Editor role on the project, ensuring it has sufficient permissions to read and write to any Cloud Storage bucket.
C.Attach a custom service account to the Vertex AI training job that has the Storage Object Viewer and Storage Object Creator roles on the bucket, and let the Vertex AI Training service use the service account's credentials automatically.
D.Create a service account key file, store it in Secret Manager, and mount it into the training container as a volume so the scikit-learn script can read the key and authenticate to Cloud Storage.
AnswerC

Vertex AI training jobs run as a service account you specify. By attaching a custom service account with least-privilege Cloud Storage roles, the training container automatically obtains credentials via the metadata server. This avoids embedding keys and follows Google-recommended security practices for accessing Cloud Storage from training jobs.

Why this answer

The correct approach is to attach a custom service account with least-privilege Cloud Storage roles to the Vertex AI training job. Vertex AI training jobs automatically use the credentials of the attached service account via the metadata server, so no key files are needed. This ensures secure, short-lived authentication and follows Google Cloud best practices for access control.

Exam trap

The trap here is assuming that a service account key file or default credentials are required, when Vertex AI training jobs natively support attaching a custom service account for automatic authentication.

359
MCQmedium

A team deploys a PyTorch model on Vertex AI for online predictions. They notice that after deployment, the latency increases over time, especially during peak hours. The model is served using a custom container. What is the most likely cause?

A.The custom container does not have a health check, causing instances to be prematurely terminated.
B.The model is not using GPU even though a GPU machine is selected.
C.The model is too large for the machine's memory, causing swapping.
D.The prediction requests are not being batched, and the model inference code is not optimized for concurrency.
AnswerD

Without request batching, each call triggers a separate forward pass, and unoptimised inference code serialises concurrent requests. Under peak load this queues requests, so latency climbs as the container saturates rather than scaling with traffic.

Why this answer

The latency increase over time, especially during peak hours, indicates that the model inference code is not handling concurrent requests efficiently. Without batching or optimized concurrency, each request is processed sequentially, causing a queue buildup under load. This is a common issue with custom containers on Vertex AI when the prediction handler is single-threaded or lacks async processing.

Exam trap

Google Cloud often tests the misconception that latency increases are always due to resource exhaustion (memory/CPU) rather than concurrency or request handling inefficiencies, leading candidates to pick Option C.

How to eliminate wrong answers

Option A is wrong because a missing health check would cause instances to be terminated and recreated, leading to intermittent failures or startup latency, not a gradual latency increase over time. Option B is wrong because selecting a GPU machine without using the GPU would result in underutilization but not necessarily increasing latency; the model would still run on CPU, and latency would be constant or high from the start. Option C is wrong because if the model were too large for memory, swapping would cause consistently high latency from the outset, not a gradual increase during peak hours.

360
MCQeasy

A data scientist wants to track the lineage of a dataset used in a training run. Which Vertex AI feature should they use?

A.Vertex ML Metadata
B.Vertex AI Feature Store
C.Vertex AI Experiments
D.Vertex AI Model Registry
AnswerA

Vertex ML Metadata records artefacts, executions and contexts, automatically capturing dataset inputs and training runs as lineage graphs. This directly satisfies the stem's requirement to track which dataset fed a training run, letting the scientist query provenance rather than reconstruct it manually from logs.

Why this answer

Vertex ML Metadata is the correct choice because it is specifically designed to track the lineage of datasets, models, and other artifacts throughout the ML lifecycle. It records metadata about each step in a pipeline, including the source dataset used for a training run, enabling full provenance tracking. This allows data scientists to trace back which data was used, how it was transformed, and which model version it produced.

Exam trap

The trap here is that candidates often confuse Vertex AI Experiments (which tracks run metrics and parameters) with lineage tracking, but Experiments does not capture the full artifact-to-execution graph that ML Metadata provides for dataset provenance.

How to eliminate wrong answers

Option B is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature values for ML models, not for tracking dataset lineage. Option C is wrong because Vertex AI Experiments is used to track and compare model training runs, hyperparameters, and metrics, but it does not natively capture the lineage of the dataset itself beyond run-level parameters. Option D is wrong because Vertex AI Model Registry is a version control system for trained models, managing model deployments and versions, but it does not track the provenance of the training data used to create those models.

361
MCQhard

A company has multiple teams working on different models. They want to enforce consistent data preprocessing steps across all teams. Which approach should they take?

A.Use Cloud Composer to orchestrate preprocessing
B.Write shared Python packages in Artifact Registry
C.Use Cloud Dataflow templates
D.Create shared Vertex AI Pipelines components
AnswerD

Shared Vertex AI Pipelines components package preprocessing logic as reusable, versioned artefacts that every team imports into its own pipeline. This enforces identical preprocessing steps across teams, satisfying the consistency constraint without duplicating code per model.

Why this answer

Vertex AI Pipelines components allow teams to define reusable, versioned, and parameterized preprocessing steps that can be shared across models and pipelines. This ensures consistent execution of data transformations because each component encapsulates the exact code and environment, and pipelines enforce the same DAG of steps regardless of which team triggers them.

Exam trap

Google Cloud often tests the distinction between 'sharing code' (e.g., packages) and 'sharing executable, environment-encapsulated pipeline steps' (e.g., components), leading candidates to choose a code-sharing option like Artifact Registry instead of the pipeline component approach that enforces consistency.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is an orchestration service for workflows (based on Apache Airflow) and does not inherently enforce consistent preprocessing logic across teams; it only schedules and monitors tasks, leaving the actual preprocessing code to be defined separately and potentially inconsistently. Option B is wrong because writing shared Python packages in Artifact Registry provides a way to distribute code, but it does not enforce a standardized execution environment or pipeline structure; teams could still call the packages with different parameters or in different orders, leading to inconsistency. Option C is wrong because Cloud Dataflow templates are used for batch and stream data processing jobs (based on Apache Beam), but they are not designed to be shared as reusable, composable steps across multiple ML pipelines; they lack the pipeline-level DAG enforcement and versioning that Vertex AI Pipelines components provide.

362
MCQhard

A company has deployed a model for image classification and wants to monitor for feature drift using XRAI attributions. However, they notice that the XRAI attribution maps are too large and are causing high latency in the monitoring pipeline. What is the most effective way to reduce the overhead of explainability monitoring for image models?

A.Disable XRAI and use integrated gradients instead
B.Reduce the sampling rate for the explainability feature
C.Use a smaller image size for the model
D.Increase the number of replicas on the endpoint
AnswerB

XRAI attribution maps are generated per prediction, so reducing the sampling rate lowers how many explanations are computed, cutting pipeline latency. This satisfies the overhead constraint while still providing statistically useful drift signal from the sampled subset.

Why this answer

Vertex AI Explainable AI supports XRAI for image models, but generating XRAI attributions can be computationally expensive. Sampling a subset of predictions reduces the number of explanations generated, lowering latency and cost.

363
MCQeasy

An organization wants to use Cloud Composer (Airflow) to orchestrate a machine learning workflow that includes running a Vertex AI Pipeline, followed by a BigQuery job, and then a Dataflow pipeline. What is the primary advantage of using Cloud Composer for this orchestration?

A.It allows orchestrating heterogeneous workflows across multiple GCP services with dependencies and retries.
B.It automatically caches the outputs of each step to avoid recomputation.
C.It integrates natively with the Vertex AI Model Registry for model versioning.
D.It provides a serverless execution environment for ML pipelines.
AnswerA

Cloud Composer runs Apache Airflow DAGs, whose operators and sensors coordinate tasks across disparate GCP services with explicit dependencies and retry policies. This satisfies the stem's need to sequence a Vertex AI Pipeline, BigQuery job and Dataflow pipeline in one workflow.

Why this answer

Cloud Composer (Apache Airflow) is designed to orchestrate heterogeneous workflows across multiple GCP services. In this scenario, it can define a Directed Acyclic Graph (DAG) that runs a Vertex AI Pipeline, then a BigQuery job, and finally a Dataflow pipeline, with built-in support for dependency management, retries, and failure handling. This is the primary advantage because it allows you to coordinate disparate services in a single, reliable workflow.

Exam trap

The trap here is that candidates may confuse Cloud Composer's orchestration capabilities with features specific to individual GCP services (like caching, model registry, or serverless execution), leading them to pick options that describe those services' features rather than the primary advantage of using an orchestrator.

How to eliminate wrong answers

Option B is wrong because Cloud Composer does not automatically cache outputs of each step; caching is a feature of specific services like Vertex AI Pipelines or Dataflow, not Airflow itself. Option C is wrong because Cloud Composer does not natively integrate with the Vertex AI Model Registry; that integration is handled by Vertex AI Pipelines or custom operators, not by Airflow's core orchestration. Option D is wrong because Cloud Composer is not serverless; it runs on a managed GKE cluster, and serverless ML pipeline execution is provided by Vertex AI Pipelines, not Cloud Composer.

364
MCQhard

A financial institution wants to use Natural Language API for sentiment analysis on customer feedback, but the domain-specific language (e.g., 'bullish', 'bearish') is not correctly classified. They have 200 labeled examples. Which approach minimizes coding effort while improving accuracy?

A.Submit a feature request to Google for domain-specific terms
B.Create a custom sentiment dictionary and pass it to the Natural Language API
C.Build a custom TensorFlow model for sentiment
D.Use AutoML Natural Language to train a custom model
AnswerD

AutoML Natural Language trains a custom sentiment model on the 200 labelled examples through the console, learning domain terms such as 'bullish' and 'bearish' that the pretrained Natural Language API misclassifies. This meets the minimal-coding constraint while adapting to the financial vocabulary.

Why this answer

AutoML Natural Language enables you to train a custom model on your 200 labeled examples without writing code, directly improving accuracy for domain-specific terms like 'bullish' and 'bearish'. This approach leverages transfer learning from Google's pre-trained models, minimizing coding effort while adapting to your unique vocabulary and sentiment patterns.

Exam trap

Google Cloud often tests the misconception that the Natural Language API supports custom dictionaries or rule-based overrides, when in fact it only offers a fixed pre-trained model, making AutoML the correct low-code path for domain adaptation.

How to eliminate wrong answers

Option A is wrong because submitting a feature request to Google for domain-specific terms is not a practical solution—Google does not provide custom term updates for individual customers, and the turnaround time is indefinite. Option B is wrong because the Natural Language API does not accept a custom sentiment dictionary; it only supports a static, built-in sentiment model, and passing a dictionary is not a supported feature. Option C is wrong because building a custom TensorFlow model requires significant coding effort, including data preprocessing, model architecture design, training, and deployment, which contradicts the goal of minimizing coding effort.

365
MCQhard

A machine learning engineer is deploying a TensorFlow model on an edge device with limited memory and compute. The model needs to perform inference with low latency. The engineer has a trained float32 model. Which model compression technique should be applied first to reduce the model size and improve inference speed without significant accuracy loss?

A.Post-training quantization to INT8
B.Knowledge distillation
C.Quantization-aware training
D.Weight pruning
AnswerA

Post-training quantization converts the trained float32 weights and activations to INT8, cutting model size roughly fourfold and accelerating inference on edge hardware with limited memory and compute. It requires no retraining and typically preserves accuracy, making it the appropriate first compression step.

Why this answer

Post-training quantization to INT8 is the correct first step because it directly reduces the model size by approximately 4x (from 32-bit floats to 8-bit integers) and speeds up inference on edge devices by leveraging integer-optimized hardware (e.g., ARM NEON or Qualcomm Hexagon). This technique requires no retraining and typically yields minimal accuracy loss for most TensorFlow models, making it the fastest path to deploy on resource-constrained devices.

Exam trap

Google Cloud often tests the misconception that quantization-aware training is always required for INT8 deployment, but the trap here is that post-training quantization is the simplest and most effective first step for reducing model size and latency on edge devices, with quantization-aware training reserved only for cases where accuracy drops below acceptable thresholds.

How to eliminate wrong answers

Option B (Knowledge distillation) is wrong because it requires training a smaller student model from scratch using the teacher model's outputs, which is computationally expensive and not a quick compression technique for an already trained model. Option C (Quantization-aware training) is wrong because it simulates quantization effects during training to preserve accuracy, but it requires retraining the model, making it a second step after post-training quantization if accuracy loss is unacceptable. Option D (Weight pruning) is wrong because it removes individual weights (often via magnitude-based pruning), which can reduce model size but typically requires retraining to recover accuracy and does not directly improve inference speed on standard edge hardware without sparse matrix support.

366
MCQmedium

You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?

A.Deploy the challenger to a separate endpoint and use a load balancer to route a percentage of requests.
B.Use Cloud Armor to route traffic based on headers.
C.Deploy both models to the same endpoint and use the traffic split feature to allocate percentages.
D.Create a new endpoint for the challenger and gradually shift DNS records.
AnswerC

Deploying both models to one Vertex AI endpoint lets the traffic split feature assign a percentage to each deployed model ID, shifting requests gradually from champion to challenger. This directly satisfies the gradual traffic-shift requirement while both versions serve live predictions, so business metrics can be compared under real traffic.

Why this answer

Vertex AI endpoints natively support traffic splitting between multiple deployed models on the same endpoint, allowing you to assign a percentage of prediction traffic to the champion and challenger. This is the built-in, supported mechanism for gradual rollouts and A/B experiments, and it integrates with Vertex AI's monitoring so you can measure business metrics per model version.

Exam trap

The trap here is assuming you need external infrastructure (load balancers, DNS, Cloud Armor) to split traffic, when Vertex AI endpoints already provide native traffic splitting as a first-class feature.

How to eliminate wrong answers

Option A is wrong because deploying to separate endpoints and adding an external load balancer bypasses Vertex AI's native traffic split and requires custom routing logic that Vertex AI already provides. Option B is wrong because Cloud Armor is a WAF/DDoS protection service operating at the edge for HTTP(S) load balancing, not a model traffic-routing mechanism for Vertex AI endpoints. Option D is wrong because DNS-based shifting is coarse, slow to propagate (TTL-dependent), and cannot provide the fine-grained percentage control or per-model metrics that Vertex AI traffic split offers.

367
Multi-Selecthard

Your team has deployed a model on Vertex AI endpoints and you are planning an A/B test to compare a new challenger model (v2) against the current champion (v1). The test should measure business metrics such as click-through rate. Which THREE steps should you take to set up the A/B test correctly? (Choose 3 correct answers)

Select 3 answers
A.Deploy the challenger model (v2) to the same endpoint as the champion (v1).
B.Modify your application to log which model version served each prediction.
C.Create a new endpoint for v2 and gradually shift DNS traffic.
D.Use Vertex AI Experiments to compare model performance.
E.Set up a traffic split between v1 and v2, e.g., 90% v1 and 10% v2.
AnswersA, B, E

Deploying both models to the same endpoint is required for A/B testing, because Vertex AI traffic splitting operates across model versions co-deployed on one endpoint. This lets requests be routed proportionally to v1 and v2 for click-through-rate comparison.

Why this answer

Option A is correct because Vertex AI endpoints natively support deploying multiple models to the same endpoint, which is the prerequisite for serving both champion v1 and challenger v2 behind a single prediction URL. Option E is correct because once both models share an endpoint, you configure a traffic split (e.g., 90% to v1 and 10% to v2) via the deployed model's traffic percentage, which is exactly how Vertex AI routes online prediction requests between model versions for an A/B test. Option B is correct because to measure business metrics such as click-through rate per variant, the application must log which model version (or deployed model ID) served each prediction so that downstream clicks can be attributed to v1 or v2.

Option C is not appropriate because shifting DNS traffic is a coarse, infrastructure-level approach that does not use Vertex AI's built-in traffic splitting and cannot reliably target a percentage of prediction requests. Option D is not appropriate because Vertex AI Experiments is designed for tracking and comparing training/evaluation runs and metrics, not for routing live prediction traffic between deployed model versions.

Exam trap

The trap here is that candidates confuse Vertex AI Experiments (for training) with endpoint traffic splitting (for serving), and they incorrectly think creating separate endpoints with DNS shifting is a valid A/B testing method, when Vertex AI's native traffic splitting is the correct and simpler approach.

368
MCQhard

You need to create a reproducible snapshot of a BigQuery table as of a specific timestamp for ML model training. The snapshot should be queryable without copying the entire dataset. Which BigQuery feature should you use?

A.BigQuery export to Cloud Storage as Parquet
B.BigQuery time travel (FOR SYSTEM_TIME AS OF)
C.CREATE TABLE AS SELECT with WHERE clause
D.BigQuery table snapshots
AnswerD

BigQuery table snapshots capture a table's contents at a specified timestamp as a lightweight, queryable reference that shares storage with the base table rather than duplicating data. This satisfies the stem's reproducibility and no-full-copy constraints.

Why this answer

BigQuery table snapshots create a lightweight, queryable copy of a table at a point in time using copy-on-write storage — only changed data consumes additional bytes, and the snapshot is a first-class table you can query directly. This satisfies the requirement for a reproducible, timestamped, queryable snapshot without duplicating the full dataset. Time travel, by contrast, is a query-time window (default 7 days) and does not persist as a durable object.

Exam trap

The trap here is conflating time travel (a transient query window) with table snapshots (a durable, queryable object) — candidates often pick FOR SYSTEM_TIME AS OF because it sounds like a snapshot, but it does not persist beyond the time-travel window.

How to eliminate wrong answers

Option A is wrong because exporting to Cloud Storage as Parquet creates a physical copy of the data outside BigQuery, which is neither queryable in-place nor storage-efficient. Option B is wrong because FOR SYSTEM_TIME AS OF is a query modifier limited to the time-travel window (default 7 days, max 7 days unless configured), and it does not create a persistent snapshot object. Option C is wrong because CREATE TABLE AS SELECT physically materializes a full copy of the data, incurring full storage cost and defeating the 'without copying the entire dataset' requirement.

369
MCQeasy

A data science team deploys a regression model to predict house prices. After one month, the mean absolute error (MAE) on the serving data increases by 20% compared to the test set. Which monitoring strategy should the team implement first to diagnose the issue?

A.Retrain the model daily with the latest data to adapt to changing patterns.
B.Monitor prediction residuals and compute serving-time MAE over sliding windows.
C.Compare the distribution of training labels with serving labels using a two-sample t-test.
D.Monitor input feature distributions for drift using the Kolmogorov-Smirnov test.
AnswerB

Monitoring residuals and computing serving-time MAE over sliding windows directly quantifies the 20% degradation against the test-set baseline, exposing whether error drifts gradually or spikes. Residual distributions also reveal bias or variance shifts in predictions, pinpointing covariate or concept drift as the root cause before invoking retraining or feature audits.

Why this answer

The first step in diagnosing a 20% MAE increase on serving data is to monitor prediction residuals over sliding windows. This directly tracks how model errors evolve in production, allowing the team to detect whether performance degradation is sudden or gradual, and to correlate it with specific time windows or data slices. Computing serving-time MAE on sliding windows provides an immediate, interpretable signal of model health without assuming the root cause.

Exam trap

Google Cloud often tests the misconception that the first step in diagnosing model degradation is to check for data drift (Option D), when in fact the correct first step is to confirm and quantify the performance drop itself using serving-time metrics like sliding-window MAE.

How to eliminate wrong answers

Option A is wrong because retraining daily without first diagnosing the cause of the MAE increase is a reactive, resource-intensive approach that may mask underlying issues like data drift or concept drift, and does not help identify whether retraining is even necessary. Option C is wrong because comparing training labels with serving labels using a two-sample t-test checks for label distribution shift, but the MAE increase could be due to feature drift, concept drift, or data quality issues unrelated to label distribution; this test is too narrow and may miss the actual cause. Option D is wrong because monitoring input feature distributions for drift using the Kolmogorov-Smirnov test is a valid technique, but it is a secondary diagnostic step; the first priority should be to confirm and characterize the performance degradation itself via residual monitoring before investigating potential causes.

370
MCQmedium

You are setting up feature monitoring in Vertex AI Feature Store to detect drift in a numerical feature. The monitoring job should run daily and alert if the Jensen-Shannon divergence exceeds 0.1. Which configuration should you use?

A.Configure feature monitoring in the feature view with a drift threshold of 0.1 using Jensen-Shannon divergence
B.Use BigQuery scheduled queries to compare distributions and send alerts
C.Set up a Cloud Composer DAG to compute drift and publish to Cloud Monitoring
D.Enable model monitoring on the Vertex AI endpoint to detect drift
AnswerA

Configuring monitoring on the feature view applies drift detection directly to the served feature data, satisfying the daily Jensen-Shannon divergence threshold of 0.1. Feature-view-level monitoring evaluates the numerical feature against its baseline distribution, so alerts trigger precisely when divergence exceeds the specified limit.

Why this answer

Vertex AI Feature Store supports native feature monitoring configured at the feature view level, where you specify the drift detection method (including Jensen-Shannon divergence) and a threshold. Setting the threshold to 0.1 with a daily schedule directly satisfies the requirement without building custom infrastructure. This is the first-party, managed approach for detecting feature drift.

Exam trap

The trap is assuming model monitoring on the endpoint covers feature drift — PMLE candidates often pick endpoint monitoring when the question is specifically about Feature Store feature-level drift detection.

How to eliminate wrong answers

Option B is wrong because BigQuery scheduled queries require you to implement drift math manually and do not integrate with Feature Store's native monitoring or alerting. Option C is wrong because a Cloud Composer DAG is a custom orchestration workaround that duplicates functionality already provided by Feature Store monitoring. Option D is wrong because model monitoring on a Vertex AI endpoint detects prediction drift and skew at serving time, not feature-level drift in the Feature Store itself.

371
MCQmedium

A data scientist uses Vertex AI Workbench to train a model and then deploys it to an endpoint. They want to automate the retraining and redeployment pipeline when new data arrives. Which service should they use?

A.Cloud Composer
B.Vertex AI Pipelines
C.Cloud Scheduler
D.Cloud Functions
AnswerB

Vertex AI Pipelines orchestrates the retraining and redeployment workflow as a repeatable DAG, triggered when new data arrives. It satisfies the automation constraint by chaining data ingestion, training, evaluation and endpoint deployment steps without manual intervention.

Why this answer

Vertex AI Pipelines is the managed orchestration service for ML workflows on Google Cloud, designed to automate training, evaluation, and deployment steps as a reproducible DAG. It integrates natively with Vertex AI endpoints and supports triggers from new data events, making it the correct choice for automating retraining and redeployment. Cloud Composer can orchestrate, but Vertex AI Pipelines is purpose-built for ML and requires less overhead for this use case.

Exam trap

PMLE often tests whether candidates pick a general orchestrator (Cloud Composer) when the ML-native, lower-overhead answer (Vertex AI Pipelines) is intended, or confuse a trigger (Cloud Scheduler) with an orchestrator.

How to eliminate wrong answers

Option A is wrong because Cloud Composer (managed Airflow) is a general-purpose workflow orchestrator — it can call Vertex AI, but it is not the ML-native pipeline service and adds operational overhead for this specific ML automation need. Option C is wrong because Cloud Scheduler is a cron-based job trigger, not a pipeline orchestrator; it can kick off a pipeline but cannot manage the multi-step training/deployment DAG. Option D is wrong because Cloud Functions is an event-driven serverless compute service for lightweight tasks, not a pipeline orchestrator capable of managing multi-step ML workflows with artifact lineage.

372
MCQmedium

A company needs to run batch predictions on 10 TB of data stored in Cloud Storage. The predictions should be written to BigQuery. Which approach should they use?

A.Export the model to Cloud Functions and trigger on file upload
B.Create a Vertex AI Batch Prediction job with GCS input and BigQuery output
C.Use Vertex AI Online Prediction with a batch job
D.Use Dataflow to read from GCS and write to BigQuery, calling the model for each record
AnswerB

Vertex AI Batch Prediction natively reads GCS input and writes results directly to BigQuery, satisfying both the 10 TB Cloud Storage source and the BigQuery destination without intermediate exports. This managed service handles large-scale batch inference, so no custom pipeline or data movement code is required.

Why this answer

Vertex AI Batch Prediction natively supports reading input from Cloud Storage and writing predictions directly to BigQuery, making it the most efficient and fully managed solution for large-scale batch inference on 10 TB of data. This approach avoids the complexity of custom infrastructure or per-record model calls, leveraging Vertex AI's optimized batch processing pipeline.

Exam trap

This question tests the distinction between batch and online prediction modes in Vertex AI. The trap is that candidates may confuse Vertex AI's batch prediction with using Dataflow or Cloud Functions, not realizing that Vertex AI natively supports BigQuery as a direct output destination for batch jobs.

How to eliminate wrong answers

Option A is wrong because Cloud Functions are designed for event-driven, lightweight processing and cannot handle 10 TB of data efficiently; exporting a model to Cloud Functions also lacks native batch prediction orchestration and BigQuery output support. Option C is wrong because Vertex AI Online Prediction is intended for real-time, low-latency inference on individual requests, not for batch jobs; there is no 'batch job' mode within online prediction. Option D is wrong because while Dataflow can read from GCS and write to BigQuery, calling the model for each record would require custom code and per-record inference, which is less efficient and more complex than using Vertex AI's built-in batch prediction with direct BigQuery output.

373
MCQhard

A media company wants to build a real-time recommendation system for articles. They have a large user base (10M+) and frequent updates to user interactions. They need to handle cold-start users and new articles. Which architecture on Vertex AI is most suitable?

A.Deploy a Deep Learning Recommendation Model (DLRM) for prediction
B.Use a contextual bandit algorithm for exploration only
C.Use matrix factorization with collaborative filtering
D.Implement a two-tower model (user and item towers) with embeddings and nearest neighbor search
AnswerD

Two-tower models can incorporate side features and enable fast retrieval.

Why this answer

The two-tower model (user and item towers) with embeddings and nearest neighbor search is the most suitable because it handles cold-start users and new articles by learning separate embeddings for users and items, enabling efficient retrieval via approximate nearest neighbor (ANN) search. This architecture supports real-time updates and scales to 10M+ users by decoupling user and item representations, allowing incremental training on new interactions without full retraining.

Exam trap

Google Cloud often tests the misconception that matrix factorization (Option C) is sufficient for cold-start scenarios, but candidates miss that it requires retraining on new data and cannot generate embeddings for unseen users or items without side features.

How to eliminate wrong answers

Option A is wrong because DLRM is a deep learning model for click-through rate prediction that requires retraining on new data and does not natively handle cold-start items or users without additional feature engineering, making it less suitable for frequent updates and real-time recommendation. Option B is wrong because a contextual bandit algorithm for exploration only lacks exploitation of known user preferences, leading to suboptimal recommendations over time, and does not provide a full recommendation system. Option C is wrong because matrix factorization with collaborative filtering cannot handle cold-start users or new articles without retraining the entire model, as it relies on existing interaction matrices and lacks a mechanism for incorporating new entities in real time.

374
MCQeasy

What does the `ML.PREDICT` command do in BigQuery ML?

A.Trains a new BigQuery ML model
B.Exports the model to Cloud Storage
C.Evaluates the model's performance
D.Makes predictions using the model
AnswerD

ML.PREDICT runs a trained BigQuery ML model against a input table and returns predicted values alongside the source rows. It satisfies the scenario's need to score new data, unlike CREATE MODEL, which trains, or ML.EVALUATE, which only reports metrics.

Why this answer

The command is likely a BigQuery ML prediction query (e.g., using `ML.PREDICT`) that uses a trained model to generate predictions on new input data, making option D correct. It does not train, export, or evaluate the model.

Exam trap

Google Cloud often tests the distinction between the four key BigQuery ML commands (`CREATE MODEL`, `ML.EVALUATE`, `ML.PREDICT`, `EXPORT MODEL`), and the trap here is confusing the prediction function with the evaluation function, especially when the exhibit shows a query that looks like it might be evaluating performance due to the presence of a model name and input data.

How to eliminate wrong answers

Option A is wrong because training a new BigQuery ML model uses the `CREATE MODEL` statement, not the `ML.PREDICT` function. Option B is wrong because exporting a model to Cloud Storage uses the `EXPORT MODEL` statement, not a prediction query. Option C is wrong because evaluating model performance uses the `ML.EVALUATE` function, which returns metrics like loss and accuracy, not predictions.

375
MCQmedium

You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?

A.Convert the model to TFLite and deploy on an edge device.
B.Build a custom container with NVIDIA Triton Inference Server and deploy it to Vertex AI.
C.Export the model to ONNX and deploy using Vertex AI's built-in TensorFlow serving.
D.Use Vertex AI Model Optimisation to automatically quantise the model.
AnswerB

Building a custom container lets you bundle Triton with the required custom op libraries, then deploy it as a Vertex AI custom container model. This satisfies the stem's constraint: Triton is not a natively supported Vertex AI serving framework, and the custom ops demand dependencies the prebuilt PyTorch containers lack.

Why this answer

Vertex AI supports custom containers for prediction, so you can package NVIDIA Triton Inference Server with your PyTorch model and any custom ops/libraries it needs, then deploy that container as a Vertex AI Model. Triton natively supports PyTorch (via TorchScript), ONNX, TensorRT, and custom backends, making it the right choice when the model relies on non-standard operators that Vertex AI's pre-built PyTorch/TensorFlow containers cannot execute.

Exam trap

The trap here is assuming that exporting to ONNX or using Vertex AI Model Optimisation will magically make unsupported custom ops work; the exam tests whether you know that custom containers are the escape hatch for non-standard runtimes and operators.

How to eliminate wrong answers

Option A is wrong because TFLite targets mobile/edge deployment and cannot run arbitrary PyTorch custom ops, and it does not address online inference on Vertex AI. Option C is wrong because exporting to ONNX does not guarantee the custom ops are supported by TensorFlow Serving, and Vertex AI's built-in TF serving container cannot execute PyTorch custom operators. Option D is wrong because Vertex AI Model Optimisation performs quantization/pruning on supported model formats but does not add support for unsupported custom ops or change the serving runtime.

Page 4

Page 5 of 11

Page 6

All pages