Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 526–600

775 questions total · 11pages · All types, answers revealed

Page 7

Page 8 of 11

Page 9
526
MCQmedium

A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?

A.Cloud Bigtable
B.Cloud CDN
C.Cloud Memorystore
D.Cloud SQL
AnswerC

Cloud Memorystore offers a managed in-memory cache that the serving path can query before invoking the model, returning stored responses for repeated identical requests. This reduces endpoint invocations, lowering both latency and cost as the stem requires.

Why this answer

Cloud Memorystore provides a managed Redis or Memcached in-memory cache that can store frequent identical prediction requests and their responses, dramatically reducing latency and backend load. It is the standard Google Cloud service for application-level caching in front of Vertex AI endpoints.

Exam trap

PMLE often tests the distinction between caching layers — candidates may pick Cloud CDN for 'caching' without realizing CDN only caches HTTP responses at the edge and cannot key on arbitrary prediction payloads.

How to eliminate wrong answers

Option A is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput analytics and time-series data, not for low-latency key-value caching of prediction responses. Option B is wrong because Cloud CDN caches HTTP content at edge locations for static assets, not dynamic prediction API responses keyed by request payload. Option D is wrong because Cloud SQL is a relational database with millisecond-to-tens-of-milliseconds latency, far slower than in-memory caching and not designed for high-QPS cache workloads.

527
MCQeasy

An ML engineer is creating a Vertex AI Pipeline that includes a component to train a model. The component requires a machine type with a GPU. The engineer wants to specify the machine type and GPU type for the component. Which approach should the engineer use?

A.Configure the GPU in the pipeline's runtime configuration, which applies to all components.
B.Use a pre-built training container from Vertex AI that automatically selects the appropriate GPU.
C.Set the machine type and accelerator in the component's container specification using the Vertex AI SDK.
D.Pass the machine type and GPU type as command-line arguments to the component's container.
AnswerC

In Vertex AI Pipelines, you can specify the machine type and accelerator (GPU) in the component's container specification when defining the component. This is done using the Vertex AI SDK's component decorator or by setting the resource requirements in the pipeline task. This ensures that the component runs on the desired hardware.

Why this answer

To specify a machine type and GPU for a Vertex AI Pipelines component, you set the resource requirements in the component's container specification. This ensures the component runs on the desired hardware. Other approaches like command-line arguments or global configuration do not correctly allocate the resources for the specific component.

Exam trap

The trap here is thinking that machine type and GPU can be passed as runtime arguments or set globally, but they must be declared in the component's resource specification.

528
MCQmedium

You are deploying a pre-trained BERT model for inference on edge devices. The model must be under 500 MB and inference latency under 50 ms. Which approach should you take?

A.Use a larger model like BERT-Large and deploy on GPU
B.Apply post-training INT8 quantization using TensorFlow Lite
C.Prune 50% of the model weights and fine-tune
D.Use knowledge distillation to train a smaller student model from scratch
AnswerB

Post-training INT8 quantization shrinks the BERT weights roughly fourfold, bringing the model comfortably under the 500 MB ceiling, while integer arithmetic accelerates inference to meet the 50 ms latency budget. TensorFlow Lite's runtime is purpose-built for edge deployment, so no retraining is required.

Why this answer

Post-training INT8 quantization reduces model size by approximately 75% (from ~440 MB to ~110 MB for BERT-Base) and accelerates inference on edge devices via integer arithmetic, easily meeting the 500 MB and 50 ms constraints. TensorFlow Lite provides hardware-optimized kernels for ARM CPUs and NPUs, making it ideal for edge deployment without requiring retraining.

Exam trap

A common pitfall is assuming that only pruning or distillation can reduce model size, but post-training INT8 quantization directly shrinks the model and speeds inference without retraining, which is ideal for deploying a pre-trained model on edge devices.

How to eliminate wrong answers

Option A is wrong because BERT-Large is ~1.3 GB, far exceeding the 500 MB limit, and GPU deployment is not feasible on most edge devices due to power and thermal constraints. Option C is wrong because pruning 50% of weights without fine-tuning would cause catastrophic accuracy loss, and fine-tuning requires the original training data and compute, which may not be available; even with fine-tuning, pruned models often need specialized hardware for speedup. Option D is wrong because knowledge distillation requires training a smaller student model from scratch, which demands significant compute, time, and access to the teacher model's logits, making it impractical for a quick deployment scenario where a pre-trained BERT model is already available.

529
Multi-Selectmedium

You are using tf.Transform to preprocess data at scale. Which TWO services are required to run tf.Transform on Google Cloud? (Choose 2)

Select 2 answers
A.Cloud Functions
B.Dataflow
C.Cloud Storage
D.Vertex AI Training
E.BigQuery
AnswersB, C

Dataflow is required to execute the Apache Beam pipeline that tf.Transform generates, running the preprocessing at scale. It satisfies the stem's requirement by providing the distributed runner that applies the transform over large datasets, complementing the service that analyses and materialises the transform artefacts.

Why this answer

tf.Transform requires Apache Beam for execution, which on GCP is typically run on Dataflow. The processed data and transform artifacts are stored in Cloud Storage.

530
MCQhard

A company has a CI/CD pipeline that retrains a model every time new training data is available. They want to automatically deploy the new model to production only if it passes a set of evaluation tests on a staging environment. Which approach best implements this?

A.Implement a two-stage pipeline: train and deploy to staging, run evaluation tests, and if passed, deploy to production using conditional logic.
B.Use Cloud Build to trigger a training job and then a separate deployment job without evaluation.
C.Use a single Vertex AI pipeline that trains and deploys to staging, then manually promote.
D.Train and deploy directly to production in one pipeline.
AnswerA

A two-stage pipeline trains and deploys to staging, runs evaluation tests, then uses conditional logic to promote to production only on passing results. This gates deployment on measured quality, satisfying the requirement to release solely when evaluation succeeds.

Why this answer

The correct approach is a two-stage pipeline: train and deploy to a staging environment, run automated evaluation tests against the staged model, and use conditional logic to promote to production only if the tests pass. This implements a proper ML CI/CD gate that prevents regressions from reaching production.

Exam trap

PMLE often tests whether candidates confuse 'automated deployment' with 'automated promotion' — the trap is choosing an option that deploys automatically but omits the evaluation gate, or one that evaluates but requires manual promotion.

How to eliminate wrong answers

Option B is wrong because a deployment job without evaluation skips the quality gate entirely, allowing untested models into production. Option C is wrong because manual promotion does not satisfy the requirement for automatic deployment based on evaluation results — it introduces a human step that the question explicitly wants to avoid. Option D is wrong because training and deploying directly to production in one pipeline bypasses staging and evaluation, which is exactly the risk the company wants to mitigate.

531
MCQhard

Your team is serving a large language model on Vertex AI using a custom container. The endpoint experiences intermittent 502 errors during traffic spikes. The autoscaling configuration uses a CPU utilization target of 60% and the model is deployed on n1-standard-4 instances. The model requires significant memory. Which combination of changes is most likely to resolve the issue?

A.Increase the target CPU utilization to 90% to allow more requests per instance.
B.Switch to a machine type with more memory, e.g., n1-highmem-8, and increase min_replica_count.
C.Enable canary traffic splitting to reduce load on the main endpoint.
D.Reduce the model batch size from 32 to 1 to lower memory per request.
AnswerB

Increasing memory headroom on n1-highmem-8 prevents the container from being killed or stalled when the model's working set exceeds the 15 GB available on n1-standard-4 during spikes, which is what produces the 502s. Raising min_replica_count keeps warm capacity ready so autoscaling lag cannot drop requests.

Why this answer

The 502 errors during traffic spikes indicate that instances are becoming unhealthy or crashing, most likely due to memory exhaustion. The model requires significant memory, and n1-standard-4 instances have only 15 GB of RAM, which is insufficient for a large language model. Switching to n1-highmem-8 (which provides 52 GB of RAM) directly addresses the memory bottleneck, and increasing min_replica_count ensures that enough instances are available to handle baseline traffic and absorb spikes without overloading any single instance.

This combination resolves the root cause: insufficient memory leading to instance failures under load.

Exam trap

The trap here is assuming that autoscaling alone can resolve performance issues, when the root cause is often insufficient per-instance resources; candidates may focus on scaling parameters rather than instance sizing.

How to eliminate wrong answers

Option A is wrong because increasing the CPU utilization target to 90% would allow more requests per instance, but this would exacerbate memory pressure and likely increase the frequency of 502 errors, as the instances are already memory-constrained. Option C is wrong because canary traffic splitting is a deployment strategy for testing new model versions, not a mechanism to reduce load on a production endpoint; it does not address the underlying memory shortage. Option D is wrong because reducing the batch size from 32 to 1 would lower memory per request but would drastically reduce throughput and increase latency, and it does not solve the fundamental issue of insufficient memory on the instances; it merely shifts the problem to higher request volume.

532
MCQmedium

A company deploys a model on Vertex AI Endpoint and expects high traffic spikes during promotional events. The current configuration uses manual scaling with 2 replicas. Which autoscaling configuration should they use to handle spikes while minimizing cost during normal traffic?

A.Keep manual scaling but increase replicas to 10.
B.Set min_replica_count=2 and max_replica_count=10 with no scaling metric.
C.Enable basic scaling with target_cpu_utilization=0.6 and set min_replica_count=2, max_replica_count=10.
D.Use custom metric scaling with a Cloud Monitoring metric for prediction latency.
AnswerC

Basic scaling with target_cpu_utilization=0.6 adds replicas automatically as CPU load rises during spikes, while min_replica_count=2 keeps baseline capacity and max_replica_count=10 caps spend. This satisfies the stem's need to absorb promotional spikes while minimising cost at normal traffic.

Why this answer

Enabling autoscaling with a CPU utilization target and setting min/max replica counts allows Vertex AI to scale replicas up during traffic spikes and down during normal traffic, balancing performance and cost. The min of 2 preserves baseline capacity while the max of 10 caps cost during spikes.

Exam trap

The trap is choosing 'more replicas' or 'custom metrics' as the answer, when the exam wants you to recognize that autoscaling requires both a scaling metric and min/max bounds to balance cost and performance.

How to eliminate wrong answers

Option A is wrong because static 10 replicas wastes money during normal traffic and still cannot scale beyond 10 during unexpected spikes. Option B is wrong because autoscaling requires a scaling metric — without one, Vertex AI cannot determine when to add or remove replicas. Option D is wrong because custom metric scaling on prediction latency is more complex and not necessary here; CPU utilization is the standard, cost-effective signal for this workload.

533
MCQeasy

A company has deployed a model to a Vertex AI Endpoint and wants to receive an email notification whenever Vertex AI Model Monitoring detects feature drift above a configured threshold. They have already set up the monitoring configuration with a training dataset baseline. What should they do next to enable email alerts?

A.Set the monitoring frequency to 5 minutes so that alerts are generated automatically.
B.Enable the Vertex AI Model Monitoring email notification toggle in the endpoint settings.
C.Configure a Pub/Sub topic as the monitoring sink and subscribe an email endpoint to it.
D.Create a Cloud Monitoring alerting policy on the Vertex AI Model Monitoring drift metric and add an email notification channel.
AnswerD

Vertex AI Model Monitoring publishes drift metrics to Cloud Monitoring. To receive email alerts, you create an alerting policy on the relevant drift metric and attach an email notification channel. This is the standard integration path and directly satisfies the requirement without custom code.

Why this answer

Vertex AI Model Monitoring emits drift metrics to Cloud Monitoring. To get email alerts, you must create a Cloud Monitoring alerting policy that watches the drift metric and includes an email notification channel. The other options either describe non-existent toggles, use an indirect Pub/Sub path, or confuse frequency with notification, so they do not achieve the goal.

Exam trap

The trap here is assuming that enabling monitoring automatically sends alerts, when notification channels must be configured in Cloud Monitoring.

534
MCQhard

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built AutoML Tables component to train a classification model. The pipeline also includes a conditional step that deploys the model to an endpoint only if the evaluation metrics exceed a threshold. Which KFP feature should be used to implement the conditional deployment?

A.dsl.ParallelFor
B.dsl.ExitHandler
C.dsl.Condition
D.dsl.Collected
AnswerC

dsl.Condition wraps pipeline steps in a conditional branch whose predicate is evaluated at runtime, so the deployment step executes only when the evaluation metrics exceed the threshold. This directly implements the stem's conditional deployment requirement within the KFP pipeline definition.

Why this answer

The `dsl.Condition` feature from KFP (Kubeflow Pipelines) is specifically designed to conditionally execute pipeline steps based on the output of a previous component. In this scenario, the AutoML Tables component produces evaluation metrics; `dsl.Condition` allows the pipeline to check whether those metrics exceed a threshold and, if true, run the deployment step. This is the correct, native KFP construct for implementing branching logic within a pipeline.

Exam trap

The trap here is that candidates often confuse `dsl.Condition` with `dsl.ExitHandler` because both involve decision-making, but `ExitHandler` is only for post-exit cleanup, not for branching based on step outputs.

How to eliminate wrong answers

Option A is wrong because `dsl.ParallelFor` is used for iterating over a collection of items and executing steps in parallel, not for conditional branching based on a single metric threshold. Option B is wrong because `dsl.ExitHandler` is a mechanism to run a cleanup or notification step when a pipeline exits (successfully or with failure), not for conditionally deploying a model based on evaluation results. Option D is wrong because `dsl.Collected` is a function used to gather outputs from parallel iterations (e.g., from `dsl.ParallelFor`) into a single list, not a control flow construct for conditional execution.

535
Multi-Selecteasy

A team has deployed a model on Vertex AI Prediction and wants to monitor for data drift. Which TWO metrics should they use to detect drift in numerical features?

Select 2 answers
A.Pearson correlation coefficient
B.Jensen-Shannon divergence (JSD)
C.Chi-squared statistic
D.Population Stability Index (PSI)
E.Kolmogorov-Smirnov (KS) statistic
AnswersB, E

JSD measures similarity between two probability distributions and works for numerical features after binning.

Why this answer

Jensen-Shannon divergence (JSD) is a symmetric, bounded (0 to 1) measure of the difference between two probability distributions, making it ideal for detecting drift in numerical features by comparing the training distribution to the serving distribution. It is a smoothed and normalized version of Kullback-Leibler divergence, and Vertex AI Prediction's Model Monitoring natively supports JSD for numerical feature drift detection.

Exam trap

Google Cloud often tests the misconception that Pearson correlation or Chi-squared are appropriate for numerical drift, when in fact Pearson measures correlation between two variables and Chi-squared is for categorical data, leading candidates to overlook the correct distribution-comparison metrics like JSD and KS.

536
MCQmedium

A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?

A.Set the 'caching' flag to 'True' in the pipeline definition using 'pipeline.caching = True'.
B.Set the environment variable 'ENABLE_CACHE' to 'true' on the pipeline run request.
C.Re-compile the pipeline with the '--enable-cache' flag.
D.Ensure that the component does not have 'dsl.cache_options(enable_cache=False)' set.
AnswerD

Caching is enabled by default in Kubeflow Pipelines; the component re-executes because cache_options(enable_cache=False) was explicitly set, disabling it. Removing that setting restores default caching behaviour, so unchanged inputs and code reuse the cached execution instead of rerunning.

Why this answer

Vertex AI pipeline caching is enabled by default for all components unless explicitly disabled using `dsl.cache_options(enable_cache=False)`. The component re-executing every time indicates that caching was likely disabled on that specific component. Removing or ensuring this setting is not present will allow the pipeline to reuse cached outputs when inputs and code have not changed.

Exam trap

The trap here is that candidates assume caching must be explicitly enabled (like in some other cloud platforms), but Vertex AI caches by default, so the issue is usually that caching was explicitly disabled on the component.

How to eliminate wrong answers

Option A is wrong because Vertex AI pipelines do not have a global `pipeline.caching` attribute; caching is controlled per component via the `@component` decorator or `ContainerComponent` definition, not at the pipeline level. Option B is wrong because there is no environment variable `ENABLE_CACHE` for Vertex AI pipeline runs; caching is configured in the pipeline definition, not via runtime environment variables. Option C is wrong because Vertex AI pipelines do not use a `--enable-cache` compilation flag; caching is a runtime feature controlled by component-level settings, not a compile-time option.

537
Multi-Selectmedium

A team is responsible for monitoring the health of a Vertex AI pipeline that runs daily. Which THREE resources should they use to gain visibility into pipeline performance and failures? (Choose 3.)

Select 3 answers
A.Cloud Trace for analyzing distributed execution
B.Cloud Composer for tracking DAGs
C.Vertex AI Experiments for comparing pipeline runs
D.Cloud Monitoring for metrics and alerts on pipeline runs
E.Cloud Logging for viewing pipeline step logs
AnswersC, D, E

Vertex AI Experiments records pipeline runs with their parameters, metrics and artefacts, enabling side-by-side comparison of successive daily executions. This satisfies the visibility requirement by exposing run-to-run performance changes and surfacing which pipeline iteration failed or degraded.

Why this answer

Vertex AI Experiments (C) is correct because it records and compares pipeline runs, letting the team track parameters, metrics, and artifacts across daily executions to evaluate performance trends. Cloud Monitoring (D) is correct because it collects pipeline run metrics and supports alerting policies so the team can be notified of failures or anomalies in the daily pipeline. Cloud Logging (E) is correct because pipeline step logs are written to Cloud Logging, providing detailed diagnostic visibility into individual step failures.

Cloud Trace (A) is not the right fit here since it targets distributed request latency tracing rather than Vertex AI pipeline run health, and Cloud Composer (B) is a separate managed Airflow service for orchestrating DAGs, not the native monitoring surface for a Vertex AI pipeline.

Exam trap

Google Cloud often tests the distinction between monitoring (observing run-level metrics and logs) and tracing (analyzing request-level latency), leading candidates to incorrectly select Cloud Trace for pipeline health visibility when it is actually intended for distributed request tracing.

538
MCQmedium

A data science team uses a shared Cloud Storage bucket to store training datasets. They notice that some team members accidentally overwrite existing datasets, causing issues with reproducibility. Which approach best prevents accidental overwrites while maintaining collaboration?

A.Use a single shared service account with strict IAM roles that allow only append operations.
B.Require team members to manually rename files before uploading.
C.Set bucket permissions to read-only for all team members except the data owner.
D.Enable object versioning on the bucket and use lifecycle rules to manage versions.
AnswerD

Object versioning preserves every overwrite as a noncurrent version, so prior datasets remain retrievable and reproducibility is maintained. Lifecycle rules then prune aged versions to control storage cost. This directly satisfies the stem's constraint of preventing accidental overwrites while keeping the bucket shared for collaboration.

Why this answer

Enabling object versioning on a Cloud Storage bucket preserves all versions of an object, so even if a team member overwrites a dataset, the previous version remains accessible. This maintains collaboration (anyone can upload) while preventing permanent data loss. Lifecycle rules can then be used to manage storage costs by automatically deleting old versions after a specified period.

Exam trap

The trap here is that candidates may think IAM roles or permissions are the only way to control data integrity, overlooking that object versioning provides a safety net without blocking collaboration.

How to eliminate wrong answers

Option A is wrong because Cloud Storage does not support 'append-only' IAM roles; objects are immutable and must be rewritten entirely, so this approach would not prevent overwrites and would break normal upload workflows. Option B is wrong because relying on manual renaming is error-prone and does not enforce any technical control, so accidental overwrites can still occur. Option C is wrong because making the bucket read-only for most team members prevents them from uploading new datasets at all, which destroys collaboration and is overly restrictive.

539
MCQmedium

A company wants to run batch predictions on millions of records stored in BigQuery. They need to preprocess the data (e.g., feature engineering) before feeding it to the model. Which approach is most scalable and cost-effective?

A.Use a large DataProc cluster to preprocess and run batch predictions.
B.Preprocess inline in the batch prediction job using a custom container.
C.Use a custom Python script on a Compute Engine instance.
D.Preprocess with Cloud Dataflow, output to Cloud Storage, then submit a Vertex AI batch prediction job.
AnswerD

Cloud Dataflow performs distributed preprocessing at scale, writing engineered features to Cloud Storage, which Vertex AI batch prediction reads directly. This decouples heavy transformation from the prediction job, satisfying the millions-of-records scalability and cost-efficiency constraint better than in-job preprocessing.

Why this answer

The most scalable and cost-effective because Cloud Dataflow (Apache Beam) provides serverless, auto-scaling preprocessing that handles large volumes of data efficiently, and Vertex AI batch predictions natively read from Cloud Storage, avoiding the need to manage infrastructure. This decouples preprocessing from prediction, allowing each to scale independently and minimizing costs by using ephemeral, pay-per-use resources.

Exam trap

A common mistake is assuming a single large cluster (Dataproc) or a single VM is sufficient for batch processing, when in fact serverless, auto-scaling services like Dataflow are more appropriate for large-scale, ephemeral preprocessing tasks.

How to eliminate wrong answers

Option A is wrong because Dataproc clusters require manual sizing, incur idle costs, and add operational overhead for a simple preprocessing task; it is overkill and less cost-effective than serverless options. Option B is wrong because preprocessing inline in a custom container for batch prediction tightly couples preprocessing with prediction, preventing independent scaling and making it harder to handle large-scale data transformations efficiently. Option C is wrong because a single Compute Engine instance cannot scale horizontally to process millions of records in a reasonable time, and managing failover, retries, and parallelization would require custom code, making it neither scalable nor cost-effective.

540
MCQeasy

You are responsible for maintaining an ML pipeline that runs daily on Vertex AI Pipelines. The pipeline preprocesses data, trains a model, and deploys it to an endpoint. Recently, the pipeline has been failing at the deployment step because the endpoint already exists and the deploy step tries to create a new endpoint instead of updating the existing one. The pipeline code is written using the Kubeflow Pipelines SDK. You need to modify the pipeline to resolve this issue with minimal changes. What should you do?

A.Change the pipeline to use a Cloud Function that triggers the deployment independently, bypassing Vertex AI Pipelines.
B.In the deployment component, add a check to verify if the endpoint exists, and if so, call the update endpoint method instead of create.
C.Set the deploy component's retry policy to infinite so it eventually succeeds.
D.Manually delete the existing endpoint before each pipeline run.
AnswerB

Checking whether the endpoint already exists and calling the update method instead of create resolves the failure directly within the deployment component, requiring minimal changes to the existing Kubeflow Pipelines SDK code while preserving the daily pipeline structure.

Why this answer

The pipeline fails because the deployment component attempts to create a new endpoint when one already exists. The minimal change is to modify the deployment component to check for the endpoint's existence and, if it exists, call the update method instead of create. This ensures idempotency and resolves the failure without altering the pipeline structure or adding external dependencies.

Exam trap

PMLE often tests idempotency in pipeline components. Candidates might choose retries or manual deletion, but the correct answer is to implement conditional logic (create vs. update) within the component, as it's the minimal code change that addresses the root cause.

How to eliminate wrong answers

Option A is wrong because using a Cloud Function bypasses Vertex AI Pipelines, adding complexity and not addressing the root cause within the pipeline. Option C is wrong because setting infinite retries will not resolve the conflict; the create operation will continue to fail. Option D is wrong because manually deleting the endpoint before each run is not automated and defeats the purpose of a pipeline; it also causes downtime and is not a code-based solution.

541
MCQhard

A machine learning engineer is training a model on Vertex AI using a custom container. The training job uses a large dataset stored in BigQuery. The engineer wants to minimize data transfer costs and maximize training speed. Which of the following approaches is most efficient?

A.Use the BigQuery Storage API to read data directly into the training program with parallel streams.
B.Run a query in BigQuery to extract the data, then save it to a local SSD on the training VM.
C.Export the BigQuery table to Cloud Storage in TFRecord format, then read it using the tf.data API.
D.Use the BigQuery client library to run a query and fetch results into memory using the to_dataframe() method.
AnswerA

The BigQuery Storage API provides high-throughput, parallel access to BigQuery data directly from the training program. It avoids exporting data to Cloud Storage, reducing costs and latency. It also supports column filtering and predicate pushdown, which can reduce the amount of data read. This is the most efficient method for training on large BigQuery datasets.

Why this answer

The BigQuery Storage API is designed for high-performance data reading, offering parallel streams and efficient data transfer. It avoids the overhead of exporting data to Cloud Storage or loading into memory all at once. This makes it the most efficient and cost-effective approach for training on large BigQuery datasets.

Exam trap

The trap here is assuming that exporting to Cloud Storage is always necessary, but the BigQuery Storage API allows direct, efficient access.

542
MCQhard

A travel booking company has a real-time recommendation system that suggests hotels and flights to users. The model is served using TensorFlow Serving on a Google Kubernetes Engine (GKE) cluster with auto-scaling enabled. The cluster uses n1-standard-4 machine types. The team has set up Cloud Monitoring dashboards and alerts. Last week, during a major holiday promotion, the team noticed that the model's inference latency P99 increased from 150 ms to 450 ms over a 30-minute period, while the request throughput increased from 500 to 1,200 requests per second. CPU utilization across the cluster rose to 95%, but memory utilization remained at 60%. The model version and the serving infrastructure configuration have not changed since the last deployment. Which action should the team take to mitigate the latency issue?

A.Implement a feature engineering pipeline that compresses the input features to reduce data size and inference time.
B.Deploy a newer version of the model that uses a more efficient architecture to reduce computational complexity.
C.Increase the number of TensorFlow Serving instances by reducing the CPU request per pod in GKE to allow more pods per node.
D.Add more nodes to the GKE cluster to increase the total CPU resources available for serving.
AnswerD

Adding nodes directly addresses the CPU saturation driving the latency spike: throughput tripled while CPU hit 95%, yet memory stayed at 60%, so the bottleneck is compute, not capacity per pod. Horizontal node scaling gives TensorFlow Serving more cores to parallelise inference across, restoring P99 latency without altering the model or serving configuration.

Why this answer

The latency spike is caused by CPU saturation (95% utilization) under increased load (500 to 1,200 RPS). Adding more nodes to the GKE cluster directly increases the total CPU resources available, allowing the existing TensorFlow Serving pods to handle the higher throughput without contention. This is the most immediate and infrastructure-appropriate fix because the model version and serving configuration have not changed, ruling out model-level or code-level optimizations.

Exam trap

Google Cloud often tests the misconception that reducing per-pod CPU requests (Option C) is a valid scaling strategy, but in reality this increases overcommitment and can worsen latency under high load, whereas adding nodes (Option D) provides dedicated resources without contention.

How to eliminate wrong answers

Option A is wrong because compressing input features may reduce data size but does not address the root cause of CPU saturation; inference latency is dominated by model computation, not I/O, and the feature engineering pipeline is not part of the serving infrastructure. Option B is wrong because deploying a newer, more efficient model is a long-term optimization, not an immediate mitigation; the question states the model version has not changed and the issue is purely resource contention under load. Option C is wrong because reducing the CPU request per pod would allow more pods per node, but this would increase CPU overcommitment and worsen contention on already saturated nodes, potentially causing further latency degradation or pod evictions.

543
MCQeasy

An MLOps engineer needs to collect ground truth labels for a deployed classification model to compare predictions against actuals. Where should the engineer store the ground truth data to enable Vertex AI model quality monitoring?

A.BigQuery
B.Firestore
C.Cloud Spanner
D.Cloud Storage
AnswerA

Vertex AI model quality monitoring reads ground truth labels from a BigQuery table, joining them against logged predictions to compute accuracy, precision and recall. Storing labels in BigQuery satisfies this requirement because the monitoring pipeline expects that source, unlike Cloud Storage objects or local files.

Why this answer

Vertex AI Model Monitoring for model quality requires ground truth labels to be stored in BigQuery. The monitoring job compares the model's online predictions (logged to BigQuery via prediction logging) against the actual observed labels, which must also reside in BigQuery so the service can join them on a shared key column. BigQuery is the only supported sink for ground truth data in Vertex AI model quality monitoring.

Exam trap

PMLE often tests the assumption that any Google database can serve as a monitoring data source, but model quality monitoring specifically requires BigQuery because the service performs SQL joins between predictions and ground truth.

How to eliminate wrong answers

Option B is wrong because Firestore is a document NoSQL database used for application state, not the analytical store Vertex AI Model Monitoring reads ground truth from. Option C is wrong because Cloud Spanner is a globally distributed relational OLTP database; it is not an integration point for Vertex AI model quality monitoring. Option D is wrong because Cloud Storage can hold prediction logs, but ground truth labels for model quality monitoring must be in BigQuery so the service can perform the join and compute quality metrics.

544
MCQmedium

A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?

A.The endpoint is not configured for autoscaling; enable min_replica=0 to allow scale-to-zero.
B.The model file is corrupted; re-upload to Vertex AI Model Registry.
C.The container has a slow initialization; set initialDelaySeconds in the health probe to give more time before considering the pod ready.
D.Use a smaller machine type (n1-standard-2) to reduce startup overhead.
AnswerC

Slow container initialisation delays readiness, so the pod receives traffic before PyTorch and model weights finish loading, producing cold-start latency after scaling. Raising initialDelaySeconds postpones the readiness probe's first check, preventing premature traffic routing until initialisation completes. This directly addresses the scaling-induced latency constraint.

Why this answer

The high latency on the first few requests after scaling up is a classic symptom of a slow container initialization. By setting `initialDelaySeconds` in the health probe, you allow the container more time to start up and become ready before it receives traffic, preventing premature routing that causes timeouts or retries. This is a common tuning parameter for custom containers on Vertex AI, where model loading or dependency initialization can take several seconds.

Exam trap

The trap here is that candidates confuse slow initialization with autoscaling misconfiguration, assuming that scale-to-zero or smaller machines would fix the latency, when in fact the root cause is the readiness probe timing.

How to eliminate wrong answers

Option A is wrong because the problem occurs after scaling up, not from a cold start with zero replicas; setting min_replica=0 would actually worsen latency by requiring full cold starts. Option B is wrong because a corrupted model file would cause persistent prediction failures or errors, not just high latency on the first few requests after scaling. Option D is wrong because using a smaller machine type (n1-standard-2) would increase startup overhead and latency, not reduce it, as it provides fewer CPU and memory resources for initialization.

545
MCQmedium

A fraud detection team trains a model on data stored in BigQuery. They want to ensure that the model can be reproduced exactly one year later, including the specific data version and training code. They use Vertex AI Pipelines for orchestration. Which practice should they implement?

A.Export the training data to a Cloud Storage bucket and enable Object Versioning on the bucket.
B.Use BigQuery table snapshots and commit the pipeline definition to a Git repository, then log the snapshot ID as a pipeline parameter.
C.Enable Vertex AI Model Registry's automatic versioning and rely on the model's artifact URI to retrieve the training data.
D.Schedule a daily BigQuery export to Cloud Storage and use the latest export for training.
AnswerB

BigQuery table snapshots provide a point-in-time immutable copy of the training data. Storing the pipeline code in Git and passing the snapshot ID as a parameter ensures both data and code are versioned. This combination enables exact reproduction of the training run.

Why this answer

Reproducibility requires capturing both the code and the exact data state. BigQuery table snapshots provide an immutable copy of the training data at a point in time. By committing the pipeline definition to Git and logging the snapshot ID as a parameter, the team can rerun the pipeline with the same code and data, ensuring identical results.

Exam trap

The trap here is assuming that model versioning alone ensures reproducibility, when in fact the training data and code must also be versioned.

546
Multi-Selecthard

A company has a prototype ML model that predicts equipment failure. They want to deploy it to production using Vertex AI. The model must be retrained weekly with new data. They also need to monitor for data drift and model performance. Which THREE components should they include in their MLOps pipeline? (Choose 3)

Select 3 answers
A.A scheduled training pipeline that retrains the model weekly.
B.A manual QA step where data scientists approve each deployment.
C.A manual review of new data before it is used for training.
D.An automated trigger that redeploys the model when performance drops below a threshold.
E.A monitoring system that checks for data drift and triggers alerts.
AnswersA, D, E

Weekly retraining satisfies the stem's explicit cadence requirement. A scheduled Vertex AI pipeline automates ingestion of new data, retrains the model, and registers an updated version, removing manual intervention. This directly fulfils the stated need to retrain weekly with fresh data.

Why this answer

Option A is correct because the requirement to retrain weekly with new data is met by a scheduled training pipeline in Vertex AI, typically implemented with Vertex AI Pipelines and a Cloud Scheduler or scheduled run that re-executes the training job on the latest data. Option D is correct because automated redeployment when performance drops below a threshold is a core MLOps practice; in Vertex AI this can be done by monitoring model performance and using a trigger (for example, a Cloud Function or Eventarc/Cloud Monitoring alert) to promote a newly trained model to an endpoint. Option E is correct because Vertex AI Model Monitoring detects data drift and can emit alerts via Cloud Monitoring, which is exactly what the scenario requires for drift and performance oversight.

Option B is not correct because a manual QA approval step for every deployment is not required by the scenario and would slow down the automated weekly retraining and redeployment workflow. Option C is not correct because manual review of all new data before training is not a required MLOps component here and would prevent the pipeline from being fully automated and scheduled weekly.

Exam trap

Google Cloud often tests the distinction between necessary manual oversight and fully automated MLOps practices, leading candidates to overestimate the need for human approval steps in a production pipeline that demands speed and scalability.

547
MCQeasy

You want to use Vertex AI Vizier for hyperparameter tuning. You have 2 categorical parameters and 3 continuous parameters. Which algorithm is best suited for this mixed parameter space?

A.Evolutionary algorithm
B.Random search
C.Bayesian optimization
D.Grid search
AnswerC

Bayesian optimisation handles mixed categorical and continuous search spaces by building a probabilistic surrogate model and selecting promising trials. It suits Vertex AI Vizier's 2 categorical and 3 continuous parameters, converging on good configurations with fewer trials than grid or random search.

Why this answer

Bayesian optimization is the default and best-suited algorithm in Vertex AI Vizier for mixed parameter spaces containing both categorical and continuous parameters. It builds a probabilistic surrogate model of the objective function and uses an acquisition function to intelligently select the next hyperparameter configuration, handling categorical and continuous dimensions natively. This makes it far more sample-efficient than random or grid search when the search space is mixed and potentially high-dimensional.

Exam trap

The trap here is assuming that any algorithm works equally well for mixed parameter spaces; candidates often pick random search because it is simple, but the exam expects recognition that Bayesian optimization is the default and most efficient choice in Vizier for mixed categorical/continuous spaces.

How to eliminate wrong answers

Option A is wrong because the evolutionary algorithm in Vizier is designed for large-scale, high-dimensional search spaces and is typically recommended when the number of trials is very large or the search space is complex; it is not the default best choice for a small mixed space of 5 parameters. Option B is wrong because random search ignores the structure of the search space and does not learn from previous trials, making it far less sample-efficient than Bayesian optimization. Option D is wrong because grid search is not a Vizier-supported algorithm and would scale exponentially with the number of parameters, making it impractical for continuous parameters.

548
MCQhard

A team is scaling their prototype inference model to handle high-throughput requests with low latency. They use a custom container on Vertex AI Prediction. They notice that latency spikes occur under heavy load. What is the most effective strategy?

A.Enable auto-scaling with a higher minimum number of replicas.
B.Optimize model serving with batching and model warm-up.
C.Use a larger machine type with more CPUs.
D.Use a GPU-based machine.
AnswerB

Batching amortises per-request overhead across concurrent inputs, while warm-up pre-loads weights and initialises the container so the first requests avoid cold-start latency. Together they directly address the latency spikes observed under heavy load on the custom container.

Why this answer

Batching groups multiple inference requests into a single forward pass, dramatically improving GPU/CPU utilization and throughput, while model warm-up pre-loads weights and initializes the runtime so the first requests after a scale-out event do not pay cold-start latency. Together they address the latency spikes under heavy load without over-provisioning. This is the standard Vertex AI Prediction optimization pattern for custom containers.

Exam trap

PMLE often tests the reflex to throw hardware (bigger machine, GPU) at latency problems, when the expected answer is serving-layer optimization like batching and warm-up.

How to eliminate wrong answers

Option A is wrong because raising the minimum replica count increases cost and does not fix per-request latency spikes caused by cold starts or unbatched inference. Option C is wrong because more CPUs help only if the model is CPU-bound and single-threaded; most inference bottlenecks are memory bandwidth or GPU-bound, and vertical scaling has diminishing returns. Option D is wrong because a GPU speeds up computation per request but does not solve cold-start latency or inefficient request handling; without batching and warm-up, GPU utilization stays low under bursty load.

549
MCQmedium

An engineer needs to compile a Kubeflow Pipeline defined in Python to a JSON format that can be run on Vertex AI Pipelines. Which command should they use?

A.kfp.compiler.Compiler().compile(pipeline_func, 'pipeline.json')
B.gcloud ai pipelines compile command.
C.kfp.Client().upload_pipeline()
D.dsl.pipeline decorator automatically compiles at runtime.
AnswerA

The KFP compiler's compile method converts a Python pipeline function into an IR YAML or JSON specification that Vertex AI Pipelines can submit and execute. Passing the pipeline function and output filename produces the required JSON artefact directly.

Why this answer

The Kubeflow Pipelines SDK provides the `kfp.compiler.Compiler().compile()` method to convert a Python-based pipeline function into a JSON or YAML format that is compatible with Vertex AI Pipelines. This JSON representation defines the pipeline's components, dependencies, and execution graph, enabling it to be submitted to Vertex AI for orchestration. The `compile()` method is the standard way to produce a portable pipeline specification from Python code.

Exam trap

The trap here is that candidates may mistakenly believe that the `gcloud ai pipelines` command or the `dsl.pipeline` decorator directly compiles the pipeline, but in Google's Vertex AI Pipelines, you must use the KFP SDK's `Compiler().compile()` method to generate the pipeline JSON specification.

How to eliminate wrong answers

Option B is wrong because `gcloud ai pipelines compile` is not a valid gcloud command; the gcloud CLI for Vertex AI uses `gcloud ai pipelines run` to submit a pre-compiled pipeline, but compilation must be done separately using the KFP SDK. Option C is wrong because `kfp.Client().upload_pipeline()` uploads a compiled pipeline package to a Kubeflow Pipelines instance, but it does not perform the compilation step itself; the pipeline must already be compiled into a JSON or YAML file before uploading. Option D is wrong because the `dsl.pipeline` decorator defines the pipeline structure and components but does not automatically compile it at runtime; explicit invocation of `Compiler().compile()` is required to generate the JSON artifact.

550
MCQhard

A media company uses a Vertex AI Endpoint to serve a video recommendation model. They have enabled Vertex AI Model Monitoring for prediction drift. After a major news event, they observe a significant increase in prediction drift alerts, but the model's recommendations remain relevant and user engagement is stable. They want to reduce unnecessary alerts without losing the ability to detect true model degradation. What should they do?

A.Switch from prediction drift to training-serving skew monitoring.
B.Reduce the monitoring frequency from daily to weekly to decrease the number of alerts.
C.Disable prediction drift monitoring and rely on business metrics instead.
D.Increase the prediction drift threshold to a higher value based on historical variance.
AnswerD

Prediction drift alerts fire when the distribution of model outputs changes beyond a threshold. A major news event can cause a legitimate shift in user behavior and thus in prediction distributions, without indicating model degradation. Raising the threshold based on observed historical variance reduces false alerts while still catching significant deviations that could signal real problems. This balances sensitivity and specificity for the current environment.

Why this answer

Prediction drift alerts are based on changes in the model's output distribution. A major news event can legitimately shift user behavior and thus prediction distributions without degrading model quality. Raising the drift threshold based on historical variance reduces false alerts while preserving the ability to detect significant deviations.

Disabling monitoring, changing frequency, or switching to skew detection do not appropriately address the root cause.

Exam trap

The trap here is treating all prediction drift alerts as indicators of model failure, when legitimate external events can cause distribution shifts that require threshold recalibration rather than disabling monitoring.

551
MCQmedium

A data science team wants to build a machine learning pipeline on Vertex AI Pipelines that preprocesses data, trains a model, and evaluates it. They need to ensure that components can be reused across multiple pipelines and that outputs from one component can be passed as inputs to another. Which approach should they take?

A.Write each component as a Cloud Composer DAG task using Python operators and manage dependencies via Airflow.
B.Use Vertex AI pre-built components exclusively and chain them using the Vertex AI SDK without a pipeline definition.
C.Define each step as a separate Cloud Build step and chain them via build triggers.
D.Use Kubeflow Pipelines SDK v2 to create Python function components decorated with @dsl.component and compose them into a pipeline using @dsl.pipeline.
AnswerD

Decorating Python functions with @dsl.component produces self-contained, reusable components whose typed inputs and outputs Vertex AI Pipelines resolves automatically, so one component's output artefact can be wired directly into the next component's input. The @dsl.pipeline decorator then composes these into a directed acyclic graph, satisfying the reuse and data-passing constraints.

Why this answer

Kubeflow Pipelines SDK v2 with @dsl.component and @dsl.pipeline decorators is the native way to define reusable, composable components in Vertex AI Pipelines. This approach allows each component to be a self-contained Python function that can be independently versioned and reused across multiple pipelines, with outputs automatically serialized and passed as inputs to downstream components via the pipeline graph.

Exam trap

Google PMLE often tests the misconception that any orchestration tool (Airflow, Cloud Build) can substitute for a purpose-built ML pipeline framework, but the key differentiator is Vertex AI Pipelines' native support for reusable components with typed artifact passing and managed execution.

How to eliminate wrong answers

Option A is wrong because Cloud Composer (Airflow) is a workflow orchestrator for general DAGs, not a purpose-built ML pipeline framework; it lacks native support for Vertex AI Pipelines' artifact tracking, component reuse, and ML-specific I/O handling. Option B is wrong because Vertex AI pre-built components cannot be chained without a pipeline definition; the Vertex AI SDK requires a pipeline specification (e.g., via Kubeflow Pipelines) to define the execution graph and pass outputs between steps. Option C is wrong because Cloud Build is a CI/CD service for building and testing code, not for orchestrating ML pipelines; it does not provide managed artifact passing, caching, or the runtime environment needed for ML training and evaluation steps.

552
Matchingmedium

Match each feature engineering technique to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Convert categorical variable into binary columns

Combine two or more features to capture interactions

Normalize numeric features to a standard range

Group continuous values into discrete intervals

Weight term frequency by inverse document frequency

Why these pairings

This matching question tests your understanding of common feature engineering techniques. Polynomial Features create new features by raising existing features to a power (e.g., x², x³). One-Hot Encoding converts categorical variables into binary columns (0/1) for each category.

Feature Scaling standardizes numerical features to have zero mean and unit variance (z-score normalization). Options D and E are incorrect: D confuses Polynomial Features with Label Encoding (which assigns numerical codes), and E confuses One-Hot Encoding with Polynomial Features.

Exam trap

A common trap is mixing up Label Encoding (assigning ordinal numbers) with One-Hot Encoding (binary columns), or confusing Polynomial Features with interaction terms. Remember that Polynomial Features only involve powers of existing features, not combinations of different features.

553
MCQhard

A data engineering team wants to orchestrate an ML pipeline that includes data preprocessing in Dataflow, AutoML training, and model deployment. They want to minimize operational overhead. Which approach is best?

A.Use Cloud Composer with Apache Airflow DAG
B.Use AI Platform Training with script
C.Use Cloud Scheduler to trigger Cloud Functions
D.Use Vertex AI Pipelines with custom components
AnswerD

Vertex AI Pipelines orchestrates Dataflow preprocessing, AutoML training, and deployment as managed, serverless steps, minimising infrastructure to maintain. Custom components wrap the Dataflow job and AutoML task, so the whole workflow runs without the team provisioning or operating separate orchestration infrastructure.

Why this answer

Vertex AI Pipelines with custom components is the best choice because it provides a fully managed, serverless orchestration service that natively integrates with Dataflow, AutoML, and model deployment. This minimizes operational overhead by eliminating the need to manage infrastructure, handle retries, or maintain a separate orchestration server, while offering built-in artifact tracking and pipeline caching.

Exam trap

The trap here is that candidates often confuse 'orchestration' with 'scheduling' and pick Cloud Scheduler, failing to recognize that a multi-step ML pipeline requires workflow orchestration with dependencies and error handling, not just a time-based trigger.

How to eliminate wrong answers

Option A is wrong because Cloud Composer with Apache Airflow DAG requires managing a Kubernetes cluster, Airflow workers, and infrastructure, which increases operational overhead rather than minimizing it. Option B is wrong because AI Platform Training with a script only handles the training step in isolation, not the end-to-end orchestration of preprocessing, training, and deployment. Option C is wrong because Cloud Scheduler to trigger Cloud Functions is a simple time-based trigger that lacks the workflow orchestration capabilities (e.g., conditional branching, parallel steps, dependency management) needed for a multi-step ML pipeline.

554
MCQeasy

A team just moved a model from prototype to production using Vertex AI. They notice prediction errors for certain inputs that were not present in training data. What should they do to detect such issues automatically?

A.Set up Vertex AI Experiments to compare predictions
B.Use BigQuery ML to analyze prediction requests
C.Enable Cloud Logging and set up alerts for error logs
D.Enable Vertex AI Model Monitoring to detect prediction anomalies
AnswerD

Vertex AI Model Monitoring compares incoming prediction requests against the training baseline and flags anomalies such as feature skew or out-of-range inputs. This automatically surfaces the unseen inputs that cause prediction errors, without manual inspection.

Why this answer

Vertex AI Model Monitoring is specifically designed to detect prediction anomalies, such as data drift and feature skew, by comparing production prediction requests against the training data distribution. This allows the team to automatically identify inputs that deviate from the training data, even if those exact inputs were not present during training, without manual inspection.

Exam trap

Google Cloud often tests the distinction between monitoring for operational errors (e.g., HTTP errors) versus monitoring for model-specific issues (e.g., data drift), leading candidates to choose Cloud Logging (Option C) when the correct answer requires a dedicated ML monitoring service.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is used for tracking and comparing model training runs and hyperparameter tuning, not for monitoring production prediction requests or detecting anomalies in real-time. Option B is wrong because BigQuery ML is a tool for creating and executing machine learning models directly in BigQuery using SQL, not for analyzing prediction requests from a deployed Vertex AI model or detecting input anomalies. Option C is wrong because while Cloud Logging can capture error logs, it only reacts to explicit errors (e.g., 4xx/5xx HTTP responses) and cannot automatically detect prediction anomalies like data drift or feature skew that do not generate error logs.

555
MCQeasy

A retail company wants to build a recommendation system to show 'frequently bought together' items. Which Recommendations AI model type should they use?

A.recently-viewed
B.frequently-bought-together
C.recommended-for-you
D.others-you-may-like
AnswerB

The frequently-bought-together model type is purpose-built to surface item pairs or sets commonly purchased in the same transaction, matching the retail 'frequently bought together' requirement. Other model types target different goals such as personalised ranking or similar-item recommendations.

Why this answer

The 'frequently-bought-together' model type in Google Cloud Recommendations AI is specifically designed to identify items that are commonly purchased in the same transaction, using co-occurrence analysis of historical purchase data. This directly matches the requirement to show items that are frequently bought together, leveraging association rule mining (e.g., Apriori algorithm) to generate recommendations.

Exam trap

In Google Cloud Recommendations AI, the key distinction is between transaction-based co-purchase models ('frequently-bought-together') and personalized recommendation models ('recommended-for-you'). Candidates may confuse 'frequently-bought-together' with 'others-you-may-like' because both relate to item similarity, but only 'frequently-bought-together' uses co-occurrence analysis of actual transactions.

How to eliminate wrong answers

Option A is wrong because 'recently-viewed' is a model type that surfaces items a user has recently browsed, not items that are frequently purchased together, and it relies on session-based user behavior rather than transaction co-occurrence. Option C is wrong because 'recommended-for-you' is a personalized model that uses user-item interaction history (e.g., collaborative filtering) to suggest items tailored to an individual, not cross-item purchase patterns. Option D is wrong because 'others-you-may-like' is a similarity-based model that recommends items similar to a given product based on content or metadata, not on transactional co-purchase frequency.

556
MCQeasy

An ML team wants to automatically retrain a model when data drift is detected. They have set up a Cloud Monitoring alert on drift. What service should they use to trigger a retraining pipeline in response to the alert?

A.Cloud Functions
B.Cloud Scheduler
C.Vertex AI Feature Store
D.Vertex AI Model Monitoring
AnswerA

Cloud Functions provides an event-driven, serverless trigger that fires when the Cloud Monitoring drift alert publishes to a Pub/Sub topic, invoking the retraining pipeline. This satisfies the requirement to react automatically to drift alerts without managing servers.

Why this answer

Cloud Functions is the correct service to trigger a retraining pipeline in response to a Cloud Monitoring alert. Cloud Monitoring can publish alerts to a Pub/Sub topic, and a Cloud Function subscribed to that topic can invoke the Vertex AI pipeline or trigger a Cloud Scheduler job that starts retraining. This event-driven pattern is the standard way to automate retraining on drift detection.

Exam trap

PMLE often tests the distinction between the service that detects drift (Vertex AI Model Monitoring) and the service that reacts to the alert (Cloud Functions), tempting candidates to pick the monitoring service itself as the trigger.

How to eliminate wrong answers

Option B is wrong because Cloud Scheduler runs jobs on a fixed time schedule, not in response to a monitoring alert; it cannot react to drift events. Option C is wrong because Vertex AI Feature Store is a feature management service for serving and sharing features, not an event-triggering mechanism. Option D is wrong because Vertex AI Model Monitoring is the service that detects drift and emits alerts, but it does not itself trigger a retraining pipeline — it is the source of the alert, not the responder.

557
MCQmedium

A healthcare company has deployed a diagnostic model on Vertex AI Endpoints. They use Vertex AI Model Monitoring to detect drift in features and predictions. The model's input features include patient age, which is a numerical feature. The monitoring job has been running for a month and has generated several alerts for age drift. However, the model's performance has not degraded. The team wants to reduce false alerts without missing critical drifts. What should they do?

A.Switch from Vertex AI Model Monitoring to a custom monitoring solution using Cloud Monitoring to have more control over alert thresholds.
B.Disable drift monitoring for the age feature entirely, since it is not causing performance degradation.
C.Analyze the distribution of age in the serving data over time to determine if the drift is due to expected seasonal variation or a data pipeline issue, then adjust the monitoring configuration accordingly.
D.Increase the drift threshold for the age feature to reduce sensitivity, and monitor other features more closely.
AnswerC

Analyzing the age distribution helps distinguish between expected variation (e.g., seasonal changes in patient demographics) and genuine issues like data pipeline errors. Based on this analysis, the team can adjust thresholds or exclude the feature from drift monitoring if it's not predictive of performance. This targeted approach reduces false alerts while maintaining vigilance for critical drifts.

Why this answer

Analyzing the age distribution over time helps determine if the drift is due to expected variability or a data issue. This allows the team to adjust monitoring configuration, such as setting a more appropriate threshold or excluding the feature if it's not performance-critical, thereby reducing false alerts without missing important drifts.

Exam trap

The trap here is to either ignore the alerts or disable monitoring, when the correct approach is to investigate the nature of the drift and tune the monitoring configuration accordingly.

558
MCQmedium

An ML engineer is training a PyTorch model on Vertex AI using a custom training job. The dataset is stored in a Cloud Storage bucket with 500,000 small JPEG images. The training job is configured with an n1-standard-8 machine and a single NVIDIA T4 GPU. The engineer observes that GPU utilization is very low (around 15%) and training is slow. The model code is not the bottleneck. What is the most likely cause and the best solution?

A.The NVIDIA T4 GPU is not powerful enough for this model; upgrade to an NVIDIA A100 GPU.
B.The training job is using a single worker and not distributing the load; switch to distributed training with multiple workers.
C.The Cloud Storage bucket is in a different region than the training job; move the bucket to the same region.
D.The data loading pipeline is inefficient due to many small files; convert the dataset to TFRecord or WebDataset format and use a larger batch size with prefetching.
AnswerD

Reading many small JPEGs individually from Cloud Storage introduces high per-file latency and I/O overhead, starving the GPU. Converting to a sharded, sequential format like TFRecord or WebDataset allows efficient streaming, and using prefetching and larger batches keeps the GPU fed, increasing utilization.

Why this answer

Low GPU utilization with many small files typically points to an I/O-bound data pipeline. Cloud Storage has high per-file access latency, so reading hundreds of thousands of small JPEGs individually creates a bottleneck. Converting to a sharded format like TFRecord or WebDataset enables efficient sequential reads, and combining with prefetching and larger batches ensures the GPU is continuously fed, improving utilization and training speed.

Exam trap

The trap here is assuming that low GPU utilization always means insufficient GPU compute and upgrading the GPU, instead of diagnosing the data input pipeline.

559
Multi-Selecthard

A team is building a batch prediction pipeline that processes raw data from Cloud Storage, performs complex preprocessing, and then runs predictions using a large model. The preprocessing step is compute-intensive and the prediction step is I/O-bound. Which TWO Google Cloud services should they combine to optimize cost and performance? (Choose 2)

Select 2 answers
A.Dataflow for preprocessing and writing results to Cloud Storage
B.Cloud Functions to preprocess data row by row
C.Cloud Run to serve the preprocessed data as an API
D.Vertex AI Batch Prediction with Cloud Storage source
E.Vertex AI Batch Prediction with BigQuery source
AnswersA, D

Dataflow provides autoscaling, distributed workers suited to the compute-intensive preprocessing stage, reading raw Cloud Storage data and writing transformed output back to Cloud Storage. This satisfies the stem's constraint by matching elastic compute to preprocessing while keeping the I/O-bound prediction stage separate.

Why this answer

Option A (Dataflow for preprocessing and writing results to Cloud Storage) is correct because Dataflow is a fully managed, autoscaling service built on Apache Beam that handles compute-intensive, parallel batch preprocessing over data in Cloud Storage, and it can write the transformed output back to Cloud Storage for the next stage. Option D (Vertex AI Batch Prediction with Cloud Storage source) is correct because it is the managed, I/O-bound batch inference service designed to read preprocessed files directly from Cloud Storage and run predictions on a large model without provisioning persistent serving infrastructure, which optimizes cost for batch workloads. Option B is wrong because Cloud Functions is event-driven and limited in execution time and memory, making it unsuitable for compute-intensive row-by-row preprocessing of large batch datasets.

Option C is wrong because Cloud Run serving preprocessed data as an API adds unnecessary request/response overhead and is designed for online serving, not batch pipeline handoff. Option E is wrong because the pipeline explicitly reads raw data from Cloud Storage and writes preprocessed results there, so a BigQuery source for batch prediction does not match the described data flow.

Exam trap

Google often tests the distinction between batch and online serving patterns, and the trap here is that candidates may choose Cloud Functions or Cloud Run for preprocessing because they are familiar serverless options, without realizing that Dataflow is purpose-built for large-scale, compute-intensive batch processing and that Vertex AI Batch Prediction is the correct service for offline inference at scale.

560
Drag & Dropmedium

Drag and drop the steps to set up a batch prediction job using Vertex AI in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct sequence for setting up a batch prediction job using Vertex AI is to first prepare your input data, then register your model, create the batch prediction job, submit the job, and finally retrieve the results. This order ensures that all necessary resources are available before job creation and submission, and that results are only requested after job completion.

561
MCQmedium

A machine learning engineer needs to create a pipeline that runs a custom container component on Vertex AI. The container expects a Cloud Storage path as input and outputs a model artifact. Which component type should they define using the Kubeflow Pipelines SDK v2?

A.Google Cloud Pipeline Components (GCPC) for custom containers
B.Python function component using @dsl.component
C.Importer component to load the container as an artifact
D.Container component using @dsl.container_component
AnswerD

The @dsl.container_component decorator lets you define a component from a custom container image, with typed inputs and outputs declared in the function signature. This satisfies the stem's need to pass a Cloud Storage path in and emit a model artifact out.

Why this answer

The Kubeflow Pipelines SDK v2 provides the @dsl.container_component decorator specifically for defining components that wrap custom container images. This allows the engineer to specify the container image, input/output paths (like a Cloud Storage path), and artifact metadata, enabling Vertex AI to execute the container as a pipeline step and capture the model artifact.

Exam trap

The trap here is that candidates confuse the Importer component (which only imports existing artifacts) with a component that runs a container to produce an artifact, or they mistakenly think Google Cloud Pipeline Components can wrap any custom container when they only provide pre-built Google service integrations.

How to eliminate wrong answers

Option A is wrong because Google Cloud Pipeline Components (GCPC) are pre-built components for Google Cloud services (e.g., AI Platform, BigQuery), not for wrapping arbitrary custom containers. Option B is wrong because @dsl.component is used for Python function components that execute inline Python code, not for running a custom container image. Option C is wrong because the Importer component is used to import existing artifacts (like a pre-trained model) into the pipeline's metadata store, not to run a container that produces an artifact.

562
MCQeasy

A retail company wants to forecast daily sales for inventory planning. They have 3 years of historical sales data with clear weekly and yearly seasonality. Which approach should they use?

A.Call a pre-built Google Cloud API for sales prediction
B.Use a linear regression model in Vertex AI
C.Use Vertex AI AutoML Tables with date as feature
D.Use BigQuery ML to train an ARIMA_PLUS model
AnswerD

ARIMA_PLUS in BigQuery ML handles time-series forecasting natively, automatically modelling weekly and yearly seasonality plus trend from the three years of historical data. This satisfies the stated seasonality requirement without manual feature engineering, and runs where the sales data already resides.

Why this answer

ARIMA_PLUS in BigQuery ML is specifically designed for time-series forecasting with multiple seasonalities (weekly and yearly). It automatically handles seasonality detection, trend decomposition, and holiday effects, making it ideal for retail sales data with clear periodic patterns.

Exam trap

The trap here is that candidates often choose AutoML Tables (Option C) thinking it can handle any structured data, but they miss that AutoML Tables is not a dedicated time-series model and requires manual feature engineering to capture seasonality, whereas ARIMA_PLUS is purpose-built for this scenario.

How to eliminate wrong answers

Option A is wrong because calling a pre-built Google Cloud API for sales prediction is vague and not a specific, integrated solution for time-series forecasting with seasonality; such APIs may not exist or may not handle custom seasonality patterns. Option B is wrong because linear regression in Vertex AI is a general-purpose model that does not inherently capture time-series dependencies like weekly and yearly seasonality without extensive feature engineering (e.g., lag features, Fourier terms). Option C is wrong because Vertex AI AutoML Tables with date as a feature treats the problem as a regression on tabular data, not as a dedicated time-series model, and may fail to properly model temporal autocorrelation and multiple seasonalities without manual time-series preprocessing.

563
MCQmedium

A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?

A.Create a Dataflow pipeline that reads from the original BigQuery table, applies Cloud DLP de-identification transforms, and writes to a new BigQuery table. Grant the data science team access to the de-identified table.
B.Enable Shielded VM on Vertex AI Workbench notebooks and use VPC-SC to restrict data access.
C.Use Cloud Key Management Service to encrypt the PHI columns in BigQuery, and share the encryption key with the data science team.
D.Use BigQuery row-level security to mask PHI columns for the data science team, and train the model directly on the original table.
AnswerA

Cloud DLP de-identification transforms applied in a Dataflow pipeline remove direct identifiers before the data lands in a separate BigQuery table, satisfying the de-identification requirement while BigQuery's default encryption at rest and TLS in transit plus audit logging cover the remaining policy constraints.

Why this answer

It uses Cloud DLP within a Dataflow pipeline to automatically de-identify PHI data as it is read from the original BigQuery table and written to a new, de-identified table. This satisfies the requirement for automated de-identification, while the original table remains encrypted at rest (BigQuery default) and in transit (TLS), and access to the original data can be logged via Cloud Audit Logs. The data science team only gets access to the de-identified table, ensuring PHI is not exposed during model development.

Exam trap

Google Cloud often tests the distinction between data masking/encryption (which still exposes PHI to authorized users) and true de-identification (which removes or transforms PHI so it is no longer considered protected health information).

How to eliminate wrong answers

Option B is wrong because Shielded VM and VPC-SC provide infrastructure security (integrity, network perimeter) but do not de-identify PHI data; the data science team would still see raw PHI in the notebooks. Option C is wrong because Cloud KMS encryption protects data at rest but does not remove or mask PHI columns; sharing the encryption key with the data science team would give them access to the raw PHI, violating the de-identification requirement. Option D is wrong because BigQuery row-level security masks columns at query time but does not de-identify the underlying data; the model training would still use the original table with PHI present in the masked columns, and the masking is not a permanent de-identification suitable for an automated pipeline.

564
Multi-Selectmedium

A data science team is configuring Vertex AI Model Monitoring for a deployed model. They want to detect both feature skew and feature drift. Which TWO configurations must they set?

Select 2 answers
A.Enable request/response logging.
B.Select the Jensen-Shannon divergence algorithm.
C.Configure the monitoring frequency (e.g., hourly).
D.Set the sampling rate to 1.0 (100%).
E.Specify a training dataset or statistics for baseline.
AnswersC, E

Required to define how often drift is computed.

Why this answer

Option C is correct because Vertex AI Model Monitoring requires a monitoring schedule/frequency (for example hourly or daily) so the job knows how often to evaluate incoming prediction data against the baseline; without a configured frequency, no monitoring analysis runs. Option E is correct because detecting feature skew and drift requires a baseline: for skew you supply the training dataset (or its generated statistics), and for drift you compare recent production data against the baseline statistics, so specifying a training dataset or statistics is mandatory. Option A is not required because request/response logging is a separate feature for capturing payloads and is not the configuration that enables skew/drift detection.

Option B is not required because Jensen-Shannon divergence is only one selectable distance metric for drift/skew detection, not a mandatory setting. Option D is not required because a 1.0 sampling rate (100%) is optional; monitoring can run on a lower sample of requests.

Exam trap

The trap is conflating 'enabling features' (logging, sampling rate) with 'required configurations' — candidates pick logging or sampling because they sound necessary, but the exam wants the two settings that directly define what to compare (baseline) and when (frequency).

565
MCQeasy

Refer to the exhibit. The team notices that the pipeline fails to read data from the specified Cloud Storage path. What is the most likely issue?

A.The bucket does not exist
B.The pipeline runner is incorrect
C.The region is mismatched
D.The service account lacks `storage.objectViewer` permission
AnswerD

A missing `storage.objectViewer` role on the service account directly prevents the pipeline from listing and reading objects at the Cloud Storage path, satisfying the stem's read-failure constraint. Without this permission, requests return 403 errors, so the pipeline cannot access the data regardless of path correctness.

Why this answer

The pipeline fails to read data from Cloud Storage because the service account lacks the `storage.objectViewer` IAM role, which grants the `storage.objects.get` and `storage.objects.list` permissions required to read objects. Without this role, the pipeline cannot authenticate or authorize the read operation, even if the bucket and path are correct.

Exam trap

Google Cloud often tests the distinction between bucket-level permissions (like `storage.objectViewer`) and project-level roles, leading candidates to overlook that the service account must have the specific IAM role on the bucket or project, not just any storage role.

How to eliminate wrong answers

Option A is wrong because if the bucket did not exist, the error would typically be a 404 'Bucket not found' or a similar explicit message, not a generic read failure. Option B is wrong because the pipeline runner (e.g., Dataflow, Apache Beam) is responsible for executing the pipeline logic, not for authenticating to Cloud Storage; a runner mismatch would cause execution errors, not permission-related read failures. Option C is wrong because Cloud Storage bucket access is global and region-mismatch errors occur only for specific operations like writing to a regional bucket from a different region, but reading is allowed across regions; a region mismatch would not block read access.

566
MCQhard

Refer to the exhibit. An engineer notices no drift alerts but the model performance has degraded. What is the likely cause?

A.Feature attribution monitoring is causing too many false positives
B.Drift threshold for income is too high
C.Skew thresholds are not configured for categorical features
D.Concept drift is occurring, which is not captured by drift or skew detection
AnswerD

Concept drift changes the input-to-output mapping itself, so incoming feature distributions still match the training baseline and drift or skew detection stays silent. The model's learned relationship is stale, which explains degraded performance without alerts.

Why this answer

Concept drift occurs when the statistical properties of the target variable change over time, causing model performance to degrade even when the input data distribution remains stable. Drift detection (e.g., data drift or skew) monitors changes in feature distributions, not the relationship between features and the target. Since no drift alerts were triggered, the input data appears unchanged, but the model's predictive relationship has shifted — this is classic concept drift, which requires performance monitoring (e.g., accuracy, F1-score) rather than drift or skew detection.

Exam trap

Google Cloud often tests the distinction between data drift (input feature changes) and concept drift (target relationship changes), trapping candidates who assume that no drift alerts mean the model is healthy, when in fact performance degradation can occur without any feature distribution shift.

How to eliminate wrong answers

Option A is wrong because feature attribution monitoring (e.g., SHAP values) explains model predictions but does not generate false positives for drift; it is unrelated to the absence of drift alerts. Option B is wrong because a drift threshold for income being too high would suppress drift alerts for that feature, but the scenario states no drift alerts at all, and concept drift is not captured by feature-level drift thresholds. Option C is wrong because skew thresholds for categorical features detect distribution shifts in those features, but the problem is a change in the target relationship (concept drift), not a change in feature distributions.

567
Multi-Selectmedium

A team wants to serve a large PyTorch model (3 GB) for online predictions with low latency. Which THREE actions should they take?

Select 3 answers
A.Use a custom container that preloads the model into memory.
B.Use batch prediction instead of online prediction.
C.Use a machine type with a GPU accelerator.
D.Optimize the model using TorchScript or quantization.
E.Deploy in multiple regions with Cloud Load Balancing.
AnswersA, C, D

Preloading the 3 GB model into memory inside a custom container eliminates per-request model loading, which otherwise dominates latency for a model this size. This directly addresses the low-latency constraint by removing cold-start and disk-read overhead from the prediction path.

Why this answer

Option A is correct because a custom container that preloads the 3 GB PyTorch model into memory eliminates the per-request model-loading overhead, which is essential for low-latency online predictions. Option C is correct because attaching a GPU accelerator provides the parallel compute throughput PyTorch inference needs, drastically reducing latency for a large model compared to CPU-only serving. Option D is correct because TorchScript (tracing/scripting for optimized execution) and quantization (e.g., FP16/INT8) reduce model size and inference cost, directly lowering latency while preserving online serving.

Option B is wrong because batch prediction is asynchronous and high-throughput, not low-latency online serving, so it contradicts the stated requirement. Option E is wrong because multi-region deployment with Cloud Load Balancing improves global availability and proximity, but does not address the model's size or per-inference latency for online predictions.

Exam trap

PMLE often tests whether candidates pick availability/scaling options (multi-region) when the real issue is model size and inference speed; the trap is confusing latency reduction with redundancy.

568
MCQhard

A company runs a Vertex AI Pipeline that trains a model and then deploys it to a Vertex AI Endpoint. The pipeline uses a conditional deployment step based on the model's evaluation metric. The team wants to ensure that if the evaluation metric falls below a threshold, the pipeline fails and no deployment occurs. Which approach should they use?

A.Add a custom component that evaluates the metric and, if below threshold, calls the Vertex AI API to cancel the pipeline.
B.Use a Condition to check the metric, and in the else branch, include a component that raises an exception or returns a failure status.
C.Configure the pipeline to use a fail-fast strategy by setting the pipeline's failure_policy to FAIL_FAST.
D.Use a Condition in the pipeline that checks if the metric is greater than the threshold, and only then execute the deployment component.
AnswerB

This approach ensures that if the metric is below threshold, the else branch executes a component that fails the pipeline. If the metric is above threshold, the deployment component runs. This provides both conditional deployment and a clear failure signal when the metric does not meet requirements.

Why this answer

To conditionally deploy and fail the pipeline when the metric is unsatisfactory, the pipeline should use a Condition with an else branch that triggers a failing component. This ensures that deployment only happens when the metric passes, and the pipeline fails otherwise. Other methods either do not cause failure or rely on unsupported features.

Exam trap

The trap here is believing that a Condition alone can fail the pipeline or that a built-in fail-fast policy exists, when in fact you must explicitly cause a failure in the else branch.

569
MCQmedium

A data science team wants to build a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it if the accuracy exceeds 0.9. They want to use the Kubeflow Pipelines SDK v2. Which construct allows them to conditionally execute the deployment step based on the evaluation metric?

A.dsl.If
B.dsl.Conditional
C.dsl.ExitHandler
D.dsl.Collected
AnswerA

dsl.If is the KFP SDK v2 conditional construct that evaluates a pipeline parameter or task output at runtime and executes its enclosed tasks only when the boolean condition holds. It satisfies the stem's requirement to deploy only when evaluation accuracy exceeds 0.9.

Why this answer

In Kubeflow Pipelines SDK v2, `dsl.If` is the correct construct for conditionally executing pipeline steps based on runtime metrics or parameters. It allows you to define a condition that, when evaluated to true, triggers the deployment step only if the model accuracy exceeds 0.9. This is the standard way to implement branching logic in v2 pipelines.

Exam trap

Candidates often confuse `dsl.If` with `dsl.Conditional`, but in Kubeflow Pipelines SDK v2 for Vertex AI, `dsl.If` is the correct construct for conditional execution based on runtime metrics.

How to eliminate wrong answers

Option B is wrong because `dsl.Conditional` is not a valid construct in Kubeflow Pipelines SDK v2; the correct name is `dsl.If`. Option C is wrong because `dsl.ExitHandler` is used to execute cleanup or notification steps when a pipeline or component exits, not for conditional branching based on evaluation metrics. Option D is wrong because `dsl.Collected` is used to gather outputs from parallel tasks (e.g., from a loop or fan-out), not to conditionally execute a step.

570
Multi-Selecteasy

Which TWO of the following can be used as input sources for Vertex AI batch prediction jobs? (Choose 2)

Select 2 answers
A.Cloud Firestore
B.Cloud SQL
C.BigQuery
D.Cloud Spanner
E.Cloud Storage
AnswersC, E

BigQuery is a supported input source for Vertex AI batch prediction jobs, letting you point the job directly at a table or query. This avoids exporting data to Cloud Storage first, satisfying the question's requirement for valid batch prediction input sources.

Why this answer

Option C (BigQuery) is correct because Vertex AI batch prediction jobs natively accept a BigQuery table as the input source, specified via the bigQuerySource field in the BatchPredictionJob, allowing you to run predictions directly against tabular data stored in BigQuery. Option E (Cloud Storage) is correct because Vertex AI batch prediction jobs also accept input files (such as JSONL, CSV, or TFRecord) stored in a Cloud Storage bucket, specified via the gcsSource field with a URI like gs://bucket/file.jsonl. Options A (Cloud Firestore), B (Cloud SQL), and D (Cloud Spanner) are not valid input sources for Vertex AI batch prediction; these are operational databases that are not directly supported as batch prediction inputs, so you would need to export their data to BigQuery or Cloud Storage first.

Exam trap

The trap here is that candidates often assume any Google Cloud database (like Firestore, Cloud SQL, or Spanner) can serve as a direct input source for batch predictions, but Vertex AI batch prediction only supports BigQuery and Cloud Storage as input sources, requiring data to be exported or staged in those services first.

571
MCQeasy

A company deploys a model on Vertex AI Prediction for real-time inference. Users report intermittent high latency during peak hours. The model is deployed on a single machine type with `min_replica_count=1` and `max_replica_count=5`. Autoscaling is enabled based on CPU utilization. What is the most likely cause of the latency spikes?

A.The model server is crashing under load due to memory issues.
B.Autoscaling based on CPU utilization does not react quickly to inference request spikes.
C.The load balancer is misconfigured and routes traffic unevenly.
D.The container image is not optimized for the model.
AnswerB

CPU utilisation lags actual inference demand: it only rises after requests queue, so the autoscaler reacts after latency has already spiked. With min_replica_count=1, peak-hour bursts hit a single replica before new ones provision, producing the intermittent spikes described.

Why this answer

CPU utilization may not be a good proxy for inference load; the system may not scale up fast enough under sudden traffic bursts. Option A is wrong because Vertex AI automatically manages container health. Option C is wrong because Vertex AI endpoints automatically distribute traffic.

Option D is wrong because the container image is built correctly.

572
MCQmedium

You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?

A.Deploy the new model as a separate endpoint and use a load balancer to split traffic.
B.Deploy the new model to the same endpoint with 100% traffic and monitor closely.
C.Create a new endpoint for the new model and use Vertex AI's traffic director to shift traffic.
D.Upload the new model as a new version and deploy it to the same endpoint with 10% traffic split.
AnswerD

Vertex AI Endpoints support multiple deployed models with configurable traffic splits. By deploying the new version to the same endpoint and assigning 10% of traffic, you can canary test it. You can then monitor metrics and gradually increase the split. This is the native, straightforward approach that integrates with Vertex AI monitoring and rollback.

Why this answer

Vertex AI Endpoints natively support deploying multiple model versions and splitting traffic by percentage. Deploying the new version to the same endpoint with a 10% split enables canary testing, and the split can be adjusted as confidence grows. Separate endpoints with external load balancing or a full cutover do not meet the controlled, gradual rollout requirement.

Exam trap

The trap here is thinking you need a separate endpoint or an external load balancer, when Vertex AI Endpoints already provide built-in traffic splitting.

573
MCQeasy

What is the primary benefit of using pipeline caching in Vertex AI Pipelines?

A.It reduces execution time and cost by reusing unchanged component outputs.
B.It encrypts data at rest.
C.It automatically scales the pipeline resources.
D.It enables parallel execution of components.
AnswerA

Pipeline caching stores each component's outputs keyed by its inputs and code version, so unchanged steps skip re-execution entirely. This directly satisfies the stem's constraint of reducing execution time and cost, since Vertex AI Pipelines avoids recomputing identical artefacts and only reruns components whose inputs or definitions changed.

Why this answer

Pipeline caching in Vertex AI Pipelines automatically detects when a component's inputs and code have not changed from a previous execution and reuses the cached output artifacts. This avoids redundant computation, directly reducing both execution time and cost by skipping re-execution of unchanged steps.

Exam trap

The exam often tests the distinction between caching (reusing outputs) and parallelization (running components concurrently), so candidates may confuse the two and incorrectly select parallel execution as the primary benefit.

How to eliminate wrong answers

Option B is wrong because encryption at rest is a data security feature managed by Cloud KMS or default Google Cloud encryption, not a benefit of pipeline caching. Option C is wrong because automatic scaling of pipeline resources is handled by Vertex AI's underlying infrastructure (e.g., node auto-scaling) or custom configuration, not by caching. Option D is wrong because parallel execution of components is achieved through pipeline design (e.g., using `dsl.ParallelFor` or independent component dependencies), not through caching; caching can actually reduce the need for parallel execution by reusing results.

574
MCQmedium

An ML team is building a feature pipeline with Dataflow that reads from BigQuery, computes features, and writes to Vertex AI Feature Store. They need to ensure that features are available for both training and serving with low latency. Which Feature Store option should they use?

A.Create a featurestore with only offline serving
B.Store features directly in BigQuery
C.Use Cloud SQL as a feature store
D.Create a featurestore with online serving enabled
AnswerD

Enabling online serving on the featurestore provisions a low-latency endpoint that serves features in real time, while the same store retains the data for training retrieval. This dual capability satisfies both training and serving requirements from one featurestore.

Why this answer

To ensure features are available for both training and serving with low latency, the featurestore must have online serving enabled. Vertex AI Feature Store provides both online serving (low-latency reads for real-time predictions) and offline serving (batch reads for training). Enabling online serving allows the same feature values to be served at low latency during prediction, ensuring consistency between training and serving.

Exam trap

The trap is assuming that storing features in BigQuery or Cloud SQL is sufficient for low-latency serving; candidates may overlook that Vertex AI Feature Store's online serving is specifically designed for millisecond-latency access and integration with training pipelines.

How to eliminate wrong answers

Option A is wrong because a featurestore with only offline serving cannot provide low-latency online serving for real-time predictions, failing the requirement. Option B is wrong because storing features directly in BigQuery does not provide the low-latency online serving needed for real-time inference; BigQuery is optimized for analytical queries, not millisecond-latency lookups. Option C is wrong because Cloud SQL is a relational database not designed as a feature store; it lacks the integration with Vertex AI and the optimized online serving capabilities of Feature Store.

575
MCQhard

You have deployed a model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with a monitoring frequency of every 24 hours. The model serves predictions with a feature called 'transaction_amount' that has a skewed distribution. After several days, you notice that the drift metrics for this feature are consistently below the threshold, but you suspect that the feature distribution has actually shifted. You want to improve the sensitivity of drift detection for this feature. What should you do?

A.Enable explainable AI on the endpoint to get feature attributions for 'transaction_amount'.
B.Change the monitoring frequency to every hour to capture more data points.
C.Increase the sampling rate to analyze a larger proportion of prediction requests.
D.Adjust the drift threshold for 'transaction_amount' to a lower value to trigger alerts more easily.
AnswerD

Lowering the threshold for the feature's drift metric makes the monitoring job more sensitive to deviations from the reference distribution. If the feature's distribution has shifted but the drift value is below the default threshold, reducing the threshold will cause alerts to fire when the shift occurs, improving sensitivity.

Why this answer

To increase sensitivity to drift for a specific feature, you can lower the drift threshold for that feature. This makes the monitoring job trigger alerts at smaller deviations from the reference distribution. Changing frequency or sampling rate affects data volume or timeliness, not the sensitivity of the drift metric itself.

Exam trap

The trap here is thinking that more frequent monitoring or higher sampling automatically increases sensitivity, while the direct way to make detection more sensitive is to adjust the threshold.

576
MCQmedium

A manufacturing company wants to predict equipment failure using sensor data stored in BigQuery. They have limited ML expertise and want to use AutoML Tables. The data includes timestamps, numerical sensor readings, and a boolean 'failure' column. The dataset is highly imbalanced with only 1% failure cases. Which of the following is the most effective approach to handle the imbalance in AutoML Tables?

A.Let AutoML Tables handle the imbalance automatically; it has built-in techniques for class imbalance.
B.Downsample the majority class to balance the dataset.
C.Use a custom loss function in the training configuration.
D.Oversample the minority class using SQL before training.
AnswerA

AutoML Tables automatically applies class weighting and other built-in techniques to address imbalanced datasets, so no manual resampling is needed. This suits the company's limited ML expertise while handling the 1% failure rate effectively during training.

Why this answer

AutoML Tables has built-in techniques to handle class imbalance, such as automatically adjusting class weights and using stratified sampling during training. This allows the model to learn from the minority class without requiring manual data preprocessing, making it the most effective and simplest approach for users with limited ML expertise.

Exam trap

The trap here is that candidates may assume manual resampling (downsampling or oversampling) is always required for imbalanced datasets, but AutoML Tables abstracts this complexity, and the exam tests whether you trust its built-in capabilities for low-code solutions.

How to eliminate wrong answers

Option B is wrong because downsampling the majority class would discard valuable data, potentially reducing model performance and losing information about normal operating conditions. Option C is wrong because AutoML Tables does not expose a custom loss function configuration; it abstracts away such hyperparameters and uses its own optimized training pipeline. Option D is wrong because oversampling the minority class using SQL before training is unnecessary and could lead to overfitting or data leakage; AutoML Tables handles imbalance internally without manual intervention.

577
MCQhard

A media streaming company uses a recommendation model deployed on Vertex AI Endpoints. The model predicts whether a user will click on a recommended item. They have set up Vertex AI Model Monitoring with a training dataset and configured drift detection for both features and predictions. Recently, they observed that the prediction drift metric (Jensen-Shannon divergence) for the 'click' prediction has exceeded the threshold, but feature drifts are within normal ranges. What is the most likely cause of this prediction drift?

A.A change in the relationship between features and the target variable (concept drift) that is not captured by feature drift alone.
B.A bug in the Vertex AI Model Monitoring configuration that incorrectly computes the prediction drift metric.
C.An increase in the overall volume of prediction requests, leading to a higher variance in the prediction distribution.
D.A data pipeline error that has corrupted the feature values, causing the model to produce different predictions.
AnswerA

Prediction drift can occur even when feature distributions remain stable if the underlying relationship between features and the target changes. This is known as concept drift. For example, user preferences may shift such that the same feature values now lead to different click behavior. Monitoring prediction distribution helps detect such changes, which may require model retraining or adaptation.

Why this answer

Prediction drift with stable feature distributions indicates concept drift, where the relationship between features and the target has changed. This is common in dynamic environments like media streaming, where user preferences evolve. Detecting this early allows for timely model updates.

Exam trap

The trap here is to assume that prediction drift must be accompanied by feature drift, overlooking concept drift as a distinct phenomenon.

578
MCQhard

A logistics company has a Vertex AI Model Monitoring job that detects feature drift on a deployed route optimization model. The model uses 50 features, and the monitoring job is configured to monitor all features. The MLOps team notices that the monitoring job is incurring high costs and taking a long time to complete. They want to reduce monitoring overhead while still detecting significant drift on the most important features. What should they do?

A.Select a subset of features to monitor based on feature importance or business relevance.
B.Disable drift detection for numerical features and only monitor categorical features.
C.Increase the sampling rate to reduce the volume of prediction requests analyzed.
D.Reduce the monitoring frequency from daily to weekly.
AnswerA

Monitoring fewer features reduces the computational load and cost per run. By focusing on the most important features, the team can still detect significant drift while lowering overhead. Vertex AI Model Monitoring allows you to specify which features to monitor. This targeted approach is the most effective way to reduce cost and runtime without sacrificing detection of critical drift.

Why this answer

Monitoring all 50 features increases cost and runtime. To reduce overhead while still detecting significant drift, the team should select a subset of features to monitor based on importance or business relevance. This reduces the computational load per run.

Reducing frequency lowers total runs but not per-run cost. Increasing sampling rate would increase cost. Disabling numerical features arbitrarily could miss critical drift.

Exam trap

The trap here is focusing on monitoring frequency or sampling rate as the primary cost driver, when the number of monitored features is often the dominant factor in per-run cost and runtime.

579
MCQhard

An engineer wants to use BigQuery ML to explain predictions from a trained boosted tree classifier for a specific set of input rows. Which function should they use?

A.ML.EVALUATE
B.ML.FEATURE_IMPORTANCE
C.ML.PREDICT
D.ML.EXPLAIN_PREDICT
AnswerD

ML.EXPLAIN_PREDICT returns predictions alongside feature attributions for the supplied rows, using explainable AI methods such as Shapley values or integrated gradients, which is exactly the per-row explanation the engineer needs from the boosted tree classifier.

Why this answer

ML.EXPLAIN_PREDICT is the BigQuery ML function specifically designed to explain predictions from a trained model for a given set of input rows. It returns the prediction along with feature attributions (e.g., Shapley values) that show how each feature contributed to the prediction. This is the correct function for interpreting individual predictions from a boosted tree classifier.

Exam trap

PMLE often tests the confusion between global explainability (ML.FEATURE_IMPORTANCE, ML.GLOBAL_EXPLAIN) and local explainability (ML.EXPLAIN_PREDICT) — candidates must know which function provides per-row feature attributions.

How to eliminate wrong answers

Option A is wrong because ML.EVALUATE computes overall model metrics (accuracy, precision, recall) on a dataset, not explanations for individual predictions. Option B is wrong because ML.FEATURE_IMPORTANCE returns global feature importance for the model, not per-row explanations. Option C is wrong because ML.PREDICT returns predictions without explanations; it does not provide feature attributions.

580
Multi-Selectmedium

You are designing a distributed training job for a very large neural network that does not fit on a single machine. You need to split the model across multiple devices. Which TWO techniques can you use?

Select 2 answers
A.ParameterServerStrategy
B.Pipeline parallelism
C.Operator-level model parallelism
D.Data parallelism with MirroredStrategy
E.MultiWorkerMirroredStrategy
AnswersB, C

Pipeline parallelism divides the model into sequential stages placed on different devices, with micro-batches flowing between them. This splits the model across devices, satisfying the stem's constraint that the neural network does not fit on a single machine.

Why this answer

Pipeline parallelism (B) is correct because it splits the model's layers into sequential stages placed on different devices, so a model too large for one machine can be distributed across devices while micro-batches flow through the pipeline. Operator-level model parallelism (C) is also correct because it partitions individual operations (e.g., splitting a large matrix multiplication or a layer's weights) across devices, which directly addresses a model that does not fit on a single machine. ParameterServerStrategy (A) is a data-parallel approach that replicates the full model on each worker and only shards the parameter updates, so it does not solve the memory problem of a model too large for one device.

Data parallelism with MirroredStrategy (D) likewise replicates the entire model on every device, which is infeasible when the model does not fit on one machine. MultiWorkerMirroredStrategy (E) is also a data-parallel strategy that requires the full model to fit on each worker, so it does not satisfy the requirement.

Exam trap

The trap is confusing data parallelism with model parallelism; candidates may pick ParameterServerStrategy or MirroredStrategy because they are familiar distributed training strategies, but the exam expects recognition that only pipeline and operator-level parallelism split the model across devices.

581
MCQeasy

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline must run on a schedule every day at 2:00 AM UTC. The engineer wants to use a fully managed Google Cloud service to trigger the pipeline. Which service should the engineer use?

A.Cloud Composer
B.Cloud Scheduler
C.Cloud Functions
D.Cloud Tasks
AnswerB

Cloud Scheduler is a fully managed cron job service that can trigger Vertex AI Pipelines via HTTP requests to the Vertex AI API. It supports cron expressions and time zones, making it ideal for scheduled pipeline runs. The engineer can create a job that invokes the pipeline at 2:00 AM UTC daily without managing any infrastructure.

Why this answer

Cloud Scheduler is the fully managed service for cron-based scheduling in Google Cloud. It can directly invoke the Vertex AI Pipelines API endpoint on a schedule, such as daily at 2:00 AM UTC. Other services either lack native scheduling or require additional components to achieve the same result.

Exam trap

The trap here is assuming that Cloud Composer is the default scheduler for Vertex AI Pipelines, when Cloud Scheduler is the simpler, fully managed option for basic cron triggers.

582
MCQhard

Your model serving endpoint on Vertex AI is experiencing increased memory usage after a recent update. The model was converted from TensorFlow to TF Lite for faster inference. You notice that the endpoint's instances occasionally get killed due to out-of-memory (OOM) errors. What is the most likely cause?

A.The TF Lite model is larger in size than the original model.
B.The Vertex AI endpoint is not configured with enough CPU.
C.The number of inference threads in the TF Lite runtime is set too high, causing memory consumption.
D.The traffic to the endpoint has increased significantly.
AnswerC

TF Lite's interpreter allocates per-thread memory arenas, so a high inference thread count multiplies working memory across concurrent threads. This satisfies the stem's OOM scenario: the converted model's runtime configuration, not the model size alone, drives the increased memory consumption killing instances.

Why this answer

The most likely cause is that the number of inference threads in the TF Lite runtime is set too high. TF Lite uses multiple threads for inference, and each thread consumes additional memory for its own stack and intermediate tensors. If the thread count is set too high relative to the available memory, it can lead to out-of-memory errors.

This is a common configuration issue when moving to TF Lite.

Exam trap

The trap is assuming that increased traffic or model size is the cause, but the question highlights the recent conversion to TF Lite, pointing to a configuration issue like thread count, which is a known memory multiplier.

How to eliminate wrong answers

Option A is wrong because TF Lite models are typically smaller than the original TensorFlow models due to quantization and optimization, so size increase is unlikely. Option B is wrong because insufficient CPU would cause slow inference, not OOM errors; memory is the constraint. Option D is wrong because increased traffic would cause more concurrent requests, but the question specifies the issue started after the update to TF Lite, and OOM is more directly related to per-instance memory configuration.

583
MCQmedium

You are deploying a scikit-learn model to a Vertex AI endpoint for real-time inference. Prediction requests arrive as JSON payloads containing a single instance per request, and the model's predict method expects a pandas DataFrame with named columns. You want to avoid writing a custom container. Which approach should you take?

A.Upload the model artifact with a pre-built Vertex AI XGBoost container, because that container internally wraps every scikit-learn estimator and feeds it a DataFrame.
B.Upload the model artifact with a pre-built Vertex AI scikit-learn container and supply a custom prediction routine (predictor.py) that converts the decoded JSON dict into a pandas DataFrame before calling predict.
C.Upload the model artifact with a pre-built TensorFlow container and wrap the scikit-learn estimator in a tf.function so the container's serving signature handles DataFrame creation.
D.Upload the model artifact with a pre-built Vertex AI scikit-learn container and rely on its default Predictor, which automatically converts JSON objects into a pandas DataFrame with matching column names.
AnswerB

A custom prediction routine packaged as a Python source distribution lets you subclass the pre-built scikit-learn Predictor and override preprocess to build a DataFrame from the decoded instance. This keeps the managed runtime image while satisfying the model's DataFrame input contract, so no custom container or Dockerfile is needed.

Why this answer

A custom prediction routine gives you a managed pre-built runtime plus a Python package that overrides preprocessing, which is the supported way to adapt request payloads to a model's expected input format. Because the estimator requires named DataFrame columns, the override must build that DataFrame explicitly; the default Predictor performs no such conversion and other pre-built containers cannot load the estimator at all.

Exam trap

The trap here is assuming the pre-built scikit-learn container automatically converts JSON payloads into a named-column DataFrame for every estimator.

584
Multi-Selectmedium

You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?

Select 2 answers
A.Increase the number of model replicas to the maximum.
B.Use TensorRT to quantize the model to FP16 or INT8.
C.Disable model caching to reduce memory usage.
D.Enable dynamic batching in Triton to aggregate requests.
E.Use a larger machine type with more vCPUs.
AnswersB, D

TensorRT applies FP16 or INT8 quantisation, reducing weight precision and memory bandwidth while exploiting GPU tensor cores for faster matrix maths. This directly satisfies the stem's inference-performance constraint on Triton, where lower-precision kernels cut latency and raise throughput versus FP32 execution.

Why this answer

Option B is correct because TensorRT can quantize a model to FP16 or INT8, reducing precision and memory bandwidth requirements while leveraging NVIDIA GPU tensor cores, which directly lowers inference latency and increases throughput on Triton. Option D is correct because Triton's dynamic batching aggregates multiple incoming inference requests into a single batch at runtime, improving GPU utilization and throughput without requiring client-side batching. Option A is not appropriate because simply maxing out replicas can waste resources and does not guarantee better performance if the model is not compute-bound or if the machine lacks capacity.

Option C is wrong because disabling model caching forces reloads and increases latency rather than improving inference performance. Option E is not ideal because adding vCPUs does not help GPU-bound inference and may not improve Triton throughput.

Exam trap

Google often tests the misconception that simply adding more replicas or CPU resources will linearly improve inference performance, ignoring the GPU-bound nature of model serving and the importance of batching and precision optimization.

585
MCQeasy

You are defining a Python function component in KFP SDK v2. Which decorator should you use?

A.@dsl.task
B.@component
C.@dsl.pipeline
D.@dsl.component
AnswerD

@dsl.component converts a plain Python function into a lightweight KFP v2 component, automatically inferring its inputs, outputs and container image. This satisfies the requirement to define a function-based component without authoring a separate component YAML specification.

Why this answer

In KFP SDK v2, the `@dsl.component` decorator is used to define a Python function as a lightweight, reusable pipeline component that can be executed independently. This decorator automatically generates a containerized component from the function's signature and type annotations, enabling type-safe inputs and outputs without requiring a separate component YAML specification.

Exam trap

The exam often tests the distinction between v1 and v2 decorators, so the trap here is that candidates familiar with KFP SDK v1 may incorrectly choose `@component` (option B) instead of the v2-specific `@dsl.component`.

How to eliminate wrong answers

Option A is wrong because `@dsl.task` is not a valid decorator in KFP SDK v2; tasks are implicitly created when a component is called within a pipeline, not via a decorator. Option B is wrong because `@component` is the decorator from KFP SDK v1 (the older `kfp.components` module) and is not used in v2, which requires the `dsl` namespace. Option C is wrong because `@dsl.pipeline` is used to define a pipeline (a DAG of components), not a single component function.

586
MCQeasy

A retail company has a model deployed on a Vertex AI Endpoint that predicts customer lifetime value. The ML team wants to monitor the model for feature drift without setting up a full Vertex AI Model Monitoring job. They need a lightweight solution that compares live prediction requests to a reference distribution and sends alerts when drift exceeds a threshold. What should they do?

A.Use Cloud Audit Logs to track prediction requests and set up log-based alerts.
B.Log prediction requests to BigQuery and use a scheduled query to compute drift metrics, then alert via Cloud Monitoring.
C.Enable request-response logging on the endpoint and manually inspect logs daily.
D.Use Vertex AI Model Monitoring with default settings and a sampling rate of 1.0.
AnswerB

Logging requests to BigQuery and running scheduled queries to compute drift metrics is a lightweight, serverless approach. It avoids the overhead of managing a Vertex AI Model Monitoring job. Cloud Monitoring can trigger alerts when computed drift exceeds a threshold. This meets the need for a simple solution that compares live data to a reference distribution and sends alerts, without full monitoring infrastructure.

Why this answer

A lightweight drift monitoring solution can be built by logging prediction requests to BigQuery, then using scheduled queries to compute drift metrics against a reference distribution. Cloud Monitoring can alert when drift exceeds a threshold. This avoids the complexity of a full Vertex AI Model Monitoring job while still providing automated alerts.

Manual inspection or using audit logs does not meet the need for automated, data-based drift detection.

Exam trap

The trap here is assuming that Vertex AI Model Monitoring is the only way to detect drift, when a custom lightweight pipeline using BigQuery and Cloud Monitoring can satisfy the requirement without a full monitoring job.

587
MCQmedium

An engineer is training a model on Vertex AI using a custom container. The training job fails with an error indicating that the container exited with a non-zero status. The engineer wants to debug the issue. What is the best way to access the logs?

A.SSH into the training container using Vertex AI's SSH feature
B.View logs in Cloud Storage under the job's output directory
C.Use Cloud Debugger to inspect the container
D.Check the logs in Cloud Logging (Logs Explorer)
AnswerD

Vertex AI writes custom container training output, including stderr and the non-zero exit trace, to Cloud Logging. Logs Explorer surfaces those entries for the failed job, satisfying the debugging requirement by exposing the container's actual error rather than only the job's status.

Why this answer

Vertex AI automatically streams all container stdout and stderr to Cloud Logging (Logs Explorer). When a custom container exits with a non-zero status, the detailed error messages, stack traces, and application logs are captured there, making it the primary and most comprehensive debugging tool. Cloud Logging provides structured, searchable logs without requiring direct access to the container.

Exam trap

A common trap is the misconception that you can SSH into a training container or that logs are stored in Cloud Storage, when in fact Cloud Logging is the centralized, default logging solution for all Vertex AI training jobs.

How to eliminate wrong answers

Option A is wrong because Vertex AI does not provide an SSH feature for training containers; training jobs run in ephemeral, isolated environments with no interactive shell access. Option B is wrong because Cloud Storage under the job's output directory stores artifacts like model checkpoints and metrics, not real-time container logs or error messages. Option C is wrong because Cloud Debugger is designed for debugging running applications in production by capturing snapshots and variable states, not for inspecting container exit errors or retrieving logs from a failed training job.

588
MCQhard

Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?

A.Set a higher minReplicaCount so the endpoint always keeps enough warm capacity for peak traffic, and combine it with a lower autoscaling metric target to trigger scaling earlier.
B.Increase the endpoint's maxReplicaCount and rely on the autoscaler's default metrics to add capacity as quickly as possible.
C.Enable request logging on the endpoint so Cloud Logging captures the 429 responses and the autoscaler can use log volume as an additional scaling signal.
D.Deploy the model to a second endpoint and split traffic between the two endpoints so each handles roughly half of the incoming requests.
AnswerA

Keeping warm replicas sized for the burst removes the cold-start gap entirely, and lowering the autoscaling metric target makes the autoscaler add replicas before saturation rather than after. Together they absorb the sudden spike immediately while still allowing scale-down when traffic subsides, which is exactly what is needed here.

Why this answer

Bursts that outpace autoscaling reaction time must be met with pre-provisioned warm capacity. Setting minReplicaCount near peak demand guarantees replicas are already serving when the spike lands, and a lower metric target causes additional replicas to be requested earlier in the ramp. Merely changing the maximum, logging, or duplicating endpoints does not remove the delay between demand and available capacity.

Exam trap

The trap here is believing that raising maxReplicaCount makes autoscaling react fast enough to absorb a sudden burst.

589
Drag & Dropmedium

Drag and drop the steps to set up model monitoring for drift detection on Vertex AI in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order to set up model monitoring for drift detection on Vertex AI is: 1. Deploy model to an endpoint. 2. Enable monitoring on the deployed endpoint. 3.

Set drift detection thresholds (e.g., for feature distributions or prediction outputs). 4. Configure alerts (e.g., email notifications). 5. Review monitoring results and alerts.

590
MCQmedium

A team develops a pipeline that trains a model and evaluates it. They want to pass the test accuracy (a float) from the evaluation component to a subsequent deployment component. Which KFP SDK type should the evaluation component output be annotated with?

A.Output[float]
B.Output[Metrics]
C.Output[Artifact]
D.Output[ClassificationMetrics]
AnswerA

Output[float] annotates the component's return value as a scalar float, which is exactly the test accuracy type the deployment component consumes. KFP serialises this primitive and passes it as a typed input, avoiding string parsing or artifact handling for a simple numeric metric.

Why this answer

In the KFP SDK, a component output annotated as Output[float] is treated as a lightweight scalar parameter that is passed directly between components via the pipeline's execution graph. Since test accuracy is a single floating-point value (not a file, dataset, or structured metric object), Output[float] is the correct type to declare so the downstream deployment component can consume it as a typed input. KFP serializes this primitive and makes it available for parameter passing without needing artifact storage.

Exam trap

PMLE often tests the distinction between KFP parameter types (Output[float], Output[int], Output[str]) and artifact types (Output[Metrics], Output[Artifact], Output[ClassificationMetrics]), catching candidates who assume any numeric output should be a Metrics artifact.

How to eliminate wrong answers

Option B is wrong because Output[Metrics] is used to emit a dictionary of named scalar metrics for visualization in the KFP UI, not to pass a single value as a typed parameter to a downstream component. Option C is wrong because Output[Artifact] declares a file- or directory-based artifact (e.g., a model, dataset, or plot) that is stored in the artifact repository, which is overkill and semantically incorrect for a single float. Option D is wrong because Output[ClassificationMetrics] is a specialized artifact type for confusion matrices, ROC curves, and calibration plots, not for passing a scalar accuracy value between components.

591
Drag & Dropmedium

Drag and drop the steps to perform a hyperparameter tuning job on Vertex AI in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Define the search space, then create and run the tuning job, monitor, and select the best parameters.

592
MCQhard

An ML engineer trained a model and registered it in Vertex AI Model Registry. They want to assign the alias 'champion' to the best-performing version for production deployment. Which gcloud command should they use?

A.gcloud ai models versions describe --model=MODEL_ID --version=VERSION_ID
B.gcloud ai models upload --model-id=MODEL_ID --display-name=champion
C.gcloud ai endpoints deploy-model --model=MODEL_ID --alias=champion
D.gcloud ai models versions update --model=MODEL_ID --version=VERSION_ID --update-aliases=champion
AnswerD

The versions update subcommand with --update-aliases assigns the 'champion' alias to a specific registered model version, enabling production deployment by alias rather than version number. Other commands set IAM policy or labels, which do not create the deployment alias the scenario requires.

Why this answer

The correct command is 'gcloud ai models versions update' with the --update-aliases flag. This command updates a specific model version's aliases, allowing you to assign or remove aliases like 'champion' to denote the best-performing version for production. Option A only describes a version, option B uploads a new model, and option C deploys a model to an endpoint, none of which assign aliases to a version.

593
Multi-Selecteasy

A company wants to use Vertex AI for hyperparameter tuning. Which three components are required to configure a hyperparameter tuning job? (Choose THREE.)

Select 3 answers
A.Algorithm (e.g., Bayesian, grid, random)
B.Machine type for each trial
C.List of hyperparameters with types and ranges
D.Objective metric name and goal (minimize or maximize)
E.Training container image
AnswersA, C, D

Required to specify how to search.

Why this answer

Vertex AI hyperparameter tuning requires specifying the search algorithm (Bayesian, grid, or random) to determine how the hyperparameter space is explored. Bayesian optimization is the default and most efficient for continuous spaces, while grid search is exhaustive and random search is simple. Without an algorithm, Vertex AI cannot decide how to sample trials.

Exam trap

Candidates often mistakenly include machine type or container image as mandatory tuning parameters, but these are optional training job settings.

594
MCQmedium

An ML engineer is using Vertex AI Training to fine-tune a large image classification model on a dataset stored in Cloud Storage. The training job uses a custom container and runs on a single NVIDIA V100 GPU. The engineer notices that GPU utilization is consistently low (around 20%) and training is slow. The data is stored as many small JPEG files. What should the engineer do to improve GPU utilization and training speed?

A.Enable Vertex AI Vizier to automatically tune hyperparameters and improve training speed.
B.Use a larger GPU instance with more memory, such as an NVIDIA A100.
C.Convert the dataset to TFRecord format and use the tf.data API with parallel interleaving and prefetching.
D.Increase the batch size to the maximum that fits in GPU memory.
AnswerC

TFRecord is a binary format that stores data efficiently and allows sequential reads, which is much faster than reading many small JPEG files individually. Using tf.data with interleave, map with num_parallel_calls, and prefetch overlaps data loading and preprocessing with GPU computation. This directly addresses the input pipeline bottleneck, leading to higher GPU utilization and faster training.

Why this answer

Low GPU utilization during training often indicates that the GPU is starved for data. When data is stored as many small files, the I/O overhead dominates. Converting to TFRecord, a binary format optimized for TensorFlow, and using tf.data with parallel reads and prefetching can dramatically speed up data loading.

This ensures the GPU receives data as fast as it can process it, improving utilization and reducing training time.

Exam trap

The trap here is attributing low GPU utilization to insufficient GPU power or needing hyperparameter tuning, rather than diagnosing the input pipeline bottleneck.

595
MCQmedium

You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?

A.Export the model as a serialized TFX pipeline.
B.Export the model to a SavedModel and deploy it using Vertex AI Edge Manager.
C.Convert the model to TensorFlow Lite using the TensorFlow Lite converter.
D.Use Vertex AI Model Optimization to compile the model for edge devices.
AnswerC

TensorFlow Lite is Google's runtime for on-device inference, and the TensorFlow Lite converter transforms a trained TensorFlow SavedModel into the compact .tflite format with optimisations such as quantisation. This suits edge devices with limited compute and memory.

Why this answer

TensorFlow Lite is specifically designed for on-device inference on edge devices, offering optimized performance and reduced model size. The TensorFlow Lite converter transforms a TensorFlow model (e.g., from a SavedModel) into the FlatBuffer format (.tflite), which is lightweight and compatible with mobile and embedded platforms. This directly addresses the requirement for real-time inference on edge devices.

Exam trap

A common misconception is that Vertex AI services like Edge Manager or Model Optimization directly produce a deployable edge format, when in fact TensorFlow Lite conversion is the required final step for on-device inference.

How to eliminate wrong answers

Option A is wrong because a TFX pipeline is a production ML workflow framework for orchestrating training, validation, and deployment, not a model format for on-device inference; serializing it does not produce a deployable edge model. Option B is wrong because Vertex AI Edge Manager is a service for managing and deploying models to edge devices, but it expects models in a compatible format like TensorFlow Lite; exporting to SavedModel alone is insufficient without conversion, and Edge Manager itself does not perform the conversion. Option D is wrong because Vertex AI Model Optimization focuses on techniques like pruning and quantization to improve model efficiency, but it does not compile the model into a format suitable for on-device inference; the output still requires conversion to TensorFlow Lite for edge deployment.

596
Multi-Selecthard

A company is building a document processing pipeline for invoices. They need to extract key fields (invoice number, date, total amount) and allow human review for invoices over $10,000. Which TWO Google Cloud services/features should they combine?

Select 2 answers
A.Cloud Vision API for OCR
B.Human-in-the-Loop (HITL) on Document AI
C.AutoML Tables to predict missing fields
D.Document AI with invoice parser processor
E.Cloud Translation API to translate invoices
AnswersB, D

Why this answer

Option D is correct because Document AI's specialized Invoice parser processor is purpose-built to extract structured fields such as invoice number, invoice date, and total amount from invoice documents, which is exactly the extraction requirement in this pipeline. Option B is correct because Document AI's Human-in-the-Loop (HITL) feature lets you define confidence thresholds and route documents for manual review, which directly satisfies the requirement to have humans review invoices over $10,000 before downstream processing. Together, the Invoice parser handles automated field extraction while HITL provides the human verification step for high-value invoices.

Option A is not the best fit because Cloud Vision API provides generic OCR without invoice-specific field extraction or built-in human review workflows. Option C is incorrect because AutoML Tables is a tabular ML service for structured data prediction, not for predicting missing fields in parsed documents. Option E is incorrect because Cloud Translation API only translates text and does not extract invoice fields or support human review.

Exam trap

PMLE often tests whether candidates pick generic OCR (Cloud Vision) instead of the domain-specific Document AI parser, and whether they recognize HITL as the mechanism for human review rather than building custom review logic.

597
MCQeasy

You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?

A.Use Cloud Run to serve each model separately.
B.Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.
C.Package all models into a single container and deploy that container.
D.Deploy each model to its own endpoint and use a load balancer.
AnswerB

Multiple models can be deployed to a single endpoint, each receiving a portion of the traffic.

Why this answer

Vertex AI Prediction supports multi-model serving, allowing you to deploy multiple models to a single endpoint and use traffic splits to route a percentage of requests to each model. This reduces costs by sharing underlying infrastructure (e.g., compute resources) across models, rather than provisioning separate endpoints or containers for each model.

Exam trap

The trap here is that candidates often confuse multi-model serving with containerization, assuming that bundling models into a single container (Option C) is equivalent to Vertex AI's native multi-model support, but this ignores the need for traffic splitting and independent model lifecycle management.

How to eliminate wrong answers

Option A is wrong because Cloud Run serves each model as a separate service, which does not consolidate models onto a single endpoint and incurs additional costs for individual scaling and networking. Option C is wrong because packaging all models into a single container violates the principle of model isolation, complicates updates, and does not leverage Vertex AI's native traffic-splitting mechanism for granular control. Option D is wrong because deploying each model to its own endpoint and using a load balancer increases operational overhead and cost, as each endpoint requires separate compute resources, defeating the purpose of cost reduction.

598
MCQeasy

A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?

A.n1-standard-4 with 1 NVIDIA Tesla T4
B.n1-standard-4
C.e2-standard-4
D.n1-highmem-8
AnswerA

NVIDIA Tesla T4 GPUs attached to n1-standard-4 instances provide the GPU acceleration the stem demands, while n1-standard-4 supplies sufficient vCPU and memory for low-latency online inference of a TensorFlow model. Vertex AI supports this accelerator pairing directly, satisfying both the GPU and latency constraints when deploying the endpoint.

Why this answer

The n1-standard-4 machine type supports attaching GPUs such as the NVIDIA Tesla T4, which provides GPU acceleration for low-latency online predictions. Vertex AI endpoints require a machine type that allows GPU attachment, and the n1-series is one of the few families that supports GPUs, while the T4 offers a good balance of cost and performance for inference workloads.

Exam trap

The trap here is that candidates may assume any machine type can be paired with a GPU, but only specific series (like n1, n2, g2) support GPU attachment, and the e2 series explicitly does not, leading to a wrong selection if the GPU requirement is overlooked.

How to eliminate wrong answers

Option B is wrong because n1-standard-4 without a GPU does not provide GPU acceleration, so it cannot meet the requirement for low-latency predictions with GPU. Option C is wrong because e2-standard-4 does not support attaching GPUs at all; the e2 series is designed for cost-optimized CPU-only workloads. Option D is wrong because n1-highmem-8, while it can support GPUs, is over-provisioned in memory for typical inference tasks and does not include a GPU by default, so it would not satisfy the explicit need for GPU acceleration unless a GPU is attached, but the option as stated lacks the GPU specification.

599
MCQeasy

A company is serving a model for their e-commerce website. They expect traffic to be low at night and very high during flash sales. They want to minimize costs while ensuring availability during spikes. Which autoscaling configuration should they use?

A.min_replica_count=5, max_replica_count=5, target_cpu=60
B.min_replica_count=1, max_replica_count=20, target_cpu=60
C.min_replica_count=10, max_replica_count=10, target_cpu=60
D.min_replica_count=0, max_replica_count=100, target_cpu=80
AnswerB

Setting min_replica_count=1 keeps costs low during quiet nights, while max_replica_count=20 allows horizontal scaling to absorb flash-sale spikes. The target_cpu=60 metric triggers additional replicas before saturation, preserving availability. This balances the stem's dual constraint: minimising idle cost while guaranteeing capacity during unpredictable demand surges.

Why this answer

Setting a high max_replica_count allows scaling to handle spikes, while a low min_replica_count saves cost during low traffic. CPU utilization target of 60% is reasonable.

600
MCQeasy

A company has a prototype ML model that works well on historical data, but when deployed to production, the model performance degrades over time. The data distribution shifts gradually. Which strategy should they implement to maintain model accuracy?

A.Increase the regularization strength to prevent overfitting.
B.Increase the amount of training data by using more historical records.
C.Implement a retraining pipeline that periodically retrains the model on recent data.
D.Switch to a more complex model architecture to better capture patterns.
AnswerC

Gradual data distribution shift causes the static model's learned patterns to diverge from production inputs, degrading accuracy. A periodic retraining pipeline ingests recent data, updating model parameters to track the shifted distribution and sustain predictive performance over time.

Why this answer

Gradual data distribution shifts (concept drift) require the model to adapt to new patterns over time. A retraining pipeline that periodically retrains on recent data ensures the model remains aligned with the current production distribution, directly addressing the degradation caused by drift without relying on static historical data.

Exam trap

Google Cloud often tests the misconception that overfitting or model complexity is the primary cause of production degradation, leading candidates to choose regularization or more complex architectures instead of recognizing that distribution shift requires data freshness.

How to eliminate wrong answers

Option A is wrong because increasing regularization strength reduces overfitting to historical noise but does not address the root cause—distribution shift—and may actually harm performance on new data by forcing the model to ignore legitimate new patterns. Option B is wrong because adding more historical records only reinforces the old distribution, making the model less responsive to recent shifts and potentially worsening drift. Option D is wrong because switching to a more complex model architecture increases capacity to fit data but does not solve the problem of stale training distribution; it may even overfit to outdated patterns and degrade faster under drift.

Page 7

Page 8 of 11

Page 9

All pages