Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 376–450

775 questions total · 11pages · All types, answers revealed

Page 5

Page 6 of 11

Page 7
376
MCQeasy

A data engineer wants to orchestrate a complex workflow that includes running a Vertex AI pipeline, then a BigQuery job, and finally a Dataflow pipeline. The workflow must handle dependencies, retries, and monitoring. Which Google Cloud service is most suitable for this orchestration?

A.Cloud Tasks
B.Cloud Composer
C.Cloud Scheduler
D.Workflows
AnswerB

Cloud Composer is managed Apache Airflow, whose DAGs natively orchestrate heterogeneous tasks across Vertex AI, BigQuery and Dataflow with dependency handling, retries and monitoring. This satisfies the stem's requirement for cross-service workflow orchestration with dependencies and retries.

Why this answer

Cloud Composer (based on Apache Airflow) is the most suitable service for orchestrating a complex workflow with dependencies, retries, and monitoring across Vertex AI, BigQuery, and Dataflow. It provides a managed Airflow environment that natively supports DAG-based orchestration, built-in retry logic, and integration with Google Cloud services via operators like VertexAIPipelineOperator, BigQueryOperator, and DataflowTemplatedJobStartOperator.

Exam trap

A common misconception is that Workflows is sufficient for complex ML orchestration, but it lacks the built-in operator integrations and retry semantics that Cloud Composer provides for multi-service pipelines.

How to eliminate wrong answers

Option A is wrong because Cloud Tasks is a distributed task queue for executing discrete, short-lived tasks with HTTP endpoints, not for orchestrating multi-step workflows with complex dependencies and retries across different services. Option C is wrong because Cloud Scheduler is a cron-based job scheduler that triggers single events at specified times, lacking the ability to manage dependencies between multiple pipeline stages or handle retries. Option D is wrong because Workflows is a low-code orchestration service for sequential or parallel steps, but it does not natively support the rich operator ecosystem, retry policies, or monitoring capabilities that Cloud Composer provides for ML pipelines involving Vertex AI, BigQuery, and Dataflow.

377
MCQmedium

An ML engineer needs to run batch predictions on 10 TB of data stored in BigQuery using a TensorFlow model. The predictions must be written to BigQuery. Which service should they use?

A.Create a Dataflow pipeline to read from BigQuery, run the model using Python, and write results to BigQuery.
B.Export BigQuery data to GCS, run batch prediction on GCS, then load results back to BigQuery.
C.Use Vertex AI online prediction with batch requests.
D.Use Vertex AI Batch Prediction with BigQuery source and sink.
AnswerD

Batch Prediction reads directly from BigQuery as the input source and writes predictions back to a BigQuery sink, avoiding 10 TB of manual export. This natively satisfies both the large-scale source and the required output destination.

Why this answer

Vertex AI Batch Prediction natively supports BigQuery as both input source and output sink, so the engineer can submit a batch job directly against the BigQuery table and have predictions written back to BigQuery without any data movement code. This is the most efficient and least error-prone option for 10 TB of data.

Exam trap

PMLE often tests whether candidates know Vertex AI Batch Prediction supports BigQuery natively — many pick the GCS export/import route or a custom Dataflow pipeline because they assume data must be staged in GCS first, which is not required.

How to eliminate wrong answers

Option A is wrong because building a custom Dataflow pipeline to load the TensorFlow model and run inference is unnecessary engineering overhead when Vertex AI Batch Prediction already handles BigQuery I/O natively, and Dataflow is not optimized for GPU/TPU model inference. Option B is wrong because exporting 10 TB to GCS, running batch prediction, then reloading to BigQuery adds significant latency, storage cost, and operational complexity compared to native BigQuery source/sink. Option C is wrong because Vertex AI online prediction is designed for low-latency, small-payload requests and has strict request size and QPS limits — it cannot handle 10 TB batch workloads efficiently.

378
MCQmedium

A team uses Vertex AI Workbench managed notebooks. They want to version control their notebook files and collaborate using Git. What is the best way to integrate Git?

A.Use Cloud Source Repositories only
B.Use the built-in Git integration in Vertex AI Workbench managed notebooks
C.Use gcloud source repos clone inside the terminal
D.Manually download notebooks and upload to GitHub via browser
AnswerB

Managed notebooks include native Git integration, exposing repository cloning, credential handling and commit or push controls through the interface. This satisfies version control and collaboration requirements without manual command-line setup, unlike unmanaged instances where git must be configured by hand.

Why this answer

Vertex AI Workbench managed notebooks include native Git integration through the JupyterLab Git extension, allowing users to clone repositories, commit, push, and pull directly from the notebook UI without leaving the environment. This is the officially supported and most seamless method for version-controlling notebooks in managed notebooks, supporting GitHub, Cloud Source Repositories, and any Git-compatible remote.

Exam trap

The trap here is confusing Cloud Storage auto-persistence of the notebook VM with actual Git version control, leading candidates to pick Cloud Source Repositories or manual uploads instead of the built-in integration.

How to eliminate wrong answers

Option A is wrong because Cloud Source Repositories is only one possible Git remote and is not required — the built-in integration supports any Git provider, so restricting to CSR is unnecessarily narrow. Option C is wrong because running 'gcloud source repos clone' in the terminal only works with Cloud Source Repositories and bypasses the integrated UI workflow, making it less flexible than the built-in Git integration. Option D is wrong because manually downloading and re-uploading notebooks via a browser is error-prone, breaks commit history, and defeats the purpose of proper version control.

379
MCQmedium

A company has a TensorFlow model for image classification that must run on edge devices with limited memory. They need to reduce the model size without significant accuracy loss. Which technique should they use?

A.Post-training quantization using TensorFlow Lite.
B.Knowledge distillation to train a smaller student model.
C.Pruning the model weights to zero out unimportant connections.
D.Use a larger VM for training.
AnswerA

Post-training quantization converts the trained model's float32 weights to 8-bit integers via TensorFlow Lite, shrinking size roughly fourfold and lowering memory use on constrained edge hardware, with minimal accuracy loss since no retraining is required.

Why this answer

Post-training quantization with TensorFlow Lite converts a trained model's weights from 32-bit floats to 8-bit integers, reducing model size by ~4x and speeding up inference on edge devices with minimal accuracy loss. TensorFlow Lite is purpose-built for edge deployment, making this the most direct and practical technique for the stated constraint.

Exam trap

PMLE often tests the distinction between model compression techniques — candidates confuse quantization (reduces precision) with pruning (removes weights) and distillation (trains a smaller model), picking the wrong one for the 'reduce size without retraining' constraint.

How to eliminate wrong answers

Option B is wrong because knowledge distillation requires training a smaller student model from scratch, which is more complex and time-consuming than post-training quantization, and it does not leverage the existing trained model directly. Option C is wrong because pruning alone reduces the number of weights but does not necessarily shrink the model file size dramatically unless combined with compression, and it often requires retraining to recover accuracy. Option D is wrong because using a larger VM for training does not reduce the deployed model size at all — it only affects training speed.

380
MCQeasy

You want to use Vertex AI JumpStart to quickly deploy a pre-built foundation model for text summarization. Which action is required?

A.Select the model from Model Garden and deploy it to a Vertex AI endpoint
B.Train the model from scratch using Vertex AI Training
C.Export the model to a Cloud Storage bucket and use batch prediction
D.Build a custom Docker container with the model and deploy to Vertex AI
AnswerA

JumpStart models are accessed through Model Garden, and deploying the selected foundation model to a Vertex AI endpoint is what provisions it for inference. This satisfies the stem's requirement to quickly deploy a pre-built summarisation model.

Why this answer

Vertex AI JumpStart provides pre-built foundation models in Model Garden that can be deployed directly to a Vertex AI endpoint with minimal configuration. Selecting the model from Model Garden and deploying it to an endpoint is the standard JumpStart workflow for getting a foundation model into production for tasks like text summarization. No training, containerization, or batch export is required because JumpStart handles the deployment plumbing.

Exam trap

PMLE often tests the distinction between JumpStart's one-click Model Garden deployment and the manual custom-container or training-from-scratch paths, so candidates overthink and pick the container option.

How to eliminate wrong answers

Option B is wrong because training from scratch is unnecessary and prohibitively expensive when JumpStart already offers pre-trained foundation models. Option C is wrong because exporting to Cloud Storage for batch prediction does not satisfy the goal of deploying a model for interactive inference and skips the JumpStart deployment path entirely. Option D is wrong because building a custom Docker container is the manual custom-container deployment route, not the JumpStart one-click deployment workflow.

381
MCQmedium

A company needs to perform sentiment analysis on streaming social media data. Which architecture should they use?

A.Dataflow → Pub/Sub → Natural Language API → BigQuery
B.Pub/Sub → Cloud Functions → Natural Language API → Cloud Storage
C.Cloud Functions → Pub/Sub → Natural Language API → BigQuery
D.Pub/Sub → Dataflow → Natural Language API → BigQuery
AnswerD

Pub/Sub ingests the continuous social media stream, Dataflow applies windowed processing for real-time sentiment scoring, and the Natural Language API performs the sentiment analysis itself. BigQuery then stores results for querying. This satisfies the streaming requirement, which batch pipelines such as Cloud Storage-triggered jobs cannot meet.

Why this answer

Streaming social media data requires a scalable, ordered ingestion pipeline. Pub/Sub ingests the stream, Dataflow processes it in real-time (e.g., windowing, deduplication), the Natural Language API performs sentiment analysis, and BigQuery stores results for querying. This decouples ingestion from processing and storage, enabling exactly-once semantics and auto-scaling.

Exam trap

Google Cloud often tests the misconception that Cloud Functions can replace Dataflow for streaming pipelines, but Cloud Functions lacks stream processing primitives (e.g., windowing, state management) and has a 9-minute timeout, making it unsuitable for continuous sentiment analysis.

How to eliminate wrong answers

Option A is wrong because Dataflow cannot directly read from a streaming source without a buffer like Pub/Sub; placing Dataflow before Pub/Sub reverses the pipeline order and breaks stream ingestion. Option B is wrong because Cloud Functions is not designed for high-throughput streaming; it has a 9-minute timeout and no built-in stream processing (e.g., windowing), making it unsuitable for continuous social media data. Option C is wrong because Cloud Functions should not be the entry point for streaming data; it lacks Pub/Sub's durability and ordering guarantees, and placing Pub/Sub after Cloud Functions would lose the stream before processing.

382
MCQeasy

When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?

A.Copy the dataset to each worker's local disk
B.Use NFS
C.Use Cloud Storage
D.Use Google Drive
AnswerC

Cloud Storage provides a shared, region-agnostic object store that every worker can read concurrently, satisfying the multi-worker distribution requirement without duplicating data. Vertex AI Training mounts or streams GCS paths directly, so each worker accesses the same dataset shards. Persistent disks or local storage would tie data to a single VM, breaking parallel access.

Why this answer

Vertex AI Training workers need shared, concurrent read access to the training dataset without manual replication. Cloud Storage (GCS) is the recommended and fully integrated solution because it provides a distributed, highly available object store that all workers can read from in parallel via the `tf.io.gfile` API or GCS connector, eliminating data duplication and ensuring consistency across the cluster.

Exam trap

The trap here is that candidates confuse 'shared storage' with 'local copies' or 'user-friendly sync tools,' assuming NFS or Drive are viable for distributed ML, when Vertex AI explicitly requires a cloud-native object store like GCS for scalability and fault tolerance.

How to eliminate wrong answers

Option A is wrong because copying the dataset to each worker's local disk introduces data duplication, increases startup latency, and risks inconsistency if workers are preempted or auto-scaled; Vertex AI does not manage local disk replication. Option B is wrong because NFS (Network File System) is not natively supported in Vertex AI Training; it would require manual setup of an NFS server, introduces a single point of failure, and adds network latency that GCS avoids with its native parallel read capabilities. Option D is wrong because Google Drive is a user-facing file sync service, not designed for high-throughput, concurrent access by distributed training jobs; it lacks the necessary IAM integration, access controls, and performance guarantees for ML workloads.

383
Multi-Selectmedium

An ML engineer is preparing to train a large model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket as a set of TFRecord files. The engineer wants to optimize the training job to reduce cost and improve performance. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Use a TPU or GPU accelerator and ensure the input pipeline is optimized to keep the accelerator busy.
B.Enable Vertex AI Model Monitoring on the training job to detect anomalies during training.
C.Store the training data in a Cloud Storage bucket in the same region as the training job.
D.Use the default Compute Engine service account with broad permissions to simplify access to Cloud Storage.
E.Increase the number of training steps to improve model accuracy, regardless of cost.
AnswersA, C

Accelerators like GPUs or TPUs can significantly speed up training for large models. However, they must be fed data efficiently; otherwise, utilization drops. Optimizing the input pipeline with tf.data, prefetching, and parallel reads ensures the accelerator is not idle. This combination reduces training time and cost, as you pay for the accelerator only while it is used effectively.

Why this answer

To optimize training cost and performance, the engineer should use accelerators and ensure the input pipeline is efficient, and store data in the same region as the training job. These actions reduce training time and avoid unnecessary network overhead. Other options like using broad service accounts or increasing steps do not contribute to efficiency and may introduce security or cost issues.

Exam trap

The trap here is thinking that using a default service account or enabling Model Monitoring on training jobs are valid optimizations, when they are not.

384
MCQmedium

A team uses Cloud Build to automatically trigger a Vertex AI pipeline when changes are pushed to the model code repository. They have a cloudbuild.yaml file that builds a container image and submits the pipeline. However, they want to run the pipeline only if the commit includes changes to the 'training/' directory. Which Cloud Build configuration option should be used to filter the trigger?

A.Add a 'ignoreFiles' field with 'training/**' to the trigger.
B.Use a 'substitutions' field with a regex pattern to filter commits.
C.Configure a Cloud Function to check the commit diff and call Cloud Build API conditionally.
D.Set the 'includedFiles' field to 'training/**' in the trigger configuration.
AnswerD

The includedFiles field with 'training/**' restricts the trigger to commits touching that directory, directly satisfying the requirement to run the pipeline only for training-code changes. Cloud Build evaluates this glob against changed paths before executing cloudbuild.yaml, avoiding unnecessary builds.

Why this answer

Cloud Build triggers support an `includedFiles` field that specifies a glob pattern. When set to `training/**`, the trigger will only fire if the commit includes changes to files under the `training/` directory. This is the native, declarative way to filter triggers based on changed file paths without additional infrastructure.

Exam trap

The trap here is that candidates confuse `ignoreFiles` with `includedFiles`, or assume that a custom solution like Cloud Functions is required when Cloud Build already provides a native, simpler mechanism for path-based filtering.

How to eliminate wrong answers

Option A is wrong because `ignoreFiles` excludes commits that match the pattern, but the requirement is to run the pipeline only when changes occur in `training/`, not to ignore them. Option B is wrong because `substitutions` are used for variable replacement in build configuration, not for filtering trigger conditions based on file changes. Option C is wrong because while a Cloud Function could achieve this, it introduces unnecessary complexity and cost; Cloud Build triggers natively support file path filtering via `includedFiles`, making a separate function an anti-pattern.

385
MCQmedium

A data engineer wants to use BigQuery ML to train a model for predicting customer churn (binary classification) using a large dataset. They want the model to be automatically tuned. Which model type should they choose?

A.LOGISTIC_REG
B.BOOSTED_TREE_CLASSIFIER
C.DNN_CLASSIFIER
D.AUTOML_CLASSIFIER
AnswerD

AUTOML_CLASSIFIER satisfies the automatic tuning constraint: BigQuery ML performs hyperparameter tuning and architecture search internally, so no manual configuration is needed. It handles binary classification directly, matching the churn prediction task, unlike boosted tree or logistic regression models, which require the engineer to specify tuning themselves.

Why this answer

(AUTOML_CLASSIFIER) is correct because it automatically performs architecture search and hyperparameter tuning to find the best model for binary classification tasks, such as customer churn prediction. This is ideal when the data engineer wants the model to be automatically tuned without manual intervention, as AutoML handles feature engineering, model selection, and tuning under the hood.

Exam trap

The trap here is that candidates often confuse 'automatically tuned' with models that have default hyperparameters (like LOGISTIC_REG or BOOSTED_TREE_CLASSIFIER), but only AUTOML_CLASSIFIER performs automated hyperparameter tuning and architecture search without requiring manual specification.

How to eliminate wrong answers

Option A (LOGISTIC_REG) is wrong because logistic regression does not support automatic tuning; it requires manual specification of hyperparameters like learning rate or regularization, and it is a simpler linear model that may not capture complex patterns in large datasets. Option B (BOOSTED_TREE_CLASSIFIER) is wrong because while it can be tuned, it does not offer fully automatic tuning; the user must manually set parameters such as tree depth, learning rate, and number of iterations. Option C (DNN_CLASSIFIER) is wrong because deep neural network classifiers require manual tuning of architecture (e.g., number of layers, neurons) and hyperparameters (e.g., learning rate, batch size), and they do not automatically search for the optimal configuration.

386
MCQhard

A financial institution needs to deploy a fraud detection model with strict latency <100ms per prediction and high throughput (1000 predictions/sec). The model is a deep neural network. Which architecture on Google Cloud meets these requirements?

A.Deploy the model on AI Platform Training with a single large VM
B.Deploy the model as a Cloud Function triggered by Cloud Pub/Sub
C.Use Vertex AI Batch Prediction with a fixed number of machines
D.Use Vertex AI Prediction with autoscaling enabled and GPU machine types
AnswerD

Vertex AI Prediction with autoscaling and GPU machine types provides the parallel compute needed for sub-100ms deep neural network inference at 1000 predictions/sec. GPUs accelerate matrix operations, while autoscaling adds nodes under load, satisfying both the latency and throughput constraints.

Why this answer

Vertex AI Prediction with autoscaling and GPU machine types is correct because it provides low-latency online serving with autoscaling to handle high throughput (1000 predictions/sec) while keeping latency under 100ms. GPUs accelerate deep neural network inference, and autoscaling ensures resources match demand without over-provisioning.

Exam trap

Google Cloud often tests the distinction between batch and online prediction services, where candidates mistakenly choose batch prediction for real-time requirements because they focus on throughput without considering latency constraints.

How to eliminate wrong answers

Option A is wrong because AI Platform Training is designed for model training, not real-time serving, and a single large VM cannot guarantee sub-100ms latency under high throughput due to resource contention and lack of autoscaling. Option B is wrong because Cloud Functions have a maximum timeout of 9 minutes (540 seconds) and are not optimized for high-throughput, low-latency ML inference; they also lack GPU support, making deep neural network inference too slow. Option C is wrong because Vertex AI Batch Prediction is for asynchronous, offline predictions on large datasets, not real-time serving with strict latency requirements; it processes jobs in batches and cannot meet sub-100ms per prediction.

387
MCQeasy

A company has deployed a model to a Vertex AI Endpoint and wants to receive an alert when the model's prediction latency exceeds a threshold. They have configured Cloud Monitoring to track the endpoint's latency metrics. They now need to create a notification channel to send alerts to their on-call team. Which Cloud Monitoring resource should they use to define the condition that triggers the alert?

A.A notification channel that sends emails to the on-call team.
B.An alerting policy with a condition based on the endpoint's latency metric.
C.A dashboard that displays the endpoint's latency over time.
D.A log-based metric that counts the number of high-latency requests.
AnswerB

Cloud Monitoring alerting policies define conditions that trigger alerts. To alert on prediction latency, you create an alerting policy with a condition that monitors the endpoint's latency metric (e.g., aiplatform.googleapis.com/prediction/latencies). This is the correct resource to define the threshold and trigger notifications.

Why this answer

Cloud Monitoring alerting policies are used to define conditions that trigger alerts. To alert on prediction latency, you create an alerting policy with a condition based on the endpoint's latency metric and attach a notification channel to it. Dashboards and notification channels do not define conditions.

Exam trap

The trap here is confusing the notification channel with the alerting policy; the notification channel only specifies where to send alerts, not when to send them.

388
MCQeasy

A data analyst wants to train a binary classification model on a BigQuery table without moving data out of BigQuery. They have limited ML expertise. Which approach should they take?

A.Use BigQuery ML with CREATE MODEL and LOGISTIC_REG model type.
B.Use Cloud Datalab to train an XGBoost model on BigQuery data.
C.Train a model using Vertex AI Workbench with a custom container.
D.Export the data to Cloud Storage and use Vertex AI AutoML Tables.
AnswerA

BigQuery ML trains models inside BigQuery using SQL, so data never leaves the warehouse. LOGISTIC_REG builds binary classification directly on the table, and CREATE MODEL keeps the workflow within SQL, matching the limited ML expertise and no-data-movement constraints.

Why this answer

BigQuery ML allows users to create and train binary classification models directly on data in BigQuery using SQL, with no need to move data or have deep ML expertise. The LOGISTIC_REG model type implements logistic regression, a standard algorithm for binary classification, and the CREATE MODEL statement handles all the underlying training infrastructure, making it ideal for a data analyst with limited ML skills.

Exam trap

Google often tests the distinction between low-code/no-code solutions (like BigQuery ML) and more advanced, infrastructure-heavy approaches (like custom containers or AutoML with data export), expecting candidates to recognize that the simplest, most integrated option is correct when the user has limited ML expertise and wants to avoid data movement.

How to eliminate wrong answers

Option B is wrong because Cloud Datalab is a deprecated interactive notebook service that requires users to write custom code and manage infrastructure, which is not suitable for someone with limited ML expertise and does not leverage BigQuery's native ML capabilities. Option C is wrong because Vertex AI Workbench with a custom container demands advanced knowledge of containerization, model training pipelines, and infrastructure management, far beyond the scope of a low-code solution for a data analyst. Option D is wrong because exporting data to Cloud Storage and using Vertex AI AutoML Tables, while low-code, introduces unnecessary data movement and additional complexity compared to the simpler, fully integrated BigQuery ML approach that keeps data in place.

389
Multi-Selecteasy

Which TWO of the following are low-code machine learning solutions on Google Cloud?

Select 2 answers
A.TensorFlow
B.scikit-learn
C.PyTorch
D.BigQuery ML
E.Vertex AI AutoML
AnswersD, E

BigQuery ML lets analysts create and train models using SQL directly inside BigQuery, with no coding or data movement. This satisfies the low-code requirement, since model training runs where the data already resides, avoiding export to separate ML environments.

Why this answer

BigQuery ML (D) is a low-code ML solution because it allows users to create, train, and deploy machine learning models using standard SQL queries directly within BigQuery, eliminating the need for custom coding in Python or other programming languages. Vertex AI AutoML (E) is also low-code as it provides a graphical interface and automated pipeline to train high-quality models with minimal manual intervention, handling feature engineering, model selection, and hyperparameter tuning automatically.

Exam trap

Google Cloud often tests the distinction between general-purpose ML frameworks (like TensorFlow, scikit-learn, PyTorch) that require significant coding versus managed services (BigQuery ML, AutoML) that provide low-code or no-code interfaces, leading candidates to mistakenly classify any ML tool on Google Cloud as low-code.

390
Multi-Selectmedium

Which TWO actions are recommended for collaborating on machine learning models using Vertex AI Model Registry?

Select 2 answers
A.Use Cloud Storage object labels to store model descriptions.
B.Use version aliases such as 'champion' and 'challenger' to manage model lifecycle.
C.Deploy all model versions to a single endpoint for comparison.
D.Attach custom metadata (e.g., training dataset, hyperparameters) to each model version.
E.Create a separate model entry for each training run.
AnswersB, D

Aliases enable controlled promotion of models.

Why this answer

Vertex AI Model Registry supports version aliases like 'champion' and 'challenger' to designate which model version should serve as the production candidate and which is under evaluation, enabling controlled lifecycle management and A/B testing without manual version tracking.

Exam trap

Google Cloud often tests the distinction between a single model entry with multiple versions versus separate model entries per run, and candidates mistakenly think separate entries provide better traceability, but the registry's versioning and alias system is specifically designed to avoid that fragmentation.

391
MCQmedium

A company wants to implement continuous delivery (CD) for ML models, where a model is automatically deployed to a staging environment and only promoted to production after passing an evaluation gate. Which combination of GCP services is BEST suited for orchestrating this CD pipeline?

A.Cloud Scheduler and Pub/Sub
B.Cloud Composer (Airflow) with Cloud Functions
C.Cloud Build with Vertex AI Pipelines and Cloud Deploy
D.Vertex AI Pipelines with Cloud Run
AnswerC

Cloud Build handles CI image and pipeline builds, Vertex AI Pipelines runs the training and evaluation gate, and Cloud Deploy manages progressive promotion to staging then production. Together they satisfy the stem's requirement that promotion occur only after the evaluation gate passes.

Why this answer

Cloud Build can trigger on code/model changes and run a pipeline that deploys to staging. After evaluation, if successful, it can promote to production using Cloud Deploy or directly update Vertex AI endpoints. Cloud Composer (Airflow) is also a good option for complex orchestration, but for CI/CD, Cloud Build is a natural fit.

The combination of Cloud Build, Cloud Deploy, and Vertex AI provides a robust CD pipeline.

392
MCQeasy

A company wants to classify support ticket text into categories. They have labeled historical tickets. Which Google Cloud service allows them to train a custom classification model with no code?

A.Vertex AI Matching Engine
B.AutoML Natural Language
C.Cloud Natural Language API
D.Document AI
AnswerB

AutoML Natural Language trains a custom text classification model from your labelled tickets through a point-and-click interface, requiring no coding. It satisfies the stem's no-code constraint directly, unlike Vertex AI custom training, which demands writing training code and containerising the model yourself.

Why this answer

AutoML Natural Language (now part of Vertex AI) is the correct service because it enables users to train custom text classification models using labeled data without writing any code. It provides a no-code interface for uploading datasets, training models, and evaluating performance, making it ideal for classifying support ticket text into custom categories.

Exam trap

The trap here is that candidates confuse the pre-trained Cloud Natural Language API (which requires no training but cannot be customized) with AutoML Natural Language (which requires labeled data but allows custom categories), leading them to select Option C incorrectly.

How to eliminate wrong answers

Option A is wrong because Vertex AI Matching Engine is designed for vector similarity search and embeddings, not for training custom classification models with labeled text data. Option C is wrong because Cloud Natural Language API is a pre-trained API that offers sentiment analysis, entity extraction, and syntax analysis, but it cannot be trained on custom labeled data for custom categories. Option D is wrong because Document AI is specialized for document processing (e.g., OCR, form parsing, invoice extraction) and is not intended for general text classification from labeled ticket data.

393
MCQhard

You are running a distributed training job on Vertex AI using PyTorch and the DistributedDataParallel (DDP) strategy across 4 nodes, each with 8 GPUs. You notice that the training loss is not decreasing as expected and the job occasionally hangs. You suspect a communication issue between nodes. Which of the following should you check first?

A.Switch to using the Horovod framework for distributed training.
B.Reduce the number of nodes to 2 to decrease communication overhead.
C.Increase the batch size per GPU to improve gradient synchronization.
D.Ensure that the MASTER_ADDR and MASTER_PORT environment variables are correctly set and that all nodes can communicate over the network.
AnswerD

In distributed training with PyTorch DDP, the MASTER_ADDR and MASTER_PORT are used for initializing the process group and coordinating communication. If these are misconfigured or if there are network issues preventing nodes from reaching each other, the job may hang or fail to synchronize gradients. Checking these first is essential for diagnosing communication problems in multi-node setups.

Why this answer

Communication issues in multi-node PyTorch DDP training often stem from incorrect MASTER_ADDR or MASTER_PORT settings, or network connectivity problems. These variables are critical for establishing the rendezvous point for all processes. Checking them first is a logical diagnostic step before considering other changes, as they are common causes of hangs and synchronization failures.

Exam trap

The trap here is assuming that changing batch size or framework will fix communication hangs, when the first step should be verifying the distributed training environment variables and network connectivity.

394
MCQhard

Refer to the exhibit. An alert policy is configured to trigger when prediction latency exceeds 500 ms for 5 consecutive minutes. The team is experiencing many false positive alerts during brief latency spikes. Which adjustment would most effectively reduce false positives while still detecting prolonged latency issues?

A.Change the comparison to less than
B.Add a condition that CPU utilization is also high
C.Increase the duration to 30 minutes
D.Increase the threshold to 1000 ms
AnswerC

Extending the condition duration to 30 minutes requires latency to remain above 500 ms continuously for longer, filtering out brief spikes while still firing on sustained degradation. This directly reduces false positives without masking prolonged latency issues.

Why this answer

Increasing the duration from 5 to 30 minutes (Option C) directly addresses the problem of false positives from brief latency spikes by requiring the latency to exceed 500 ms for a longer continuous period before triggering an alert. This ensures that only sustained, prolonged latency issues—not transient spikes—activate the policy, aligning with the goal of detecting genuine degradation while ignoring noise.

Exam trap

Google Cloud often tests the distinction between threshold and duration adjustments, trapping candidates who think raising the threshold (Option D) is the only way to reduce false positives, when in fact increasing the evaluation window is more precise for filtering out transient spikes without compromising detection of sustained issues.

How to eliminate wrong answers

Option A is wrong because changing the comparison to 'less than' would invert the logic, triggering alerts when latency is below 500 ms, which is the opposite of detecting high latency and would generate false positives for normal or low-latency conditions. Option B is wrong because adding a condition that CPU utilization is also high introduces an unnecessary dependency that may miss prolonged latency issues caused by other factors (e.g., network bottlenecks, memory pressure, or I/O wait), and it does not address the core problem of brief latency spikes. Option D is wrong because increasing the threshold to 1000 ms would allow sustained latency between 500 ms and 1000 ms to go undetected, failing to capture prolonged issues that still violate the original 500 ms requirement, and it does not filter out brief spikes.

395
MCQeasy

A company needs to serve a model for real-time predictions with a strict latency SLA of 100ms at the 99th percentile. The model is lightweight and traffic patterns are highly variable with occasional spikes. Which deployment strategy best meets the SLA while controlling cost?

A.Deploy the model as a Cloud Run service with autoscaling to zero.
B.Deploy to Vertex AI Endpoint with manual scaling and a fixed number of replicas.
C.Use Vertex AI Batch Prediction.
D.Deploy to Vertex AI Endpoint with min_replica_count=3 and autoscaling enabled.
AnswerD

Setting min_replica_count=3 keeps three replicas warm, eliminating cold-start latency that would breach the 100ms p99 SLA during traffic spikes. Autoscaling then absorbs variable load beyond that floor, so you pay for three always-on nodes rather than provisioning peak capacity permanently.

Why this answer

Setting a minimum number of replicas ensures baseline capacity to handle initial spikes without cold start delays, while autoscaling handles larger spikes. Option A is wrong because Cloud Run with autoscaling to zero may cause cold start delays, which could violate the strict latency SLA. Option B is wrong because manual scaling with a fixed number of replicas may lead to over-provisioning or under-provisioning.

Option C is wrong because batch prediction is not real-time.

396
Multi-Selectmedium

A company uses Vertex AI for AutoML training. Which THREE are best practices for managing model versions?

Select 3 answers
A.Deploy each model version to a separate endpoint
B.Use Vertex AI Model Registry to version models
C.Use evaluation metrics to compare versions
D.Use labels to tag models for tracking
E.Automatically delete old versions after 30 days
AnswersB, C, D

Correct: Centralized model versioning.

Why this answer

Vertex AI Model Registry is the central repository for managing and versioning models, allowing you to track iterations, compare performance, and control deployments. It provides a structured way to organize models, roll back to previous versions if needed, and maintain lineage for compliance and reproducibility.

Exam trap

The trap here is that candidates may think deploying each version to a separate endpoint is necessary for isolation, but Vertex AI's traffic splitting on a single endpoint is the correct and cost-effective approach for managing multiple model versions.

397
MCQhard

A financial services company uses Document AI to process loan applications. They want to ensure that any documents the model cannot process with high confidence are reviewed by a human before finalizing the decision. Which Document AI feature should they enable?

A.AutoML Tables model retraining
B.Cloud DLP for data inspection
C.Increase the number of processors
D.Human-in-the-Loop (HITL)
AnswerD

Human-in-the-Loop routes low-confidence extractions to human reviewers before decisions finalise, directly satisfying the requirement that uncertain documents receive manual review. Document AI assigns confidence scores per field, and HITL triggers review when those scores fall below your configured threshold, ensuring no unreliable output reaches the loan decision.

Why this answer

Human-in-the-Loop (HITL) in Document AI allows you to route documents that the model processes with low confidence to human reviewers for validation or correction before finalizing. This directly matches the requirement to have a human review any documents the model cannot process with high confidence.

Exam trap

PMLE often tests the difference between model improvement features (retraining, AutoML) and operational review features (HITL) — candidates may choose retraining when the requirement is for a human review step, not model accuracy improvement.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for tabular data, not document processing, and retraining does not provide a human review workflow. Option B is wrong because Cloud DLP is for data inspection and redaction (PII discovery), not for human review of document extraction confidence. Option C is wrong because increasing the number of processors scales throughput but does not add a human review step for low-confidence documents.

398
Multi-Selectmedium

A machine learning engineer is monitoring a deployed churn prediction model that has shown a gradual decline in accuracy over the past month. The engineer wants to diagnose the root cause of the performance degradation. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the model's learning rate and fine-tune it on the latest data.
B.Immediately retrain the model using all available historical data to improve accuracy.
C.Deploy a second model in parallel to compare predictions.
D.Use Vertex AI Model Monitoring to detect data drift by comparing the distribution of recent input features against the training data distribution.
E.Monitor the model's prediction accuracy by comparing recent predictions against newly collected ground truth labels.
AnswersD, E

Vertex AI Model Monitoring compares recent serving inputs against the training baseline, surfacing feature distribution shifts that explain gradual accuracy decay. This directly satisfies the stem's diagnostic goal: identifying data drift as the root cause of the churn model's month-long degradation, rather than retraining blindly.

Why this answer

Option D is correct because Vertex AI Model Monitoring is designed to detect training-serving skew and prediction drift by comparing the statistical distribution of recent input features against the baseline training data distribution, which directly diagnoses whether the accuracy decline stems from data drift. Option E is correct because comparing recent predictions against newly collected ground truth labels measures actual model performance degradation and confirms whether the accuracy drop is real, providing the diagnostic evidence needed before remediation. Option A is wrong because increasing the learning rate and fine-tuning is a remediation action, not a diagnostic step, and could worsen the model without first identifying the root cause.

Option B is wrong because immediately retraining on all historical data is a premature fix that does not diagnose the cause and may reintroduce stale patterns. Option C is wrong because deploying a parallel model for comparison does not identify the root cause of the existing model's degradation and adds operational complexity without diagnostic value.

Exam trap

The trap here is that candidates often confuse reactive retraining (Option B) with diagnostic monitoring, failing to recognize that the first step in troubleshooting performance degradation is to identify the root cause through drift detection and ground truth comparison, not to immediately modify or retrain the model.

399
MCQhard

You need to perform a large-scale feature computation on streaming data from Pub/Sub, transforming raw events into features, and writing results to Vertex AI Feature Store for online serving. Which Google Cloud architecture is most appropriate?

A.Use Dataproc with Spark Streaming to read from Pub/Sub and write to Feature Store
B.Use Cloud Functions triggered by Pub/Sub to compute features and update Feature Store
C.Use Dataflow streaming pipeline with Apache Beam to read from Pub/Sub, compute features, and write to Feature Store
D.Use Cloud Run to consume Pub/Sub messages and update Feature Store via a service
AnswerC

Dataflow runs Apache Beam, providing the horizontal scaling and windowing needed for high-volume Pub/Sub streams, and its native Feature Store sink writes computed features directly for low-latency online serving, meeting the streaming transformation and serving constraints.

Why this answer

Dataflow with Apache Beam is the purpose-built managed service for streaming ETL on Google Cloud, natively integrating with Pub/Sub as a source and Vertex AI Feature Store as a sink. Beam's windowing, watermarks, and exactly-once semantics handle the large-scale, continuous feature computation required for online serving. It provides autoscaling and managed infrastructure, which is exactly what a production streaming feature pipeline needs.

Exam trap

PMLE often tests the misconception that any compute service (Cloud Functions, Cloud Run) can handle streaming feature engineering, when in fact managed Apache Beam on Dataflow is the canonical answer for large-scale streaming transformations into Feature Store.

How to eliminate wrong answers

Option A is wrong because Dataproc with Spark Streaming requires manual cluster management and lacks first-class integration with Vertex AI Feature Store, adding operational overhead for a streaming pipeline. Option B is wrong because Cloud Functions has execution timeouts (up to 60 minutes for 2nd gen but typically short), limited concurrency, and is not designed for sustained high-throughput streaming feature computation. Option D is wrong because Cloud Run is a request/container-based service without native streaming windowing semantics, making it unsuitable for continuous large-scale feature computation from Pub/Sub.

400
MCQmedium

A financial services firm wants to predict loan default risk using a dataset with 30,000 labeled examples and 25 numeric and categorical features. Their team includes SQL analysts but no Python developers, and they want to minimize operational overhead. They decide to use BigQuery ML. Which model type should they use to achieve the best predictive performance while keeping the solution low-code?

A.BOOSTED_TREE_CLASSIFIER
B.DNN_CLASSIFIER
C.KMEANS
D.LOGISTIC_REG
AnswerA

BOOSTED_TREE_CLASSIFIER is an ensemble model that often achieves higher accuracy on tabular data with non-linear relationships, such as loan default prediction. It is fully supported in BigQuery ML, requires only SQL to train, and automatically handles feature engineering like categorical encoding. With 30,000 examples, it has sufficient data to train effectively without overfitting, making it ideal for this low-code, high-performance scenario.

Why this answer

For a binary classification task with a moderate-sized tabular dataset and a low-code requirement, BigQuery ML's BOOSTED_TREE_CLASSIFIER offers an excellent balance of accuracy and ease of use. It automatically handles categorical features and non-linear relationships, often outperforming linear models. It requires only SQL to train and deploy, aligning with the team's skills and minimizing operational overhead.

Exam trap

The trap here is assuming that a simple linear model like logistic regression is sufficient for all classification tasks, when actually boosted trees often yield better performance on tabular data without added complexity.

401
MCQmedium

A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?

A.Set min_replicas to an estimated baseline and max_replicas to a higher number
B.Set min_replicas and max_replicas equal to a fixed number
C.Set min_replicas to 0 and max_replicas to a high number
D.Do not set min_replicas; let Vertex AI automatically determine
AnswerA

Pinning min_replicas to an estimated baseline keeps warm capacity always available, eliminating cold starts, while max_replicas absorbs traffic spikes. This satisfies the stem's high-throughput, strict-latency and consistent-performance requirements, unlike autoscaling from zero, which reintroduces cold-start delays.

Why this answer

Setting min_replicas to an estimated baseline ensures that a minimum number of instances are always running, eliminating cold starts for baseline traffic. Setting max_replicas to a higher number allows the service to scale up to handle traffic spikes while maintaining consistent performance. This configuration balances cost and latency by avoiding the overhead of scaling from zero while still accommodating bursts.

Exam trap

Google Cloud often tests the misconception that setting min_replicas to 0 is cost-effective, but the trap here is that it ignores the strict latency requirement and the reality of cold start delays in model serving with Vertex AI.

How to eliminate wrong answers

Option B is wrong because setting min_replicas and max_replicas equal to a fixed number prevents any autoscaling, leading to either over-provisioning (waste) or under-provisioning (latency spikes) under variable load. Option C is wrong because setting min_replicas to 0 means the service can scale down to zero, causing cold starts on every request when traffic resumes, which violates the strict latency requirement. Option D is wrong because not setting min_replicas and relying on Vertex AI's automatic determination may result in the service scaling to zero or having unpredictable baseline capacity, introducing cold starts and inconsistent performance.

402
MCQhard

A company needs to maintain an audit trail of model changes for compliance. Multiple teams will be updating models. What is the best approach to track who created, modified, or deployed each model version?

A.Enable Cloud Storage audit logs and require all model files to be stored in a bucket
B.Use Cloud Logging to collect logs from all services and search for model names
C.Use Vertex AI Experiments and Metadata to track model lineage and audit logs
D.Ask team members to maintain a shared spreadsheet of changes
AnswerC

Vertex AI Experiments and Metadata record lineage automatically, capturing parameters, artifacts and the identity behind each run. This satisfies the compliance requirement for an audit trail showing who created, modified or deployed every model version across multiple teams.

Why this answer

Vertex AI Experiments and Metadata provide built-in lineage tracking that records which experiment, run, artifact, and model version were created, by which user, and with which parameters and metrics. Combined with Cloud Audit Logs for Vertex AI API calls, this gives a complete, queryable audit trail of who created, modified, or deployed each model version — exactly what compliance requires.

Exam trap

PMLE often tests whether candidates confuse generic logging (Cloud Logging, Cloud Storage audit logs) with purpose-built ML lineage (Vertex AI Metadata) — the trap is picking a logging answer that lacks model-version context.

How to eliminate wrong answers

Option A is wrong because Cloud Storage audit logs only track object-level access to model files in buckets; they do not capture model creation, versioning, or deployment events in Vertex AI, so the audit trail is incomplete. Option B is wrong because Cloud Logging collects service logs but does not provide structured model lineage — searching logs for model names is ad hoc, error-prone, and not a reliable compliance mechanism. Option D is wrong because a shared spreadsheet is manual, tamper-prone, and not authoritative; it fails audit requirements for integrity and completeness.

403
MCQmedium

An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?

A.Deploy the model on a larger machine type to handle duplicate requests faster.
B.Enable Vertex AI endpoint caching by setting the `enable_cache` flag.
C.Implement a cache layer using Cloud Memorystore for Redis, hashing prediction requests.
D.Use Cloud CDN in front of the endpoint.
AnswerC

Cloud Memorystore for Redis supplies a low-latency shared cache; hashing each prediction request produces a deterministic key so identical requests return the stored response instead of recomputing. This directly removes the redundant computation and cost the stem describes.

Why this answer

Vertex AI does not provide built-in request caching; instead, the recommended pattern is to implement an external cache like Cloud Memorystore for Redis. By hashing the prediction request payload and using it as a cache key, identical requests within the short time window can be served from Redis, eliminating redundant model inference and reducing both latency and cost.

Exam trap

The trap here is that candidates assume Vertex AI has a native caching feature (like an `enable_cache` flag) because other Google Cloud services offer caching, but Vertex AI endpoints require an external cache layer like Cloud Memorystore for Redis.

How to eliminate wrong answers

Option A is wrong because deploying on a larger machine type increases throughput but does not eliminate redundant computation for identical requests; it still performs the same inference multiple times, wasting resources. Option B is wrong because Vertex AI endpoints do not support an `enable_cache` flag; this is a fictitious feature, and Vertex AI has no built-in request caching mechanism. Option D is wrong because Cloud CDN caches static content at the edge based on HTTP cache headers, but prediction requests are typically POST with dynamic payloads that are not cacheable by CDN, and CDN cannot inspect or hash request bodies for deduplication.

404
MCQmedium

A company needs to forecast product demand for the next 12 months using historical sales data. They want to use BigQuery ML with minimal coding. Which model type is most suitable?

A.K_MEANS
B.MATRIX_FACTORIZATION
C.ARIMA_PLUS
D.LINEAR_REG
AnswerC

ARIMA_PLUS handles time-series forecasting natively in BigQuery ML, requiring only a single CREATE MODEL statement on the historical sales column. It automatically detects seasonality, trends and holidays, satisfying the 12-month horizon and minimal-coding constraint without exporting data or writing Python.

Why this answer

ARIMA_PLUS is BigQuery ML's purpose-built time-series forecasting model, designed for exactly this scenario: forecasting future values from historical time-ordered data with minimal SQL coding. It automatically handles seasonality, holidays, and trend decomposition, and supports features like forecasting multiple time series at once. LINEAR_REG could technically model time as a feature, but it cannot capture seasonality or autocorrelation, making it unsuitable for demand forecasting.

Exam trap

The trap here is that candidates see 'forecast' and reach for LINEAR_REG because it is a regression model, forgetting that time-series forecasting requires specialized models like ARIMA_PLUS that handle seasonality and autocorrelation.

How to eliminate wrong answers

Option A is wrong because K_MEANS is an unsupervised clustering algorithm used for segmentation, not for predicting future numeric values over time. Option B is wrong because MATRIX_FACTORIZATION is a collaborative-filtering technique for recommendation systems (e.g., user-item matrices), not time-series forecasting. Option D is wrong because LINEAR_REG assumes independent observations and cannot model seasonality, trends, or autocorrelated time-series structure, so it would produce poor demand forecasts.

405
MCQmedium

A company deploys a custom ML model on Vertex AI to predict customer churn. The model retrains weekly, and predictions are served via a Vertex AI endpoint. After a recent retraining, the monitoring dashboard shows a sudden increase in prediction requests but a decrease in predicted churn probabilities. The model's accuracy on the validation set remains stable. What is the most likely cause of the observed behavior?

A.A training-serving skew exists between the training pipeline and the serving endpoint.
B.Concept drift has occurred, changing the relationship between features and churn.
C.The incoming data distribution has changed, e.g., due to a new marketing campaign attracting different customers.
D.Data leakage during training caused the model to overfit to historical patterns.
AnswerC

A covariate shift in the live feature distribution explains both symptoms: new marketing-driven customers produce different input patterns, inflating request volume while shifting predicted probabilities downward. Validation accuracy stays stable because the validation set still reflects the old distribution, so the drift is invisible there.

Why this answer

A sudden increase in prediction requests alongside a decrease in predicted churn probabilities, while validation accuracy remains stable, indicates a shift in the incoming data distribution (covariate shift). This is typical when a new marketing campaign attracts a different customer segment that inherently has lower churn risk. The model itself hasn't degraded; it's simply seeing a different population than it was trained on, which changes the base rate of churn in the live traffic.

Exam trap

Google Cloud often tests the distinction between covariate shift (data distribution change) and concept drift (relationship change), trapping candidates who assume any change in predictions must be due to model degradation or data leakage.

How to eliminate wrong answers

Option A is wrong because training-serving skew refers to a mismatch in feature preprocessing or data format between training and serving, which would typically cause a drop in accuracy or anomalous predictions, not a stable validation accuracy with a shift in prediction distribution. Option B is wrong because concept drift would change the relationship between features and the target (churn), leading to a decline in model accuracy on the validation set, which is explicitly stated as stable. Option D is wrong because data leakage during training would cause overfitting to historical patterns, resulting in poor generalization and a drop in validation accuracy, not a stable accuracy with a shift in prediction probabilities.

406
MCQmedium

A team wants to deploy two versions of a model (v1 and v2) on Vertex AI Endpoint to conduct an A/B test. They need to split traffic so that 10% of requests go to v2. Which configuration achieves this?

A.Deploy both versions on the same endpoint and use the `traffic_split` parameter to allocate 90% to v1 and 10% to v2.
B.Configure a global load balancer in front of two endpoints and set the weight.
C.Create two separate endpoints, one for each version, and have the client randomly select the endpoint.
D.Deploy v2 as a canary deployment and set the canary rollout to 10% in Cloud Deployment Manager.
AnswerA

One endpoint can host both versions as separate DeployedModels, and the traffic_split parameter maps each version's ID to a percentage. Assigning 90% to v1 and 10% to v2 delivers exactly the required A/B distribution through a single endpoint.

Why this answer

Vertex AI Endpoints allow distributing traffic between deployed models using the traffic_split parameter. Setting 90% to v1 and 10% to v2 achieves the A/B test traffic split. Option B is incorrect because a global load balancer in front of two endpoints adds unnecessary complexity and is not the native Vertex AI method.

Option C is incorrect because the client randomly selecting endpoints introduces client-side logic and does not leverage Vertex AI's built-in traffic splitting. Option D is incorrect because Cloud Deployment Manager's canary rollout is for infrastructure deployment, not for model traffic splitting; Vertex AI provides traffic splitting for A/B testing.

407
Multi-Selectmedium

Which THREE practices improve collaboration when using Cloud Composer for ML pipelines?

Select 3 answers
A.Keep all pipeline logic in a single large DAG for simplicity.
B.Use a shared Cloud Storage bucket for intermediate artifacts with appropriate permissions.
C.Store DAGs in a version-controlled repository and use CI/CD to deploy them.
D.Embed service account keys directly in DAG code for authentication.
E.Use Airflow variables and connections to parameterize DAGs.
AnswersB, C, E

Facilitates handoff between pipeline steps and teams.

Why this answer

Cloud Composer workflows often require sharing intermediate data (e.g., transformed datasets, model checkpoints) across multiple DAGs or team members. A shared Cloud Storage bucket with fine-grained IAM permissions enables secure, centralized artifact exchange without duplicating data or exposing it to unauthorized users. This practice avoids hard-coded paths and ensures that all pipeline stages can reliably access the same artifacts, which is critical for reproducibility and collaboration in ML pipelines.

Exam trap

Google Cloud often tests the misconception that a single monolithic DAG simplifies collaboration, when in fact it creates bottlenecks and merge conflicts; the trap is that candidates confuse 'simplicity' with 'ease of collaboration' without considering modularity and CI/CD practices.

408
MCQhard

Your team uses a Vertex AI Pipeline that reads from a BigQuery table, trains a model, and registers it. A teammate wants to know which BigQuery table snapshot was used for a specific registered model version so they can reproduce the training data exactly. Which action should they take?

A.Query Vertex ML Metadata for the dataset artifact linked as an input to the execution that produced the model version.
B.List the model's versions in Vertex AI Model Registry and inspect the version's description field for the table name.
C.Check the Vertex AI Feature Store for the entity type that corresponds to the BigQuery table.
D.Open the pipeline run in the Vertex AI console and read the BigQuery table name from the run's parameters.
AnswerA

Vertex ML Metadata records the dataset artifact as an input to the training execution, and that artifact can carry a URI or metadata identifying the BigQuery snapshot. Traversing from the model artifact backward to this input artifact gives the exact data reference needed for reproduction.

Why this answer

Reproducing training data requires the exact dataset reference, which is captured as an input artifact in Vertex ML Metadata. Traversing the lineage from the model artifact to the training execution and its input dataset artifact yields that reference, whereas table names, descriptions, or feature stores do not guarantee snapshot-level accuracy.

Exam trap

The trap here is assuming that a table name in pipeline parameters or a description field is equivalent to an immutable dataset snapshot.

409
MCQmedium

You are an MLOps engineer at a retail company. Your team has deployed a demand forecasting model to a Vertex AI Endpoint. The model uses 12 numerical features and outputs a single numeric value representing predicted units sold. You have configured Vertex AI Model Monitoring with a training dataset that includes the full feature schema and prediction distribution. After a week, you observe that the feature 'promotion_flag' has a Jensen-Shannon divergence of 0.15, while all other features remain below 0.05. The model's prediction distribution has also shifted. Which action should you take first to diagnose the cause of the drift?

A.Examine the distribution of the 'promotion_flag' feature in the recent serving data and compare it with the training distribution to determine if the feature's meaning or data collection process has changed.
B.Increase the monitoring frequency to 15 minutes and lower the drift threshold for 'promotion_flag' to 0.05 to get more granular alerts.
C.Ignore the drift because a Jensen-Shannon divergence of 0.15 is below the typical threshold of 0.2 used for numerical features.
D.Immediately trigger a retraining pipeline using the most recent data to update the model and reduce the drift.
AnswerA

The 'promotion_flag' feature shows the highest divergence, indicating a significant shift. Comparing recent serving data distribution to the training distribution helps identify whether the feature's meaning or data collection process has changed. This diagnostic step is appropriate before considering retraining or other actions, as it pinpoints the root cause of drift and informs subsequent mitigation steps.

Why this answer

The feature 'promotion_flag' has the highest drift, so examining its recent distribution compared to training helps identify if the feature's meaning or data collection has changed. This diagnostic step is crucial before retraining or adjusting monitoring, as it reveals whether the drift is due to a data pipeline issue or a genuine change in the underlying pattern.

Exam trap

The trap here is assuming that any drift requires immediate retraining, when in fact the first step should be to diagnose the cause of the drift to ensure the right corrective action.

410
MCQmedium

A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?

A.Add the SQL query as a pipeline parameter and pass it to the component as an input.
B.Set the component's `enable_caching` argument to False in the pipeline definition.
C.Modify the component's container image tag to a new version and rebuild the pipeline.
D.Clear the pipeline's cache by deleting the pipeline run's metadata from Vertex ML Metadata.
AnswerA

Vertex AI Pipelines computes the cache key from the component's inputs (including code and parameters). By promoting the SQL query from hardcoded code to an input parameter, any change to the query alters the input, invalidating the cache and triggering re-execution. This maintains caching benefits for unchanged queries and aligns with pipeline best practices of parameterizing variable logic. It directly solves the problem without globally disabling caching.

Why this answer

Vertex AI Pipelines caching uses a fingerprint derived from the component's code and input parameters. When a component's internal logic changes but its inputs and code hash remain the same, the cache is considered valid. To force re-execution without disabling caching entirely, the engineer should expose the changing element—the SQL query—as an input parameter.

This changes the cache key only when the query changes, preserving caching for other runs. Rebuilding the image or disabling caching are less precise and more disruptive.

Exam trap

The trap here is assuming that modifying the component's source code automatically changes the cache key when the component is not rebuilt or its inputs are not changed.

411
MCQhard

A recommendation system model is updated daily via a retraining pipeline. After each update, the online prediction latency increases significantly for about 30 minutes before returning to normal. What is the most likely cause and solution?

A.The Vertex AI endpoint autoscaling policy is too aggressive, causing scale-down during retraining.
B.The retraining pipeline runs on a GKE cluster that shares resources with the serving endpoint.
C.The model is being switched from CPU to GPU at deployment.
D.The new model version causes cold start in the serving infrastructure; pre-warm the model by sending a dummy request after deployment.
AnswerD

Cold start occurs because the newly deployed model version's serving containers must load weights and initialise before handling traffic, inflating latency until caches warm. Sending a dummy request immediately after deployment pre-warms the infrastructure, eliminating the 30-minute degradation window.

Why this answer

The most likely cause is a cold start when the new model version is deployed. The serving infrastructure needs to load the model into memory, which takes time and causes increased latency for the first requests. Pre-warming the model by sending dummy requests after deployment can mitigate this.

Option A is incorrect because autoscaling policy affecting scale-down would not cause a temporary spike after each update; it would be more persistent. Option B is incorrect because GKE cluster sharing resources would cause consistent latency issues, not just for 30 minutes after update. Option C is incorrect because CPU-to-GPU switching is not a typical cause for temporary latency spikes.

412
MCQhard

A team uses Vertex AI Pipelines to automate training and deployment. They need to ensure that only models that pass a set of quality checks (e.g., accuracy > 0.9, latency < 100ms) are deployed to production. How should they implement this?

A.Manually review each model before promotion
B.Use Cloud Functions to deploy only if accuracy is reported in BigQuery
C.Set up Cloud Build triggers to deploy every model version
D.Add a Pipeline component that evaluates metrics and uses a conditional gate to deployment
AnswerD

A pipeline component computes the quality metrics, and a conditional gate evaluates them against thresholds such as accuracy above 0.9 and latency under 100ms, only allowing deployment when all checks pass. This enforces automated quality control within Vertex AI Pipelines itself.

Why this answer

Vertex AI Pipelines supports conditional logic through components and pipeline control flow, allowing a component to evaluate model metrics and a downstream conditional gate to promote the model only if thresholds (accuracy > 0.9, latency < 100ms) are met. This embeds quality gates directly into the MLOps pipeline, ensuring automated, repeatable promotion decisions.

Exam trap

PMLE often tests whether candidates choose manual or external mechanisms instead of native pipeline control flow, tricking them into selecting Cloud Functions or manual review when the correct answer is an in-pipeline conditional gate.

How to eliminate wrong answers

Option A is wrong because manual review does not scale, is error-prone, and defeats the purpose of automating training and deployment with Vertex AI Pipelines. Option B is wrong because using Cloud Functions triggered by BigQuery metrics introduces an external, loosely coupled mechanism that is not integrated into the pipeline's control flow and adds latency and complexity. Option C is wrong because deploying every model version unconditionally bypasses quality checks entirely, which is the opposite of the requirement.

413
Multi-Selecthard

A team is collaborating on a Vertex AI model using Vertex AI Model Registry. They need to ensure that model versions are properly managed and that deployments are reproducible. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Use a single default version for all deployments to simplify endpoint configuration.
B.Rely on the model's auto-generated resource name to identify versions, as it includes a timestamp.
C.Grant all team members the Vertex AI User role on the project to allow unrestricted model uploads.
D.Assign each model version a unique alias that maps to a specific model artifact and container image.
E.Store the training pipeline's parameters and data snapshot URI in the model's description or labels.
AnswersD, E

Aliases in Vertex AI Model Registry provide a stable reference to a specific model version, enabling reproducible deployments. By mapping an alias to a model artifact and its container image, teams can deploy the exact version regardless of later updates. This supports collaboration because team members can refer to the alias without tracking version numbers manually.

Why this answer

Using aliases to map model versions to specific artifacts and container images, and recording pipeline parameters and data snapshot URIs in metadata, together ensure that deployments are reproducible and that team members can collaborate effectively. Aliases provide stable references, while metadata captures the training context for audits and retraining.

Exam trap

The trap here is assuming that auto-generated resource names or broad permissions are sufficient for version management, when in fact deliberate aliasing and metadata capture are required for reproducibility.

414
MCQhard

You are an ML engineer at a logistics company. The company uses a Vertex AI Pipeline with BigQuery ML to train a model that predicts delivery delays based on weather, traffic, and historical order data. The pipeline runs daily and includes steps: (1) data extraction from BigQuery, (2) feature engineering using Dataflow, (3) model training with BigQuery ML (logistic regression), (4) model evaluation, and (5) conditional deployment to a Vertex AI Endpoint if accuracy > 0.85. Recently, the pipeline has been failing at step 5 with the error: "Vertex AI Endpoint creation failed: Quota limit of 1 endpoint per region exceeded." The company has already created one endpoint in the same region for another model. The pipeline is configured to create a new endpoint each time a model is deployed. The engineer needs to fix this with minimal changes to the pipeline code. Which course of action should the engineer take?

A.Submit a quota increase request to Google Cloud for Vertex AI Endpoints in the current region.
B.Change the region in the pipeline configuration to a region with available endpoint quota.
C.Remove the accuracy threshold and deploy every model automatically to a pre-created endpoint.
D.Modify the deployment step to check if an endpoint already exists and, if so, deploy a new model version to the existing endpoint instead of creating a new one.
AnswerD

Reusing the existing endpoint sidesteps the one-endpoint-per-region quota, since Vertex AI caps endpoints per region rather than deployed models. Deploying the new model as an additional model version on that endpoint satisfies step 5's conditional deployment while requiring only a change to the deployment step, leaving the pipeline's other components untouched.

Why this answer

It directly addresses the root cause: the pipeline fails because it tries to create a new endpoint each time, exceeding the regional quota of one endpoint. By modifying the deployment step to check for an existing endpoint and deploying a new model version to it, the engineer avoids quota issues without altering the pipeline's core logic or requiring external approvals. This approach leverages Vertex AI's model versioning capability, which allows multiple model versions under a single endpoint, aligning with minimal code changes.

Exam trap

The trap here is that candidates may focus on quota limits as a resource issue (Option A) or a region issue (Option B), rather than recognizing that the pipeline's deployment logic is architecturally flawed by creating a new endpoint per deployment, which is both inefficient and violates best practices for model serving.

How to eliminate wrong answers

Option A is wrong because submitting a quota increase request is a slow, administrative process that does not constitute a minimal code change and may not be approved quickly, leaving the pipeline broken in the meantime. Option B is wrong because changing the region introduces additional complexity (e.g., data residency, latency, and potential BigQuery dataset location mismatches) and does not address the underlying design issue of creating a new endpoint per deployment. Option C is wrong because removing the accuracy threshold undermines the model quality gate, potentially deploying poor models, and still requires creating a new endpoint each time, which would still hit the quota limit.

415
MCQmedium

An organization wants to trigger a Vertex AI pipeline whenever new data arrives in a Cloud Storage bucket. Which approach should they use?

A.Configure a Vertex AI pipeline trigger directly on the bucket using the GCP Console.
B.Use Pub/Sub notifications from the bucket and a Dataflow job to start the pipeline.
C.Set up a Cloud Scheduler job that runs every minute and checks for new files in the bucket.
D.Use Cloud Functions triggered by Cloud Storage events to call the Vertex AI pipeline API.
AnswerD

A Cloud Storage-triggered Cloud Function receives the object-finalise event and calls the Vertex AI Pipelines API to launch a run, giving event-driven execution without polling. This satisfies the stem's requirement to trigger the pipeline whenever new data lands in the bucket.

Why this answer

Cloud Functions can be directly triggered by Cloud Storage events (e.g., `google.storage.object.finalize`) and can then call the Vertex AI pipeline API using the Cloud SDK or client libraries. This provides a serverless, event-driven architecture that reacts immediately to new data without polling or additional infrastructure.

Exam trap

The trap here is that candidates may assume Vertex AI pipelines have a built-in Cloud Storage trigger (Option A) or over-engineer the solution with Dataflow (Option B), when the simplest and most native serverless approach is Cloud Functions.

How to eliminate wrong answers

Option A is wrong because Vertex AI pipelines do not support configuring a trigger directly on a Cloud Storage bucket via the GCP Console; there is no native bucket-to-pipeline trigger. Option B is wrong because while Pub/Sub notifications from the bucket are possible, adding a Dataflow job introduces unnecessary complexity and latency—Dataflow is a batch/stream processing engine, not a lightweight event router. Option C is wrong because a Cloud Scheduler job that runs every minute and checks for new files is inefficient (polling), introduces up to 60 seconds of latency, and does not scale well; it also requires custom code to track file states.

416
MCQeasy

A developer wants to add text translation to a mobile app. They need to translate user-generated content into multiple languages, and latency is critical. Which pre-built API should they use?

A.Translation API
B.Vision API
C.Text-to-Speech API
D.Natural Language API
AnswerA

The Translation API is a pre-built, low-latency service supporting many language pairs, so it meets the requirement to translate user-generated content into multiple languages without training custom models. Custom translation would add latency and development overhead.

Why this answer

Translation API provides fast, real-time translation for text. Natural Language API is for analysis. Text-to-Speech is for audio.

Vision API is for images.

417
Multi-Selectmedium

A company wants to build a model to predict housing prices using BigQuery ML. They have a dataset with features like area, number of bedrooms, and location. Which TWO model types are appropriate for this regression task?

Select 2 answers
A.LOGISTIC_REG
B.K_MEANS
C.MATRIX_FACTORIZATION
D.BOOSTED_TREE_REGRESSOR
E.LINEAR_REG
AnswersD, E

Why this answer

BOOSTED_TREE_REGRESSOR (D) is appropriate because it is a tree-based ensemble method specifically designed for regression tasks, and BigQuery ML supports it via the `CREATE MODEL` statement with `model_type='BOOSTED_TREE_REGRESSOR'`. It handles non-linear relationships and interactions between features like area, bedrooms, and location, making it suitable for predicting continuous housing prices.

Exam trap

The PMLE exam often tests the distinction between regression and classification models, leading candidates to mistakenly choose LOGISTIC_REG for regression tasks because of the word 'regression' in its name, but it is actually a classification algorithm.

418
MCQmedium

A data scientist needs to retrieve training data from Vertex AI Feature Store that exactly matches the feature values as they were at a specific historical timestamp to avoid label leakage. Which feature view configuration should they use?

A.Enable point-in-time retrieval on the feature view.
B.Use the offline store without point-in-time and rely on data ordering.
C.Use the online store with a timestamp filter.
D.Create a new feature view with only historical data.
AnswerA

Point-in-time retrieval returns feature values as they existed at the supplied timestamp, joining each training row to its historical feature state. This directly prevents label leakage, satisfying the requirement that retrieved features exactly match values at a specific historical timestamp rather than current values.

Why this answer

Point-in-time retrieval is a feature of Vertex AI Feature Store that returns feature values as of a specified timestamp.

419
Multi-Selectmedium

A team of data scientists and ML engineers is collaborating on a shared feature store in Vertex AI Feature Store. They need to ensure that feature definitions are versioned and that changes are reviewed before being used in production pipelines. Which TWO practices should they implement?

Select 2 answers
A.Allow data scientists to edit feature definitions directly in the Vertex AI Feature Store console.
B.Require code reviews for all changes to feature definitions before merging to the main branch.
C.Define multiple feature views in Vertex AI Feature Store for different environments and manage access via IAM.
D.Store feature definition code in a version-controlled repository such as Cloud Source Repositories.
E.Use scheduled batch jobs to synchronize feature definitions from a shared spreadsheet to Vertex AI Feature Store.
AnswersB, D

Code reviews ensure quality and approval.

Why this answer

Requiring code reviews for all changes to feature definitions before merging to the main branch enforces a peer-review gate, ensuring that modifications are validated for correctness, consistency, and compliance before they reach production. This aligns with MLOps best practices for governance and reduces the risk of introducing errors or breaking changes into the feature store.

Exam trap

Google Cloud often tests the distinction between environment isolation (IAM and multiple feature views) and the actual versioning/review process, leading candidates to mistakenly select Option C as a versioning practice when it only addresses access control and environment separation.

420
MCQhard

A logistics company uses Vertex AI AutoML Tables to predict delivery delays based on order attributes, weather data, and traffic data. The model is retrained weekly using a Vertex AI Pipeline that runs a BigQuery query to get training data, then triggers AutoML training. Recently, the pipeline fails with the error 'Dataset not found' when the AutoML training step starts. The BigQuery query runs successfully and outputs a table. Which is the most likely cause?

A.The AutoML training step is referencing a different dataset location.
B.The training data has been manually deleted from Cloud Storage.
C.The pipeline's IAM permissions are insufficient to access BigQuery.
D.The BigQuery output table is not being passed as a Vertex AI Dataset resource.
AnswerD

AutoML training requires a Vertex AI Dataset resource, not a raw BigQuery table reference. The pipeline's BigQuery step succeeds, but the training step cannot locate a registered dataset, producing 'Dataset not found'. The output table must first be imported into a Vertex AI Dataset before AutoML training can consume it.

Why this answer

The error 'Dataset not found' occurs because AutoML Tables requires a Vertex AI Dataset resource (a metadata wrapper) to reference the training data, not just a BigQuery table. The pipeline's BigQuery query produces a table, but if that table is not explicitly converted into or passed as a Vertex AI Dataset resource (via the `aiplatform.Dataset` creation step), AutoML training cannot locate it. Option D correctly identifies this missing step as the root cause.

Exam trap

Google Cloud often tests the distinction between a raw data source (BigQuery table) and a Vertex AI Dataset resource, trapping candidates who assume AutoML can directly consume a BigQuery table without the required metadata wrapper.

How to eliminate wrong answers

Option A is wrong because the error is 'Dataset not found', not a location mismatch; AutoML Tables uses Dataset resource IDs, not direct paths, so a different dataset location would cause a different error (e.g., 'Permission denied' or 'Table not found'). Option B is wrong because the training data is stored in BigQuery, not Cloud Storage, and the error occurs at the AutoML step, not during data retrieval; manual deletion of a Cloud Storage file would not affect a BigQuery-sourced dataset. Option C is wrong because the BigQuery query runs successfully, proving the pipeline's IAM permissions to access BigQuery are sufficient; insufficient permissions would fail at the query step, not at the AutoML training step.

421
MCQhard

A data science team is deploying a PyTorch model for real-time inference using Vertex AI Endpoints. The model requires a custom container with specific CUDA drivers and Python packages. They have created a Docker image and pushed it to Artifact Registry. The pipeline should automatically retrain the model every week and deploy the new version if it passes validation. However, the deployment step fails intermittently with the error 'The container image is not compatible with the machine type.' What is the most likely cause?

A.The service account does not have permission to pull the container from Artifact Registry.
B.The container image requires GPU support but the machine type specified in the endpoint is a CPU-only machine.
C.The container's health check endpoint is not responding correctly.
D.The model artifact size exceeds the maximum allowed for the machine type.
AnswerB

Vertex AI validates the container against the endpoint's machine type; a CUDA-dependent image cannot start on a CPU-only node, producing the incompatibility error. Selecting a GPU-backed machine type aligns the runtime environment with the image's driver requirements.

Why this answer

The error 'The container image is not compatible with the machine type' indicates a mismatch between the container's hardware requirements and the machine type selected for the Vertex AI Endpoint. Since the custom container requires specific CUDA drivers, it is built for GPU acceleration. If the endpoint is configured with a CPU-only machine type (e.g., n1-standard-4), the container will fail to run because the GPU drivers cannot initialize, triggering this incompatibility error.

Exam trap

Google Cloud often tests the distinction between deployment-time compatibility errors and runtime health check failures, tricking candidates into confusing a misconfigured machine type with a failing health probe.

How to eliminate wrong answers

Option A is wrong because a permission issue (e.g., missing artifactregistry.reader role) would produce an 'unauthorized' or 'access denied' error when pulling the image, not a compatibility error. Option C is wrong because a failing health check would cause the deployment to succeed initially but then report the container as unhealthy, not a pre-deployment compatibility error. Option D is wrong because Vertex AI has no per-machine-type artifact size limit; model size constraints are separate and would manifest as a resource-exhausted error, not a compatibility error.

422
Multi-Selecthard

You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?

Select 2 answers
A.Use Cloud Functions to transform each file individually.
B.Use Cloud SQL to store intermediate results.
C.Run Vertex AI batch prediction job with GCS source pointing to the processed TFRecord files.
D.Use Dataflow to read CSV, perform feature engineering, and write to GCS in TFRecord format.
E.Use Dataflow to read CSV, perform feature engineering, and write to BigQuery.
AnswersC, D

Running a Vertex AI batch prediction job against GCS-hosted TFRecord files lets the pipeline consume the preprocessed, engineered features at scale, avoiding repeated transformation of 50 TB of raw CSV. TFRecord input suits large batch workloads efficiently.

Why this answer

Option D is correct because Dataflow is the managed, horizontally scalable Apache Beam service designed for large-scale ETL like 50 TB of CSV, and it can perform the required complex transformations (datetime parsing, one-hot encoding) before writing the engineered features to GCS in TFRecord format, which is an efficient binary format for ML training and prediction. Option C is correct because a Vertex AI batch prediction job can then consume those processed TFRecord files directly from GCS as its input source, letting the model run inference at scale without re-doing feature engineering. Option A is not appropriate because Cloud Functions is an event-driven, short-lived serverless compute service with limited memory and execution time, making it unsuitable for transforming 50 TB of data file-by-file.

Option B is not appropriate because Cloud SQL is a relational OLTP database, not a scalable store for massive intermediate ML feature data. Option E is not appropriate because writing engineered features to BigQuery does not produce the TFRecord input that Vertex AI batch prediction expects, and BigQuery is not the target format for this pipeline's model input.

Exam trap

Candidates often assume that any scalable service can handle batch processing, but Cloud Functions and Cloud SQL are unsuitable for 50TB. The trap is to think that more than two services are needed, but a Dataflow pipeline and Vertex AI batch prediction are sufficient.

423
MCQhard

A global e-commerce company uses BigQuery ML to forecast daily sales for 10,000 products. They use a time-series model with a horizon of 7 days. Recently, forecasts for a specific product category have been consistently too high. They suspect the model is not capturing a new seasonal pattern. Which action should they take first to diagnose the issue?

A.Retrain the model with minimal additional data
B.Run ML.EVALUATE on the recent sales data and compare accuracy metrics
C.Increase the forecast horizon to 14 days
D.Switch to AutoML forecasting via Vertex AI AutoML
AnswerB

ML.EVALUATE computes accuracy metrics such as MAE and RMSE against recent labelled sales, revealing whether the systematic over-prediction stems from the model failing to capture the new seasonal pattern before any retraining or feature changes are attempted.

Why this answer

Running ML.EVALUATE on recent sales data allows you to compute accuracy metrics (e.g., MAE, MAPE) specifically for the period where the model is failing. This isolates whether the error is due to a new seasonal pattern or another cause, without retraining or changing the model architecture. It is the standard first diagnostic step in BigQuery ML for time-series models.

Exam trap

Google Cloud often tests the principle that diagnosis must precede action—candidates mistakenly jump to retraining or switching tools instead of evaluating the existing model's performance on the problematic data window.

How to eliminate wrong answers

Option A is wrong because retraining with minimal additional data does not diagnose why forecasts are too high; it only incorporates more data without identifying the root cause. Option C is wrong because increasing the forecast horizon to 14 days would worsen the problem by extending predictions further into the uncertain future, not addressing the seasonal pattern miss. Option D is wrong because switching to AutoML forecasting via Vertex AI AutoML is a premature architectural change that bypasses the diagnostic step; you should first evaluate the current model to understand the error before migrating.

424
MCQmedium

A team is using Delta Lake on Dataproc for their data lake with ACID transactions. They want to version data for ML experiments and roll back to a previous version if needed. Which Delta Lake feature should they use?

A.Delta Lake streaming
B.Delta Lake schema enforcement
C.Delta Lake time travel
D.Delta Lake optimization (Z-order)
AnswerC

Time travel queries Delta table snapshots by version or timestamp, letting the team reproduce an exact training dataset and restore an earlier state via RESTORE. This directly satisfies the versioning and rollback requirement, unlike schema evolution or Z-ordering, which address structure and query performance.

Why this answer

Delta Lake time travel allows querying previous versions of a Delta table using either a version number or a timestamp, enabling reproducibility for ML experiments and rollback to a prior state. This is the native Delta Lake feature designed for versioning and auditing data changes. It directly satisfies the requirement to version data and roll back.

Exam trap

PMLE often tests confusion between Delta Lake features — candidates may pick schema enforcement or Z-order thinking they provide versioning, when only time travel enables historical queries and rollback.

How to eliminate wrong answers

Option A is wrong because Delta Lake streaming is about ingesting and processing streaming data, not versioning or rollback. Option B is wrong because schema enforcement prevents incompatible writes but does not provide historical version access. Option D is wrong because Z-order optimization improves query performance by co-locating related data, not versioning or rollback.

425
MCQeasy

An ML engineer has a prototype scikit-learn model that must be served on Vertex AI. The model requires a custom preprocessing step that cannot be expressed in a scikit-learn Pipeline. The engineer wants to package the model with this preprocessing logic and deploy it to a Vertex AI Endpoint for online predictions. Which approach should they take?

A.Use Vertex AI Model Monitoring to apply the preprocessing transformations before the model receives the request.
B.Save the model with joblib, upload it as a Vertex AI Model, and deploy it to an Endpoint; Vertex AI automatically applies the preprocessing.
C.Deploy the model as a batch prediction job, which automatically applies preprocessing from a saved scikit-learn Pipeline.
D.Create a custom container that includes the scikit-learn model and the preprocessing code, push it to Artifact Registry, import it as a Vertex AI Model, and deploy to an Endpoint.
AnswerD

A custom container allows the engineer to bundle the model artifact, the preprocessing logic, and the required dependencies into a single image. Vertex AI runs this container for online predictions, so the preprocessing executes exactly as written before the model inference. This is the supported and recommended way to serve models with custom preprocessing on Vertex AI.

Why this answer

When a model requires custom preprocessing that is not part of the model artifact itself, the most reliable way to serve it on Vertex AI is to build a custom container. The container can include the model, the preprocessing code, and all dependencies, ensuring consistent behavior between training and serving. Vertex AI then deploys this container as a Model resource and makes it available via an Endpoint.

Exam trap

The trap here is assuming that Vertex AI automatically applies preprocessing defined in a scikit-learn Pipeline or that Model Monitoring can transform requests.

426
MCQeasy

A marketing agency uses Vertex AI AutoML Vision to classify social media images into brand logos and generic content. They have 5,000 images per class. The model achieves 95% accuracy on validation set, but in production it misclassifies many images that contain logos in unusual angles or lighting. They have limited ML expertise and want to improve robustness. Which action should they take?

A.Switch to a custom CNN model trained with data augmentation.
B.Augment the training set with images that have varied angles and lighting.
C.Deploy the model with a lower confidence threshold.
D.Use Vertex AI Matching Engine for similarity search instead.
AnswerB

Adding training images with varied angles and lighting exposes the model to the same distribution shift causing production errors, letting AutoML Vision learn rotation- and illumination-invariant features. This directly addresses the robustness gap without requiring ML expertise.

Why this answer

The core issue is a domain shift between the training data (likely clean, canonical logo images) and production data (logos at unusual angles and lighting). Augmenting the training set with those specific variations directly addresses the lack of robustness by exposing the model to the missing edge cases during training, which is the most effective and simplest fix for a team with limited ML expertise using AutoML Vision.

Exam trap

The trap here is that candidates often assume a more complex model (custom CNN) is needed for robustness, when in fact the problem is a data distribution mismatch that can be fixed with simple data augmentation, which is the most practical solution for a team with limited ML expertise using a managed service like AutoML.

How to eliminate wrong answers

Option A is wrong because switching to a custom CNN model requires significant ML expertise to design, train, and tune, which contradicts the team's limited ML expertise; AutoML Vision already uses a CNN-based architecture under the hood, so the issue is data quality, not model architecture. Option C is wrong because lowering the confidence threshold would increase the number of false positives (misclassifying generic content as logos), which does not fix the model's inability to correctly recognize logos at unusual angles—it only changes the decision boundary, not the model's feature representation. Option D is wrong because Vertex AI Matching Engine is designed for similarity search (e.g., finding nearest neighbors in an embedding space), not for classification; it would require generating embeddings for all images and does not directly solve the classification robustness problem, nor does it leverage the existing labeled training data.

427
MCQmedium

A team is training a large TensorFlow model that requires more memory than a single GPU provides. They have access to multiple GPUs on a single machine. Which distributed training strategy should they use to split the model layers across GPUs?

A.tf.distribute.experimental.MultiWorkerMirroredStrategy
B.tf.distribute.experimental.ParameterServerStrategy
C.Manual device placement using tf.device to assign layers to specific GPUs
D.tf.distribute.MirroredStrategy
AnswerC

tf.device lets you pin individual layers to named GPUs, so a model too large for one device is partitioned layer-by-layer across the machine's GPUs. This directly satisfies the stem's requirement to split model layers rather than replicate the model.

Why this answer

When a single model's layers exceed one GPU's memory, the model itself must be partitioned across devices — this is model parallelism. Manual device placement with tf.device('/GPU:0'), tf.device('/GPU:1'), etc. is the TensorFlow-native way to assign specific layers or operations to specific GPUs, splitting the model across them.

Exam trap

The trap is assuming any 'distributed strategy' solves memory limits — most strategies (Mirrored, MultiWorker, ParameterServer) replicate the model and only help with speed, not with fitting an oversized model.

How to eliminate wrong answers

Option A is wrong because MultiWorkerMirroredStrategy replicates the full model on each worker and synchronizes gradients — it addresses throughput across machines, not the case where the model does not fit on one device. Option B is wrong because ParameterServerStrategy also replicates the model across workers with parameter servers coordinating updates; it does not split layers within a single model. Option D is wrong because MirroredStrategy performs data parallelism — it creates a full replica of the model on every GPU, which is impossible when the model exceeds single-GPU memory.

428
MCQeasy

A marketing team wants to build a model that predicts whether a customer will click on an ad, using a dataset in BigQuery. They have limited ML expertise and want to avoid writing complex code. They decide to use BigQuery ML with a logistic regression model. Which SQL statement should they use to create the model?

A.CREATE MODEL `project.dataset.model` OPTIONS(model_type='logistic_reg', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
B.CREATE MODEL `project.dataset.model` OPTIONS(model_type='linear_reg', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
C.CREATE MODEL `project.dataset.model` OPTIONS(model_type='logistic_reg') AS SELECT * FROM `project.dataset.training_data`
D.CREATE MODEL `project.dataset.model` OPTIONS(model_type='kmeans', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
AnswerA

This statement correctly creates a logistic regression model in BigQuery ML. It specifies the model type as logistic_reg and identifies the label column as 'clicked' using the input_label_cols option. The SELECT statement provides the training data, which includes both features and the label. BigQuery ML will automatically use all other columns as features, making it suitable for users with limited ML expertise.

Why this answer

The correct SQL statement creates a logistic regression model in BigQuery ML with the proper model type and label column specification. Logistic regression is ideal for binary classification tasks such as predicting ad clicks. The input_label_cols option identifies the target variable, and the SELECT statement supplies the training data.

This approach allows users with limited ML expertise to build a model using only SQL.

Exam trap

The trap here is confusing linear regression with logistic regression for binary classification tasks.

429
Multi-Selectmedium

An ML pipeline must run a set of preprocessing tasks for each data shard in parallel. Which KFP SDK features should they use to implement this? (Choose two.)

Select 2 answers
A.dsl.ParallelFor
B.dsl.PipelineParam
C.dsl.Collected
D.dsl.Condition
E.dsl.ExitHandler
AnswersA, C

dsl.ParallelFor iterates over the shard list at pipeline-compile time, generating one preprocessing task instance per shard that KFP schedules concurrently. This directly satisfies the requirement to run preprocessing tasks for each data shard in parallel.

Why this answer

dsl.ParallelFor (A) is correct because it is the KFP SDK construct that iterates over a list (such as the set of data shards) and fans out the loop body into parallel task executions, which is exactly what is needed to run preprocessing per shard concurrently. dsl.Collected (C) is correct because it is used with dsl.ParallelFor to gather the outputs of all parallel iterations into a single list, allowing downstream pipeline steps to consume the aggregated results of the per-shard preprocessing. dsl.PipelineParam (B) is not the mechanism for parallel iteration; it only represents a runtime parameter passed into a pipeline or component. dsl.Condition (D) implements conditional branching (if/else) rather than parallel fan-out, and dsl.ExitHandler (E) defines cleanup logic that runs when a scope exits, neither of which provides the required parallel execution over shards.

Exam trap

In the Google PMLE exam, note that dsl.ParallelFor is used for parallel iteration over data shards, and dsl.Collected gathers outputs from all iterations. Candidates often confuse these with dsl.Condition (for branching) or dsl.PipelineParam (for parameters), so carefully read whether the question asks for parallel processing or conditional logic.

430
MCQeasy

What is the primary benefit of using a centralised model registry in MLOps?

A.Governance and version control of models
B.Better hyperparameter tuning
C.Faster model training
D.Automatic model deployment
AnswerA

A centralised model registry provides governance and version control, tracking each model's lineage, approvals, and deployment stage. This gives teams a single authoritative source for model artefacts, satisfying auditability and reproducibility requirements across the MLOps lifecycle.

Why this answer

A centralised model registry provides governance, versioning, and lineage tracking, enabling collaboration and auditability.

431
MCQmedium

Your ML pipeline uses Vertex AI Feature Store to serve features for online predictions. You need to monitor the freshness of features in the online store. Which approach is most effective?

A.Set up a Cloud Monitoring alert for feature store entity count.
B.Schedule a nightly BigQuery batch job to compare feature values.
C.Create a custom metric in Cloud Monitoring that tracks the time since last feature update, and set an alert threshold.
D.Enable detailed audit logs in Feature Store and export to BigQuery.
AnswerC

Feature freshness is a temporal property not exposed by default Vertex AI Feature Store metrics. A custom Cloud Monitoring metric recording time since the last feature update, paired with an alert threshold, directly detects stale online values before they degrade predictions.

Why this answer

Cloud Monitoring custom metrics allow you to track the timestamp of the last feature update in Vertex AI Feature Store and set an alert threshold for staleness. This directly measures feature freshness, which is critical for online predictions where stale features can degrade model accuracy. Other options either measure unrelated metrics (entity count), are too slow (nightly batch), or focus on auditing rather than real-time monitoring.

Exam trap

The trap here is that candidates confuse monitoring entity count (a capacity metric) with freshness, or assume that batch comparison or audit logs provide real-time monitoring, when only a custom staleness metric with alerting directly addresses the requirement.

How to eliminate wrong answers

Option A is wrong because monitoring the entity count in the feature store tracks the number of stored feature values, not the time since they were last updated, so it cannot detect staleness. Option B is wrong because a nightly BigQuery batch job introduces latency of up to 24 hours, making it unsuitable for real-time freshness monitoring required for online predictions. Option D is wrong because enabling detailed audit logs and exporting to BigQuery provides an after-the-fact record of changes but does not offer real-time alerting on feature staleness.

432
MCQmedium

A data science team wants to version control their datasets along with code using Git. They need a tool that integrates with Git and tracks changes to large data files. Which tool should they use?

A.BigQuery table snapshots
B.Delta Lake
C.Git LFS
D.DVC
AnswerD

DVC stores large dataset files outside Git while committing small metafiles that capture content hashes and versions, so Git tracks dataset changes without bloating the repository. This pointer-based mechanism satisfies the requirement to version data alongside code using Git.

Why this answer

DVC (Data Version Control) is purpose-built to integrate with Git and version large datasets and ML models by storing metadata and pointers in Git while keeping the actual data in remote storage. It provides commands like dvc add, dvc push, and dvc pull that mirror Git workflows, making it the right choice for versioning datasets alongside code.

Exam trap

The trap is confusing Git LFS with DVC — both handle large files, but only DVC provides dataset versioning, pipeline reproducibility, and remote storage abstraction for ML workflows.

How to eliminate wrong answers

Option A is wrong because BigQuery table snapshots version BigQuery tables for time travel and recovery, not Git-integrated dataset versioning for arbitrary files. Option B is wrong because Delta Lake provides ACID transactions and time travel on data lakes but is not a Git-integrated version control tool for datasets. Option C is wrong because Git LFS tracks large files via pointers but is designed for binary assets, not dataset pipelines with reproducibility, metrics, and experiment tracking like DVC.

433
MCQhard

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to use a low-code solution that can automatically handle feature engineering and model selection. They also need to deploy the model for online predictions with low latency. Which Google Cloud service should they use?

A.Vertex AI AutoML Tables
B.Vertex AI Pipelines with Kubeflow
C.Cloud Functions with a custom Python script
D.BigQuery ML with a logistic regression model
AnswerA

Vertex AI AutoML Tables automatically performs feature engineering, including one-hot encoding, normalization, and feature selection, and selects the best model architecture. It supports online prediction with low latency when deployed to an endpoint. This aligns perfectly with the need for a low-code solution that handles feature engineering and provides real-time fraud detection. It is the most suitable choice for this scenario.

Why this answer

Vertex AI AutoML Tables is the correct choice because it provides automatic feature engineering and model selection, and supports low-latency online predictions when deployed to an endpoint. It is designed for tabular data like financial transactions and requires minimal coding. Other options either lack automatic feature engineering, require significant coding, or are not intended for building predictive models, making AutoML Tables the best fit for real-time fraud detection.

Exam trap

The trap here is assuming that BigQuery ML models can be directly used for low-latency online predictions without additional deployment steps, when in fact Vertex AI AutoML Tables offers a more streamlined path.

434
Multi-Selectmedium

Which THREE actions should be taken to automate a machine learning pipeline using Cloud Build and Vertex AI?

Select 3 answers
A.Write a cloudbuild.yaml that builds a training container and submits a Vertex AI PipelineJob
B.Use Cloud Functions to retrain the model each time a build completes
C.Set up a Cloud Scheduler job to poll for new build artifacts
D.Define the training and deployment steps in a Vertex AI Pipeline and submit it from Cloud Build
E.Configure a Cloud Build trigger to run on commits to the source repository
AnswersA, D, E

Cloud Build executes declarative build steps from cloudbuild.yaml, so this file can both build the training container image and invoke the Vertex AI API to submit a PipelineJob, providing the end-to-end automation the scenario demands.

Why this answer

Option A is correct because a cloudbuild.yaml file is the declarative configuration Cloud Build uses to define build steps, and one of those steps can build a custom training container image (e.g., via docker build) and then submit a Vertex AI PipelineJob using the gcloud ai pipelines run command or the Vertex AI SDK. Option D is correct because Vertex AI Pipelines is the native service for orchestrating ML workflow steps such as training, evaluation, and deployment as a directed acyclic graph, and Cloud Build can invoke that pipeline as part of CI/CD, giving a fully automated, reproducible ML pipeline. Option E is correct because a Cloud Build trigger bound to the source repository (e.g., a GitHub or Cloud Source Repositories trigger on push or pull request) is what actually initiates the automated build/pipeline run on commits, which is the core of CI/CD automation.

Option B is not appropriate because Cloud Functions is not the intended mechanism for retraining models on build completion; Cloud Build triggers or Pub/Sub notifications are the correct integration points, and this adds unnecessary complexity. Option C is not appropriate because polling for build artifacts with Cloud Scheduler is an anti-pattern; event-driven triggers (Cloud Build triggers or Pub/Sub) are preferred over scheduled polling for reacting to new builds.

Exam trap

Google Cloud often tests the distinction between event-driven triggers (Cloud Build triggers, Pub/Sub) and polling mechanisms (Cloud Scheduler, Cloud Functions) — the trap here is that candidates may think polling or separate functions are needed for automation, when in fact Cloud Build's native triggers and pipeline submission are the correct, integrated approach.

435
MCQmedium

A company uses Vertex AI Vector Search (Matching Engine) for a product recommendation system. The product embeddings are updated hourly. Which index update method should they use to ensure low latency for new items?

A.Batch rebuild the index every hour
B.Use streaming updates to add new embeddings incrementally
C.Create a new index each hour and swap endpoints
D.Use brute-force index to simplify updates
AnswerB

Streaming updates let the index ingest new product embeddings incrementally without a full rebuild, so hourly-refreshed items become searchable almost immediately. This satisfies the low-latency requirement for new items, whereas batch rebuilds would delay their availability.

Why this answer

Vertex AI Vector Search supports streaming updates, allowing new embeddings to be added incrementally without rebuilding the entire index. This ensures low latency for new items by making them searchable almost immediately after update, which is critical for hourly refresh cycles where batch rebuilds would introduce significant delay.

Exam trap

The trap here is that candidates often assume batch rebuilds are the only reliable method for consistency, overlooking that streaming updates in Vertex AI Vector Search are designed specifically for low-latency incremental ingestion without sacrificing search quality.

How to eliminate wrong answers

Option A is wrong because batch rebuilding the index every hour incurs high latency and computational cost, as the entire index must be reconstructed from scratch, delaying availability of new items. Option C is wrong because creating a new index each hour and swapping endpoints is inefficient and introduces downtime during the swap, plus it requires managing multiple index versions unnecessarily. Option D is wrong because brute-force indices do not simplify updates; they perform exhaustive linear scans, which are slow and unscalable for large embedding sets, and they lack the optimized approximate nearest neighbor (ANN) search that Vector Search provides.

436
Multi-Selecteasy

An ML team is deploying a model to Vertex AI for the first time. Which THREE are best practices for scaling from prototype to production?

Select 3 answers
A.Manually scale instances based on historical traffic patterns.
B.Store all features in a Feature Store for consistency.
C.Use a single large instance to simplify management.
D.Monitor model performance for drift and accuracy degradation.
E.Automate model retraining and deployment using Vertex AI Pipelines.
AnswersB, D, E

A Feature Store provides a centralised, versioned repository so training and serving pipelines draw identical feature values, eliminating training-serving skew. This directly satisfies the production-consistency requirement when scaling beyond a prototype, where ad-hoc feature engineering typically diverges between environments.

Why this answer

Option B is correct because storing features in a Vertex AI Feature Store ensures the same feature transformations are used at training and serving time, preventing training-serving skew and enabling feature reuse and consistency across models. Option D is correct because production models degrade over time due to data drift and concept drift, so continuous monitoring of prediction quality, accuracy, and drift metrics is essential to detect degradation and trigger remediation. Option E is correct because automating retraining and deployment with Vertex AI Pipelines creates a repeatable, auditable MLOps workflow that reduces manual error and lets the model adapt to new data efficiently.

Option A is not a best practice because manual scaling based on historical patterns is reactive and error-prone; production should use autoscaling driven by live metrics such as CPU, GPU, or request load. Option C is not a best practice because a single large instance creates a single point of failure and does not scale horizontally, whereas Vertex AI supports distributed serving with multiple replicas for resilience and elasticity.

Exam trap

Google Cloud often tests the misconception that manual scaling or single-instance architectures are simpler and more reliable, but the PMLE exam emphasizes automated, resilient, and consistent practices like autoscaling and feature stores for production ML workloads.

437
MCQeasy

A team has developed a prototype of a recommendation model using a small dataset on a single VM. They need to scale to a larger dataset for production training. They plan to use Vertex AI training with a custom container. What is the best practice for handling the increased data volume?

A.Increase the batch size to maximum.
B.Use TFRecord format and streaming reads.
C.Store all data in memory before training.
D.Use a single powerful VM with high memory.
AnswerB

TFRecord with streaming reads avoids loading the full dataset into memory, letting Vertex AI training workers read shards incrementally from Cloud Storage. This satisfies the increased data volume constraint, whereas downloading the entire dataset or using in-memory NumPy arrays would exhaust worker memory during distributed training.

Why this answer

When scaling to a larger dataset for production training on Vertex AI with a custom container, the best practice is to use TFRecord format and streaming reads. TFRecord is a binary format optimized for TensorFlow, enabling efficient data serialization and streaming from Cloud Storage without loading the entire dataset into memory. This approach supports large-scale distributed training and reduces I/O bottlenecks.

Exam trap

PMLE often tests the misconception that simply scaling up hardware (more memory, bigger batch) solves data volume issues, but the exam expects knowledge of efficient data formats and streaming for distributed training.

How to eliminate wrong answers

Option A is wrong because increasing batch size to maximum can lead to out-of-memory errors and does not address data loading efficiency; it may also harm model convergence. Option C is wrong because storing all data in memory is not scalable for large datasets and will fail as data volume grows. Option D is wrong because using a single powerful VM with high memory does not solve the scalability issue and is cost-inefficient; distributed training with streaming is preferred for production.

438
MCQmedium

An ML engineer is designing a pipeline that should run only when new training data arrives in a Cloud Storage bucket. Which event-driven approach should they use to trigger the Vertex AI Pipeline?

A.Use Cloud Storage Pub/Sub notifications to send events to a Cloud Function that triggers the pipeline.
B.Use Cloud Tasks to queue a pipeline run whenever a new file is uploaded.
C.Configure the pipeline to run on a schedule and check for new data inside the pipeline.
D.Set up a Cloud Scheduler job that runs every minute to check for new files.
AnswerA

Cloud Storage Pub/Sub notifications emit an event whenever an object lands in the bucket, and a Cloud Function subscribed to that topic invokes the Vertex AI Pipeline. This satisfies the stem's requirement to run only when new training data actually arrives, avoiding polling.

Why this answer

The best approach is to use Cloud Storage notifications via Pub/Sub, then a Cloud Function that receives the event and calls the Vertex AI API to create a pipeline job. This is a common event-driven pattern. Cloud Scheduler is for scheduled triggers, not event-driven.

Cloud Tasks and Cloud Run are not typically used for this purpose.

439
MCQeasy

Your company deploys batch prediction jobs using Vertex AI Batch Prediction. You need to monitor the jobs for failures and performance. What is the recommended approach?

A.Use Cloud Logging to export batch prediction logs and create log-based metrics.
B.Set up email alerts in the Vertex AI console for failed jobs.
C.Use Cloud Monitoring to create custom dashboards and alerts based on Vertex AI batch prediction metrics.
D.Enable the Recommender to get optimization suggestions for batch jobs.
AnswerC

Vertex AI Batch Prediction emits metrics into Cloud Monitoring, so custom dashboards and alerting policies can track job failures and performance without bespoke instrumentation. This directly satisfies the requirement to monitor batch jobs for failures and performance using the platform's native observability surface.

Why this answer

Cloud Monitoring (formerly Stackdriver) is the native Google Cloud service for collecting, visualizing, and alerting on metrics from Vertex AI, including batch prediction job success rates, latency, and resource utilization. It provides pre-built dashboards and the ability to create custom alerts, making it the recommended approach for monitoring failures and performance in a centralized, scalable way.

Exam trap

Google Cloud often tests the misconception that Cloud Logging is the primary monitoring tool for metrics, when in fact Cloud Monitoring is the dedicated service for metrics and alerting, while Cloud Logging is for logs and log-based metrics only.

How to eliminate wrong answers

Option A is wrong because Cloud Logging is designed for log data, not structured metrics; while you could create log-based metrics from batch prediction logs, this is an indirect, less efficient method that lacks the pre-built performance metrics and alerting capabilities of Cloud Monitoring. Option B is wrong because email alerts in the Vertex AI console are not a native feature; Vertex AI does not provide a built-in email alerting mechanism for job failures—alerts must be configured through Cloud Monitoring or Cloud Logging. Option D is wrong because the Recommender provides optimization suggestions (e.g., machine type, resource allocation) but does not monitor job failures or performance in real time; it is a post-hoc analysis tool, not a monitoring solution.

440
MCQeasy

A retail company has deployed a machine learning model using Vertex AI Endpoints to predict inventory demand. The model was trained on data from the past two years and has been in production for six months. The team has enabled Vertex AI Model Monitoring to track prediction drift with an alert threshold of 0.2. Last week, they received an alert that the prediction drift score reached 0.35, exceeding the threshold. The engineer checks the monitoring dashboard and sees that the distribution of predictions has shifted noticeably compared to the training data. The engineer also notices that the model's accuracy metrics, computed from weekly ground truth data, have remained within acceptable range. What should the engineer do first?

A.Investigate the input feature distributions for the recent serving requests to identify if data drift is the underlying cause of the prediction drift.
B.Increase the prediction drift alert threshold to 0.4 to reduce the number of false alerts.
C.Retrain the model using the latest three months of data to incorporate recent trends.
D.Roll back to an earlier model version that had lower prediction drift.
AnswerA

Prediction drift with stable accuracy points to changed inputs rather than model degradation, so examining recent serving feature distributions confirms whether data drift underlies the shift. This satisfies the stem's need to identify the root cause before retraining or altering thresholds.

Why this answer

Prediction drift is a downstream symptom, and the first diagnostic step is to determine whether it stems from input data drift, concept drift, or a benign shift in the input mix. Since accuracy on weekly ground truth is still acceptable, the model itself is not yet degraded, so the engineer should investigate the recent serving feature distributions against the training baseline before taking any corrective action. Vertex AI Model Monitoring can surface feature-level drift metrics that pinpoint which inputs moved.

Exam trap

The trap here is conflating prediction drift with model degradation and jumping to remediation (retrain, rollback, or raise threshold) instead of first diagnosing whether input data drift is the root cause.

How to eliminate wrong answers

Option B is wrong because raising the threshold to 0.4 merely suppresses the alert without diagnosing or addressing the underlying shift, which could mask a real degradation. Option C is wrong because retraining is a remediation step that should only follow root-cause analysis; retraining on three months of data could also introduce new bias or overfit to a transient trend. Option D is wrong because rolling back to an earlier model version does not fix the drift (the input data has changed, not the model) and discards a model that still meets accuracy targets.

441
MCQmedium

A company is using Vertex AI Vizier for hyperparameter tuning of a model with 5 integer hyperparameters, each with a range of 10-100. They have a budget of 50 trials and want to maximize the chance of finding the best configuration. Which Vizier algorithm should they use?

A.Grid search
B.Simulated annealing
C.Bayesian optimization (GP bandit)
D.Random search
AnswerC

Bayesian optimisation with a Gaussian process bandit models the objective surface and selects trials that balance exploration against exploitation, converging efficiently within a limited budget. With 50 trials across five integer parameters, it maximises the chance of locating the best configuration.

Why this answer

Bayesian optimization (GP bandit) is Vizier's default and most sample-efficient algorithm, using a Gaussian Process surrogate model to balance exploration and exploitation across the 5-dimensional hyperparameter space. With only 50 trials over a large search space, it converges on promising regions far faster than uninformed methods. This makes it the best choice for maximizing the chance of finding the optimal configuration within a limited budget.

Exam trap

The trap here is assuming that random search is 'good enough' for high-dimensional tuning or that grid search guarantees coverage — PMLE often tests whether candidates understand that Bayesian optimization is the sample-efficient choice when the trial budget is constrained.

How to eliminate wrong answers

Option A is wrong because grid search scales exponentially with dimensionality and would exhaust the 50-trial budget long before covering a meaningful fraction of the 5-dimensional space. Option B is wrong because simulated annealing is a local search heuristic not natively offered as a Vizier algorithm and is poorly suited to expensive black-box hyperparameter evaluations. Option D is wrong because random search ignores prior trial results and requires many more samples than Bayesian optimization to find competitive configurations.

442
MCQmedium

An ML engineer wants to containerize a custom training script and use it as a component in a Vertex AI Pipeline. The component should accept a dataset URI and a learning rate parameter, and output a trained model artifact. Which approach should the engineer use to define the component?

A.Use a pre-built Google Cloud Pipeline Component for Vertex AI Training with custom container configuration.
B.Use ContainerComponent from kfp.v2.components to define the container, its inputs, and outputs.
C.Define a Python function component with @dsl.component and include the container code inline.
D.Use the importer component to import the script and then run it as a task.
AnswerB

ContainerComponent from kfp.v2.components lets the engineer declare a custom container image plus typed inputs (dataset URI, learning rate) and outputs (model artifact), satisfying the stem's requirement to containerise a custom training script as a pipeline component.

Why this answer

ContainerComponent from kfp.v2.components allows you to define a custom container component by specifying the container image, command, inputs, and outputs directly. This is the appropriate approach when you have a custom training script that you want to containerize and use as a component in a Vertex AI Pipeline, as it gives you full control over the container configuration and artifact handling.

Exam trap

Candidates often confuse Python function components (@dsl.component) with container components. The trap here is that they may think a Python function component can containerize a custom script, but it cannot directly specify a container image and artifact outputs like ContainerComponent does.

How to eliminate wrong answers

Option A is wrong because pre-built Google Cloud Pipeline Components for Vertex AI Training are designed for standard training jobs with built-in algorithms or custom containers, but they do not allow you to define custom inputs and outputs as artifacts in the same declarative way as ContainerComponent; they are more rigid and less suited for a fully custom component with a dataset URI and learning rate parameter. Option C is wrong because @dsl.component is used for Python function components that run Python code directly, not for containerized components; including container code inline would mix the container definition with Python function logic, which is not the intended use and would not properly handle container image specification and artifact outputs. Option D is wrong because the importer component is used to import existing artifacts (like models or datasets) into the pipeline, not to run a training script; it cannot execute a custom training script or produce a trained model artifact from scratch.

443
MCQhard

You are scaling a prototype recommendation model to production on Vertex AI. The model is trained with a custom container and uses a large embedding table that must be updated frequently. You need to serve predictions with low latency and also support online updates to the embeddings without redeploying the model. Which Vertex AI feature should you use?

A.Deploy the model to a Vertex AI Endpoint with a custom container that loads the embedding table from a shared store such as Cloud Storage or a database at prediction time, and update the shared store independently.
B.Deploy the model to a Vertex AI Endpoint with a pre-built container and use Vertex AI Feature Store to serve the embeddings as features at prediction time.
C.Deploy the model to a Vertex AI Endpoint and use the endpoint's built-in support for model updates through Vertex AI Model Registry versioning.
D.Use Vertex AI Batch Prediction with a large batch size so that embeddings are computed in bulk and can be updated between batches.
AnswerA

A custom container can load or refresh the embedding table from an external store, allowing the embeddings to be updated without redeploying the model. This supports low-latency serving if the store is fast and the container caches or periodically refreshes the table. It is the standard way to combine a custom model with frequently changing embeddings on Vertex AI.

Why this answer

To serve low-latency predictions while allowing online embedding updates, the model should run in a custom container that can read and refresh the embedding table from an external store. This keeps the model deployed while the embeddings change independently. The other options either do not update internal embeddings, require redeployment, or are batch-oriented and therefore unsuitable for low-latency online serving.

Exam trap

The trap here is confusing feature serving with model-internal embedding updates, or assuming that Model Registry versioning gives you online weight updates without redeployment.

444
MCQmedium

A team deploys a model using Vertex AI and wants to monitor for concept drift. What should they track?

A.Number of prediction requests
B.Model prediction latency
C.Changes in input data distribution
D.Changes in the relationship between inputs and outputs
AnswerD

Concept drift is defined by a change in the statistical relationship between input features and target outputs, so tracking that mapping is the only signal that captures it. Feature distribution shifts alone indicate data drift, not concept drift.

Why this answer

Concept drift refers to a change in the underlying relationship between the input features and the target variable over time, which degrades model performance. In Vertex AI, monitoring this requires tracking the statistical relationship between inputs and outputs (e.g., via prediction residuals or model performance metrics), not just the input distribution alone. Option D correctly identifies this need, as concept drift is fundamentally about the input-output mapping shifting, even if the input distribution remains stable.

Exam trap

Google Cloud often tests the distinction between data drift (input distribution changes) and concept drift (input-output relationship changes), and the trap here is that candidates confuse the two, picking Option C because they think monitoring input data is sufficient for detecting all model degradation.

How to eliminate wrong answers

Option A is wrong because the number of prediction requests measures traffic volume, not data or concept drift; it is a scaling or operational metric, not a model quality metric. Option B is wrong because prediction latency measures inference speed, which is a performance indicator unrelated to the statistical properties of data or model relationships. Option C is wrong because changes in input data distribution represent data drift (covariate shift), not concept drift; while data drift can cause concept drift, monitoring only input distribution misses shifts in the input-output relationship that occur without distributional changes.

445
Multi-Selecteasy

Which TWO are best practices for building ML pipelines on Vertex AI Pipelines?

Select 2 answers
A.Store all trained models in Cloud Storage without versioning
B.Use Cloud Build as the pipeline orchestrator
C.Use a container-based approach for each component
D.Define pipelines using the Kubeflow Pipelines SDK
E.Use Cloud Composer as the primary pipeline tool
AnswersC, D

Containerising each component gives every pipeline step a reproducible, dependency-isolated runtime, so Vertex AI Pipelines can execute steps consistently across environments. This satisfies the portability and reproducibility requirements of production ML pipelines, since each component runs as a self-contained image.

Why this answer

Option C is correct because Vertex AI Pipelines executes each pipeline step as a containerized component, so packaging each component as a Docker container (with its dependencies and code) ensures reproducibility, portability, and consistent execution across environments. Option D is correct because Vertex AI Pipelines is built on Kubeflow Pipelines, and the Kubeflow Pipelines SDK (the `kfp` package, e.g., `@dsl.pipeline` and `@component` decorators) is the supported, idiomatic way to define, compile, and submit pipelines to Vertex AI. Option A is wrong because storing models without versioning breaks lineage, rollback, and reproducibility; models should be registered in Vertex AI Model Registry with versioning.

Option B is wrong because Cloud Build is a CI/CD build service, not a pipeline orchestrator for ML workflows. Option E is wrong because Cloud Composer (managed Apache Airflow) is a general workflow orchestrator, not the primary tool for authoring and running Vertex AI Pipelines.

Exam trap

Google Cloud often tests the distinction between general-purpose orchestration tools (Cloud Composer, Cloud Build) and ML-specific pipeline services (Vertex AI Pipelines), expecting candidates to recognize that container-based components and the Kubeflow Pipelines SDK are the correct building blocks for ML pipelines on Vertex AI.

446
MCQeasy

You are deploying a scikit-learn model to Vertex AI for online prediction. The model expects a JSON payload with a single feature vector. You need to ensure the endpoint can handle bursts of traffic up to 1000 requests per second while maintaining low latency. What should you do?

A.Deploy the model to a Vertex AI Endpoint with a single machine type and enable autoscaling based on CPU utilization.
B.Deploy the model to a Vertex AI Endpoint with a fixed number of replicas equal to the expected peak traffic.
C.Use a batch prediction job to process incoming requests every minute.
D.Deploy the model to a Vertex AI Endpoint with a GPU machine type to accelerate inference.
AnswerA

Vertex AI Endpoint supports autoscaling, which automatically adjusts the number of replicas based on traffic. Setting a minimum and maximum replica count with CPU utilization as the metric allows the endpoint to handle bursts while keeping latency low. This is the standard and recommended approach for scaling online predictions.

Why this answer

Vertex AI Endpoints provide autoscaling to dynamically adjust replicas based on traffic. By enabling autoscaling with CPU utilization as the metric, the endpoint can scale out during bursts and scale in during low traffic, ensuring low latency and cost efficiency. This is the correct way to handle variable online prediction traffic.

Exam trap

The trap here is thinking that provisioning for peak load or using GPUs will solve latency issues, when the key is elastic scaling based on actual demand.

447
Multi-Selectmedium

Which TWO are best practices when deploying AutoML models to production?

Select 2 answers
A.Monitor for data drift
B.Train the model on a disk to reduce latency
C.Enable Vertex AI Explainability
D.Deploy on sole-tenant nodes
E.Use TPUs for model serving
AnswersA, C

Data drift can degrade performance; monitoring is essential.

Why this answer

Monitoring for data drift (Option A) is a best practice because production models can degrade over time as the statistical properties of input data change. Vertex AI provides a Model Monitoring service that automatically detects skew and drift by comparing serving data distribution against training data distribution, triggering alerts when anomaly thresholds are breached. This ensures model reliability and performance in production.

Exam trap

Google Cloud often tests the misconception that TPUs are suitable for model serving, but TPUs are optimized for training and not supported for Vertex AI AutoML serving, which uses CPUs or GPUs for inference.

448
Multi-Selecthard

A company is deploying a Vertex AI pipeline that trains a model and then runs a custom evaluation component. The evaluation component must only run if the training component succeeds and the model's accuracy exceeds a threshold. The pipeline must also support retries for transient errors in the training component. The engineer needs to configure the pipeline to meet these requirements. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Configure the pipeline to run the evaluation component in parallel with training to reduce latency.
B.Set `enable_caching=False` on the evaluation component to ensure it always runs after training.
C.Set the `retry` policy on the training component to retry on specific exit codes or exceptions.
D.Use a `dsl.Condition` to wrap the evaluation component and check the accuracy metric against the threshold.
E.Use a `dsl.ExitHandler` to catch failures in the training component and retry it manually.
AnswersC, D

Vertex AI Pipelines supports retry policies on individual components. By configuring a retry policy with a maximum retry count and backoff, the training component can automatically retry on transient failures such as resource exhaustion or network timeouts. This satisfies the requirement to support retries for transient errors. The retry policy can be specified using the `retry` argument when defining the component or task, and it applies only to that task.

Why this answer

The evaluation component must run only when training succeeds and accuracy exceeds a threshold, so a dsl.Condition is used to gate its execution based on the accuracy metric. To handle transient errors in training, a retry policy on the training component is configured. These two actions together meet the requirements.

Other options either do not enforce the condition, do not provide retries, or are not applicable.

Exam trap

The trap here is confusing caching with conditional execution, or assuming that an ExitHandler can serve as a retry mechanism.

449
MCQeasy

A team is using Vertex AI Pipelines to automate their ML workflow. They want to ensure that pipeline runs are reproducible and that artifacts are tracked. Which feature should they use?

A.Vertex AI Feature Store
B.Vertex AI Experiments
C.Vertex AI Model Registry
D.Vertex AI Endpoints
AnswerB

Vertex AI Experiments records parameters, metrics and artefacts for each pipeline run, giving lineage and comparison across executions. This satisfies the reproducibility and artefact-tracking requirement, unlike raw pipeline execution alone, which does not persist experiment-level metadata.

Why this answer

Vertex AI Experiments is the correct feature because it captures parameters, metrics, and artifacts for each pipeline run, enabling reproducibility and lineage tracking. This directly supports the team's need to ensure runs are reproducible and artifacts are tracked, as Experiments automatically logs metadata for every execution.

Exam trap

The trap here is that candidates confuse artifact tracking with model management or deployment features, leading them to select Model Registry or Endpoints instead of recognizing that Experiments provides the run-level metadata and lineage required for reproducibility.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store is designed for managing and serving feature data for ML models, not for tracking pipeline runs or artifacts. Option C is wrong because Vertex AI Model Registry focuses on managing model versions and deployment, not on capturing run-level metadata or artifact lineage. Option D is wrong because Vertex AI Endpoints are for deploying models to serve predictions, not for tracking reproducibility or artifacts in pipeline runs.

450
MCQeasy

An ML engineer needs to track the costs incurred by Vertex AI prediction endpoints. Which tool should they use to set budget alerts and monitor spending?

A.Vertex AI Model Monitoring
B.Cloud Logging
C.Cloud Monitoring with custom metrics
D.Google Cloud Budgets & Alerts
AnswerD

Google Cloud Budgets & Alerts scopes budgets to projects, folders, or labels, and triggers notifications when actual or forecast spend crosses thresholds. Labelling the Vertex AI endpoints lets the engineer isolate their prediction costs, satisfying the requirement to set budget alerts and monitor spending on those endpoints specifically.

Why this answer

Google Cloud Budgets & Alerts is the dedicated billing-scoped tool that lets you define a budget amount (or filter by project/service/label) and attach alert thresholds at percentages of actual or forecasted spend. It is the only option that operates on billing data and can trigger Pub/Sub notifications or email alerts when Vertex AI endpoint costs cross a threshold. Model Monitoring, Cloud Logging, and Cloud Monitoring track operational metrics, not monetary spend.

Exam trap

The trap here is confusing operational monitoring tools (Cloud Monitoring, Model Monitoring) with billing-scoped tools — candidates pick Cloud Monitoring because it 'alerts,' but only Budgets & Alerts works on cost data.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Monitoring tracks model drift, skew, and prediction quality metrics — it has no visibility into billing or cost data. Option B is wrong because Cloud Logging captures log entries (audit, application, platform logs) and cannot aggregate or alert on dollar amounts. Option C is wrong because Cloud Monitoring with custom metrics can alert on operational telemetry you push, but it does not natively ingest Cloud Billing cost data as a metric for budget thresholding.

Page 5

Page 6 of 11

Page 7

All pages