Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 451–525

775 questions total · 11pages · All types, answers revealed

Page 6

Page 7 of 11

Page 8
451
MCQeasy

You need to set up monitoring for a Vertex AI model that serves predictions in real-time. The model is expected to have a latency SLA of under 100ms. Which metric should you configure an alert on to ensure the SLA is met?

A.p50 latency of prediction requests
B.Prediction drift score
C.p99 latency of prediction requests
D.Number of prediction requests per second
AnswerC

A latency SLA is a tail-latency commitment, so p99 captures the slowest 1% of prediction requests that breach the 100ms threshold. Alerting on p99 detects SLA violations that average latency would hide behind fast responses.

Why this answer

P99 latency measures the worst-case latency experienced by 99% of requests, which is the standard metric for enforcing a strict SLA like under 100ms. Monitoring p99 ensures that even the slowest 1% of requests do not violate the threshold, providing a robust guarantee for real-time predictions.

Exam trap

Google Cloud often tests the misconception that median (p50) latency is sufficient for SLAs, but the trap is that SLAs require tail-latency guarantees (p99 or p999) to catch performance outliers that violate the threshold.

How to eliminate wrong answers

Option A is wrong because p50 latency (median) ignores the tail latency, meaning half of the requests could exceed 100ms without triggering an alert, failing the SLA. Option B is wrong because prediction drift score measures changes in model input/output distributions over time, not latency, and is irrelevant for SLA compliance. Option D is wrong because the number of prediction requests per second (throughput) does not measure individual request latency; high throughput can occur even if latency spikes above 100ms.

452
MCQhard

An ML engineer is authoring a Vertex AI pipeline where a custom training component must read a dataset from a BigQuery table and write the trained model to a Cloud Storage bucket. The engineer wants the component to be reusable across projects and environments without hardcoding project IDs or bucket names. Which design should the engineer use?

A.Use the Vertex AI SDK's aiplatform.init() with no arguments inside the component and rely on the default project and bucket.
B.Read the project ID and bucket name from environment variables set inside the component's container image at build time.
C.Hardcode the production project ID and bucket name in the component, then override them with a pipeline-level parameter only when running in non-production.
D.Pass the BigQuery table URI and the Cloud Storage output URI as component input parameters, and let the pipeline caller supply them at runtime.
AnswerD

Making the dataset URI and output URI component inputs parameterizes the component so the same definition can be reused across projects and environments. The pipeline caller supplies the concrete values at runtime, which is the standard Vertex AI Pipelines pattern for portability. Hardcoding or deriving them inside the component would tie the component to one project and defeat reusability.

Why this answer

Component reusability in Vertex AI Pipelines comes from parameterizing all environment-specific values as inputs. Passing the BigQuery table URI and Cloud Storage output URI as component inputs lets the same component definition run in any project or environment, with the pipeline caller providing concrete values. The other options embed environment-specific values in the container or rely on implicit defaults, which breaks portability and can cause cross-environment mistakes.

Exam trap

The trap here is thinking that relying on default credentials and default project inside a component is equivalent to passing explicit inputs, when implicit defaults make the component environment-dependent.

453
MCQeasy

A company has deployed a TensorFlow model on Vertex AI Prediction for real-time inference. They notice that during peak hours, the prediction latency increases significantly, and some requests time out. The model requires GPU acceleration. Which action should they take to reduce latency and avoid timeouts?

A.Enable autoscaling with min replicas set to the base load and max replicas set to handle peak load, and ensure GPU quota is sufficient.
B.Switch to a larger machine type with more vCPUs.
C.Increase the number of replicas in the Vertex AI Prediction endpoint statically to handle peak load.
D.Use Cloud Functions to invoke the model asynchronously.
AnswerA

Enabling autoscaling with appropriate min and max replicas allows the endpoint to dynamically scale up during peak traffic and scale down during low traffic, ensuring sufficient GPU resources to handle the load without manual intervention. Ensuring adequate GPU quota is also critical to prevent resource exhaustion.

Why this answer

Enabling autoscaling with appropriate min and max replicas allows the endpoint to dynamically scale up during peak traffic and scale down during low traffic, ensuring sufficient GPU resources to handle the load without manual intervention. Ensuring adequate GPU quota is also critical to prevent resource exhaustion. Option B is wrong because switching to a larger machine type with more vCPUs does not address GPU-bound inference latency; the bottleneck is GPU acceleration, not CPU.

Option C is wrong because statically increasing replicas leads to over-provisioning and waste during off-peak hours, and cannot react quickly to sudden traffic spikes. Option D is wrong because Cloud Functions are serverless and do not support GPU acceleration; they add additional latency and are not suitable for real-time GPU inference.

454
MCQmedium

A data scientist has a TensorFlow 2.x model trained on a single GPU. They want to scale training to multiple GPUs on a single Vertex AI machine without code changes. Which strategy should they use?

A.MultiWorkerMirroredStrategy
B.TPUStrategy
C.CentralStorageStrategy
D.MirroredStrategy
AnswerD

MirroredStrategy replicates the model across all GPUs on one machine using all-reduce synchronisation, distributing the existing single-GPU code across devices with no code changes. This satisfies the stem's single-machine, multi-GPU scaling requirement, unlike distributed strategies spanning multiple hosts.

Why this answer

MirroredStrategy is TensorFlow's default single-machine, multi-GPU strategy. It replicates the model on each GPU and synchronizes gradients via all-reduce, and it can be enabled with minimal code changes (typically just wrapping model creation in strategy.scope()), making it the right choice for scaling on one Vertex AI machine.

Exam trap

The trap is choosing MultiWorkerMirroredStrategy for a single-machine multi-GPU scenario — candidates confuse 'multiple GPUs' with 'multiple workers' and overlook that MirroredStrategy is the single-host default.

How to eliminate wrong answers

Option A is wrong because MultiWorkerMirroredStrategy is designed for multiple machines (workers), not a single machine with multiple GPUs. Option B is wrong because TPUStrategy targets Tensor Processing Units, not GPUs, and requires TPU-specific infrastructure. Option C is wrong because CentralStorageStrategy places variables on CPU or one GPU and is intended for scenarios where variables are too large to replicate — it is not the standard no-code-change path for multi-GPU scaling on one machine.

455
MCQeasy

A startup wants to add sentiment analysis to their customer feedback app without any labeled data or custom model training. Which Google Cloud service should they use?

A.Cloud Natural Language API
B.AutoML Natural Language with manual labeling
C.Use BigQuery ML to train a text classification model
D.Train a custom sentiment model on Vertex AI
AnswerA

Cloud Natural Language API provides pre-trained sentiment analysis through a REST call, requiring no labelled data or custom training. This satisfies the startup's constraint of adding sentiment scoring immediately without building or training a model.

Why this answer

The Cloud Natural Language API provides pre-trained models for sentiment analysis that require no labeled data or custom training. It offers a ready-to-use sentiment analysis feature via a simple API call, making it ideal for a startup that wants to add sentiment analysis without any machine learning expertise or data preparation.

Exam trap

Google Cloud often tests the distinction between pre-trained APIs and custom training services, where candidates mistakenly choose AutoML or Vertex AI because they think any ML task requires custom training, overlooking the existence of fully managed, pre-trained APIs like Cloud Natural Language API.

How to eliminate wrong answers

Option B is wrong because AutoML Natural Language with manual labeling requires labeled data and custom model training, which contradicts the requirement of no labeled data or custom training. Option C is wrong because BigQuery ML is designed for training models using SQL queries on structured data, not for pre-trained sentiment analysis, and it still requires labeled data and training. Option D is wrong because training a custom sentiment model on Vertex AI involves building, training, and deploying a custom model, which requires labeled data and significant ML effort, not a pre-built solution.

456
MCQeasy

A machine learning engineer wants to deploy a trained model to Vertex AI for online predictions. Which Vertex AI resource is required to serve the model and provide an endpoint URL?

A.Vertex AI Pipeline
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Endpoint
AnswerD

An Endpoint is the managed Vertex AI resource that hosts the deployed model and exposes a serving URL for online predictions. Deploying the model to an Endpoint allocates compute and returns the endpoint URL the engineer needs to send prediction requests.

Why this answer

Vertex AI Endpoint is the required resource to deploy a trained model for online predictions, as it provides a dedicated endpoint URL that accepts prediction requests and routes them to the model. Without an endpoint, the model cannot be accessed via HTTP/HTTPS for real-time inference, which is the core requirement for online serving.

Exam trap

The trap here is that candidates confuse the Model Registry (which stores and versions models) with the actual serving infrastructure, assuming that registering a model automatically creates an endpoint, when in fact a separate Endpoint resource must be created and the model must be deployed to it.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipeline is used for orchestrating and automating ML workflows (e.g., training, evaluation), not for serving models or providing an endpoint URL. Option B is wrong because Vertex AI Model Registry is a central repository for managing model versions and metadata, but it does not itself expose an endpoint for predictions; models must be deployed to an endpoint for serving. Option C is wrong because Vertex AI Feature Store is designed for storing, serving, and sharing feature data for training and prediction, not for hosting models or providing inference endpoints.

457
MCQmedium

An ML engineer is building a Vertex AI Pipeline that includes a data validation component. The component should fail the pipeline if the input data does not meet certain statistical thresholds. The engineer wants to ensure that the pipeline stops immediately and does not proceed to training if validation fails. Which mechanism should the engineer use in the component?

A.Raise an exception in the component code and catch it in the pipeline definition.
B.Return a non-zero exit code from the component's container.
C.Write a warning to the logs and continue execution.
D.Use a conditional branch in the pipeline to skip training if validation fails.
AnswerB

In Vertex AI Pipelines, a component failure is indicated by a non-zero exit code from the container. This causes the pipeline task to fail, and by default, the pipeline stops executing subsequent tasks. This is the standard way to enforce validation gates and prevent downstream tasks from running with invalid data.

Why this answer

A non-zero exit code from a component's container is the standard way to signal failure in Vertex AI Pipelines. When a component fails, the pipeline task fails, and the pipeline stops by default, preventing downstream tasks from executing. This enforces the validation gate effectively.

Exam trap

The trap here is thinking that logging a warning or using a conditional branch is sufficient, but only a non-zero exit code causes the pipeline to fail and stop immediately.

458
MCQeasy

A startup is deploying a scikit-learn model to Vertex AI for online predictions. They want to minimize the effort required to containerize the model and ensure it can handle HTTP requests. What should they do?

A.Use the pre-built scikit-learn container provided by Vertex AI and deploy the model by specifying the model artifact in Cloud Storage.
B.Convert the scikit-learn model to TensorFlow SavedModel format and use the TensorFlow pre-built container for deployment.
C.Deploy the model using a custom container that includes the scikit-learn library and a simple HTTP server implemented with Python's http.server module.
D.Write a custom Flask application to load the model and expose a /predict endpoint, then build a Docker image and push it to Artifact Registry.
AnswerA

Vertex AI provides pre-built containers for popular frameworks like scikit-learn. These containers handle HTTP serving, request parsing, and model loading automatically. By simply pointing to the model artifact, the startup avoids containerization effort and ensures compatibility with Vertex AI's prediction API.

Why this answer

The pre-built scikit-learn container on Vertex AI eliminates the need for custom containerization and HTTP server code. It automatically handles model loading and prediction requests, making it the lowest-effort solution for deploying scikit-learn models.

Exam trap

The trap here is assuming that a custom container is always needed, overlooking the convenience and compatibility of Vertex AI's pre-built containers for common frameworks.

459
MCQeasy

Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

A.Cloud Memorystore for Redis
B.Cloud CDN
C.Bigtable
D.Cloud Storage with object versioning
AnswerA

Memorystore for Redis provides a low-latency in-memory store that sits in front of the endpoint, letting identical feature vectors return cached predictions instead of re-invoking the model. This directly cuts both response latency and per-prediction cost, matching the stem's caching goal.

Why this answer

Cloud Memorystore for Redis is an in-memory data store that provides sub-millisecond latency, making it ideal for caching prediction responses. By caching identical prediction requests, you can reduce the number of calls to the Vertex AI endpoint, lowering latency and cost. Redis supports key-value storage with TTL, perfect for caching.

Exam trap

The trap is selecting Cloud CDN because it is a caching service, but CDN caches HTTP responses at the edge and is not suitable for caching dynamic API prediction results that require custom key logic.

How to eliminate wrong answers

Option B is wrong because Cloud CDN caches HTTP content at edge locations, but it is designed for static or cacheable web content, not for caching API prediction responses that may require authentication and dynamic inputs. Option C is wrong because Bigtable is a NoSQL database for large-scale analytical workloads, not a low-latency cache. Option D is wrong because Cloud Storage with object versioning is for storing objects with version history, not for caching API responses.

460
MCQmedium

A team is using Vertex AI Model Monitoring to detect prediction drift on a deployed model. They have configured the monitoring job to run every 6 hours. After a week, they notice that the drift metric has been consistently high but no alerts have been triggered. They have verified that the alerting policy in Cloud Monitoring is correctly configured. What is the most likely cause?

A.The sampling rate is too low, causing the drift metric to be inaccurate.
B.The drift threshold is set higher than the observed drift metric values.
C.The monitoring job is not writing the drift metrics to Cloud Monitoring.
D.The monitoring frequency is too high, causing the system to ignore the drift.
AnswerB

If the drift threshold is set higher than the actual drift values, the alert will not fire even though drift is present. The team sees high drift, but it may not exceed the threshold. Adjusting the threshold to a lower value would trigger alerts appropriately.

Why this answer

Alerts are triggered when the drift metric exceeds the configured threshold. If the threshold is set higher than the observed drift values, no alert will fire. The team should review and lower the threshold to match their tolerance for drift, ensuring alerts are generated when drift is significant.

Exam trap

The trap here is assuming that high drift always triggers alerts, but alerts depend on the threshold set in the monitoring configuration.

461
MCQmedium

The exhibit shows part of a Vertex AI Pipeline definition. The pipeline fails at the training step with an error: 'Missing required input: train_data'. What is the most likely cause?

A.The evaluation step expects a metric output but training does not produce it
B.The training step uses the wrong image tag
C.The container command for data_processing is incorrect
D.The data_processing step does not define any outputs
E.The pipeline is missing a deployment step
AnswerD

Vertex AI Pipelines passes data between steps through declared component outputs. If the data_processing component declares no output artifact, no train_data reference exists for the training step, so compilation or runtime reports the missing required input.

Why this answer

The error 'Missing required input: train_data' indicates that the training step expects an input artifact named 'train_data', but no upstream step provides it. In Vertex AI Pipelines, a component's output must be explicitly defined and connected to the downstream component's input. Since the data_processing step does not define any outputs, it cannot produce the 'train_data' artifact, causing the training step to fail.

Exam trap

Google Cloud often tests the distinction between runtime errors (e.g., container image issues) and graph validation errors (e.g., missing input/output connections), leading candidates to confuse a missing output definition with a container or command misconfiguration.

How to eliminate wrong answers

Option A is wrong because the error is about a missing input, not a missing metric output; the evaluation step's expectations are irrelevant to the training step's input requirement. Option B is wrong because an incorrect image tag would cause a container runtime error (e.g., 'ImagePullBackOff'), not a 'Missing required input' error, which is a pipeline graph validation issue. Option C is wrong because an incorrect container command for data_processing would cause that step to fail, but the error specifically points to the training step's missing input, not a failure in data_processing.

Option E is wrong because a missing deployment step would not cause a training step input error; deployment occurs after training and evaluation, and its absence would not affect the training step's input requirements.

462
MCQeasy

A data science team is using Vertex AI Pipelines to orchestrate their ML workflows. They want to ensure that each pipeline run is reproducible and that artifacts are versioned. Which Vertex AI feature should they use to track and manage pipeline artifacts?

A.Vertex AI Model Registry
B.Vertex ML Metadata
C.Vertex AI Experiments
D.Vertex AI Feature Store
AnswerB

Vertex ML Metadata automatically tracks artifacts, executions, and contexts produced by Vertex AI Pipelines runs. It provides lineage and versioning, enabling reproducibility by recording parameters and artifacts. This is the native service for managing pipeline artifacts and their relationships, making it the correct choice for the team's requirement.

Why this answer

Vertex ML Metadata is the dedicated service for capturing and managing metadata and artifacts from Vertex AI Pipelines. It records executions, artifacts, and their relationships, enabling lineage tracking and reproducibility. The other options serve different purposes: Experiments for run comparison, Model Registry for models, and Feature Store for feature serving.

Exam trap

The trap here is confusing Vertex AI Experiments with Vertex ML Metadata; while both track metadata, only Vertex ML Metadata provides comprehensive artifact lineage and versioning for pipeline executions.

463
MCQmedium

Refer to the exhibit. What is the purpose of this query?

A.To detect data drift
B.To find prediction errors in Cloud Logging
C.To count all prediction requests
D.To monitor model latency
AnswerB

Cloud Logging filters entries by severity and payload fields, surfacing mismatches between predicted and actual values logged by the model. This satisfies the stem's requirement to identify prediction errors, since the query targets error-level log entries rather than training metrics or infrastructure events.

Why this answer

The query filters Cloud Logging entries for the string 'prediction failed', which directly indicates prediction errors logged by the ML prediction service. This is a common pattern for monitoring model inference failures in production, not for measuring drift, counting requests, or measuring latency.

Exam trap

Google Cloud often tests the distinction between log-based monitoring (for errors) and metric-based monitoring (for counts, latency, drift), so candidates mistakenly choose 'count all prediction requests' when the query clearly filters for failures, not all requests.

How to eliminate wrong answers

Option A is wrong because data drift detection requires comparing feature distributions over time, not searching for error log messages. Option C is wrong because counting all prediction requests would require a metric like `prediction_count` or a log-based metric counting all prediction entries, not filtering for failures. Option D is wrong because monitoring model latency requires timing metrics (e.g., `prediction_latency_ms`) or log entries with duration fields, not a search for 'prediction failed'.

464
MCQeasy

A financial services company wants to extract text and structured data from scanned loan application forms. They need a fully managed, low-code solution that can handle various form layouts and requires minimal machine learning expertise. Which Google Cloud service should they use?

A.Document AI
B.Vision API
C.BigQuery ML
D.AutoML Natural Language
AnswerA

Document AI is a fully managed service that uses pre-trained and customizable models to parse documents, extract text, and identify structured fields. It supports form parsing and can be used with minimal ML expertise via the console or API. It is ideal for processing loan applications with varying layouts.

Why this answer

Document AI provides specialized document parsing capabilities, including form understanding and entity extraction, with pre-built models and the ability to train custom extractors with minimal effort. It is fully managed and designed for low-code document processing, making it the best fit for loan application forms.

Exam trap

The trap here is assuming that Vision API's OCR is sufficient, but it lacks structured extraction and form parsing capabilities.

465
MCQeasy

Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?

A.Cloud CDN
B.Cloud Memorystore
C.Cloud Spanner
D.BigQuery
AnswerB

Cloud Memorystore provides a managed Redis or Memcached layer that stores prediction results keyed by request parameters, so identical requests return cached responses without invoking the endpoint. This cuts both latency and endpoint load, matching the repeated-prediction pattern described.

Why this answer

Cloud Memorystore (B) is correct because it provides a managed in-memory cache (Redis or Memcached) that can store the results of identical prediction requests, reducing latency and load on the prediction endpoint. By caching responses keyed on the user profile or request parameters, subsequent identical requests can be served directly from Memorystore in microseconds, avoiding redundant model inference.

Exam trap

The trap here is that candidates confuse caching at the edge (CDN) with caching at the application layer (Memorystore), assuming any cache service works for dynamic API responses, but Cloud CDN cannot cache POST requests or application-specific payloads without significant configuration and still lacks the fine-grained key-value semantics needed for identical prediction requests.

How to eliminate wrong answers

Option A (Cloud CDN) is wrong because it caches static content (e.g., images, CSS) at edge locations, not dynamic API responses for identical requests; it cannot cache POST request payloads or application-level prediction results without complex workarounds. Option C (Cloud Spanner) is wrong because it is a globally distributed relational database designed for transactional consistency and high availability, not for low-latency caching of ephemeral prediction responses. Option D (BigQuery) is wrong because it is a serverless data warehouse for analytical queries on large datasets, not a caching layer for real-time inference results.

466
MCQmedium

You have a Python training script that reads a 500 GB CSV dataset from a Cloud Storage bucket. You submit a Vertex AI custom training job using a pre-built container, specifying a machine with 16 vCPUs and 60 GB RAM. The job fails after a few minutes with an out-of-memory error. You need to scale the prototype to handle this dataset without changing the model architecture. What should you do?

A.Increase the machine type to one with 128 GB RAM and retry the job.
B.Convert the CSV files to TFRecord format and use the tf.data API to stream the data in batches during training.
C.Split the CSV files into smaller files and use a larger number of workers with data parallelism.
D.Use Vertex AI Pipelines to preprocess the data and store it in BigQuery, then read from BigQuery during training.
AnswerB

Converting to TFRecord and using tf.data enables efficient streaming from Cloud Storage, so the full dataset is never loaded into memory. This directly addresses the out-of-memory error by decoupling data size from machine memory. It is the recommended approach for large-scale training on Vertex AI and requires only changing the input pipeline, not the model.

Why this answer

The out-of-memory error occurs because the training script attempts to load the entire 500 GB dataset into memory. Using TFRecord and tf.data allows streaming data in batches, so memory usage remains bounded. This is the standard method for scaling data input on Vertex AI and requires minimal changes to the training code.

Exam trap

The trap here is assuming that simply increasing machine memory or splitting files will solve the problem, when the real fix is to stream data instead of loading it all at once.

467
MCQmedium

A company has a large dataset of labeled images (e.g., different species of plants). They want to train a custom image classification model with minimal effort and no prior ML experience. Which Google Cloud service should they use?

A.Cloud TPU
B.AutoML Vision
C.Vertex AI Workbench with a custom TensorFlow model
D.Vision API
AnswerB

AutoML Vision trains custom image classifiers through a point-and-click interface, using transfer learning on Google's pretrained models. This satisfies the stem's constraints: labelled plant images, minimal effort, and no ML expertise. Vertex AI Vision now supersedes it, but AutoML Vision remains the purpose-built answer for codeless custom classification.

Why this answer

AutoML Vision is the correct choice because it allows users with no prior ML experience to train a custom image classification model using a simple graphical interface, requiring only labeled images as input. It automates model architecture search, hyperparameter tuning, and deployment, minimizing manual effort while delivering a production-ready model.

Exam trap

The trap here is that candidates confuse AutoML Vision (custom model training with minimal effort) with Vision API (pre-trained, no custom training), often picking D because both involve 'Vision' and seem low-code, but Vision API cannot be retrained on custom data.

How to eliminate wrong answers

Option A is wrong because Cloud TPU is a hardware accelerator for training custom models, requiring users to write and manage their own ML code (e.g., TensorFlow/PyTorch), which demands significant ML expertise and effort. Option C is wrong because Vertex AI Workbench with a custom TensorFlow model requires users to write, debug, and train a model from scratch using notebooks, which is not minimal effort and assumes ML experience. Option D is wrong because Vision API is a pre-trained API for general image recognition tasks (e.g., label detection, OCR) and cannot be trained on custom labeled datasets like plant species; it offers no customization for specific classification needs.

468
MCQhard

An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?

A.Increase the request timeout on the interactive endpoint and add more replicas.
B.Enable autoscaling on the bulk endpoint so it absorbs the interactive traffic as well.
C.Move both endpoints to a machine type with more vCPUs and more memory.
D.Deploy a container variant tuned for the interactive endpoint: load the model once at startup, keep it resident, and avoid per-request batch assembly, then serve it on a latency-optimized machine type.
AnswerD

Latency on small requests is dominated by fixed per-request overhead, so removing repeated model loading and batch assembly from the request path gives the largest gain. A container that loads the model once and keeps it warm in memory, paired with a machine type chosen for single-request speed rather than bulk throughput, directly targets the 200 ms SLO while leaving the bulk endpoint's throughput-optimized configuration untouched.

Why this answer

Interactive single-record scoring and bulk batched scoring have opposite performance profiles, so they should not share one serving configuration. The interactive path is dominated by fixed per-request overhead such as model loading and batch assembly, so a container that loads the model once, keeps it resident, and skips per-request batching removes that cost. Pairing it with a machine type selected for low-latency single requests, while leaving the bulk endpoint tuned for throughput, resolves the SLO breach without over-sizing the bulk tier.

Exam trap

The trap here is treating a latency SLO breach as a pure capacity problem and adding replicas or vCPUs, when the fixed per-request overhead of batch-oriented model loading is the actual cause.

469
Multi-Selecteasy

A data analyst wants to build a binary classification model using a low-code ML solution on Google Cloud. The dataset is stored in BigQuery and contains 500,000 rows with 20 features, including categorical and numerical columns. The analyst has minimal coding experience and needs to deploy the model as an API endpoint for real-time predictions. Which two Google Cloud services should the analyst use to accomplish this task with minimal code? Choose two options.

Select 2 answers
A.BigQuery ML
B.Vertex AI Endpoints
C.Cloud Functions
D.Vertex AI Workbench
E.AutoML Tables
AnswersB, E

Vertex AI Endpoints provides a serverless option to deploy trained models as REST APIs with autoscaling, ideal for real-time predictions without code.

Why this answer

Vertex AI Endpoints is correct because it provides a managed service to deploy trained models as REST API endpoints for real-time predictions with minimal code. The analyst can deploy an AutoML Tables model directly to a Vertex AI Endpoint, enabling low-code deployment and serving.

Exam trap

Google Cloud often tests the distinction between model training services (BigQuery ML, AutoML Tables) and model deployment services (Vertex AI Endpoints), leading candidates to incorrectly select BigQuery ML for real-time API deployment when it only supports batch inference.

470
MCQeasy

A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?

A.Text-to-Speech and Natural Language API
B.Speech-to-Text and Translation API
C.Video Intelligence API
D.Speech-to-Text and Natural Language API
AnswerD

Speech-to-Text streams audio into text with low latency, and the Natural Language API then performs sentiment analysis and entity detection on that text. Together they deliver real-time transcription plus sentiment and urgency identification without training custom models.

Why this answer

Real-time call transcription requires Speech-to-Text to convert audio to text with low latency, and Natural Language API to analyze that text for sentiment and entity/urgency detection. Together they form the standard streaming audio-analysis pipeline on Google Cloud. Speech-to-Text supports streaming recognition, and Natural Language provides sentiment and entity analysis on the resulting transcript.

Exam trap

PMLE often tests whether candidates confuse the direction of the audio APIs, so picking Text-to-Speech (which synthesizes speech) instead of Speech-to-Text (which transcribes) is the classic wrong-answer trap for transcription scenarios.

How to eliminate wrong answers

Option A is wrong because Text-to-Speech converts text to audio — the opposite direction — and cannot transcribe incoming calls. Option B is wrong because Translation API translates between languages; it does not perform sentiment analysis or urgency detection, so it fails the stated requirement. Option C is wrong because Video Intelligence API analyzes video content (labels, shots, explicit content) and is not designed for real-time audio transcription or sentiment analysis.

471
MCQeasy

A company uses Cloud Composer to orchestrate their ML pipelines. They notice that tasks are being queued but not executed, causing delays. What is the most likely cause?

A.The Airflow web server is down
B.The DAG file is corrupted
C.The Cloud Storage bucket containing DAGs is not accessible
D.The Airflow worker resources are exhausted
AnswerD

Queued-but-unexecuted tasks mean the scheduler has work but no slot to run it. Exhausted worker CPU or memory on the Composer cluster prevents Airflow from starting queued task instances, so the queue grows while nothing executes. Scaling worker count or resources resolves the backlog.

Why this answer

When tasks are queued but not executed, it typically indicates that the Airflow workers have no available slots to pick up new tasks. In Cloud Composer, the Celery executor distributes tasks to workers; if all worker concurrency slots are saturated or the worker node pool is under-provisioned, tasks remain in the 'queued' state until a worker becomes free. This is the most likely cause given the symptom of tasks being queued without execution.

Exam trap

The trap here is that candidates confuse the roles of Airflow components (web server, scheduler, worker) and assume a UI or DAG access issue causes queued tasks, when in reality the worker capacity is the bottleneck.

How to eliminate wrong answers

Option A is wrong because the Airflow web server is responsible for the UI and DAG parsing, not for executing tasks; if it were down, the UI would be inaccessible but tasks could still be queued and executed by workers. Option B is wrong because a corrupted DAG file would cause a parse error, preventing the DAG from being scheduled or appearing in the UI, not leaving tasks in a queued state. Option C is wrong because if the Cloud Storage bucket containing DAGs were not accessible, the DAGs would not be synced to the Airflow environment at all, resulting in missing DAGs rather than queued tasks.

472
MCQeasy

You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?

A.n1-standard-2 with NVIDIA Tesla T4
B.e2-standard-2
C.n1-standard-2
D.n1-highmem-2
AnswerA

NVIDIA Tesla T4 GPUs deliver low-latency inference for online predictions, and the n1-standard-2 shape supplies adequate vCPU and memory for a single model replica. This satisfies the stem's latency-minimisation constraint, since T4s accelerate matrix operations that CPU-only machine types would process far more slowly.

Why this answer

For online predictions with minimal latency, a machine type with a GPU is essential because it accelerates model inference. The n1-standard-2 with NVIDIA Tesla T4 provides GPU acceleration, which significantly reduces latency compared to CPU-only instances. The other options lack GPUs and are not optimized for low-latency inference.

Exam trap

The trap is overlooking the need for a GPU for low-latency inference and selecting a CPU-only machine type based on cost or memory, when the question explicitly asks to minimize latency.

How to eliminate wrong answers

Option B is wrong because e2-standard-2 is a general-purpose CPU-only machine type, which would result in higher latency for model inference. Option C is wrong because n1-standard-2 is CPU-only and lacks the parallel processing power of a GPU for inference. Option D is wrong because n1-highmem-2 is optimized for memory-intensive workloads but still CPU-only, so it does not minimize latency for predictions.

473
Multi-Selecthard

Which THREE actions should be taken to manage model versions effectively?

Select 3 answers
A.Delete old versions immediately
B.Use Vertex AI Model Registry
C.Set up model evaluation alerts
D.Use the same model name for all versions
E.Assign version aliases like 'champion' and 'experiment'
AnswersB, C, E

Model Registry provides versioning and deployment control.

Why this answer

Vertex AI Model Registry is a centralized repository that tracks, versions, and manages ML models. It enables you to organize models, assign aliases (like 'champion' or 'experiment'), and control deployment, ensuring reproducibility and governance across the model lifecycle.

Exam trap

Google Cloud often tests the misconception that deleting old versions is a best practice for storage optimization, when in reality versioning requires retaining history for reproducibility and rollback, and that aliases are the correct mechanism for labeling model stages.

474
MCQeasy

A data scientist has deployed a classification model on a Vertex AI Endpoint and wants to monitor for feature drift in the serving data compared to the training data. Which Vertex AI service should be used?

A.Vertex AI Explainable AI
B.Vertex AI Model Evaluation
C.Vertex AI Continuous Training
D.Vertex AI Model Monitoring
AnswerD

Vertex AI Model Monitoring compares serving-time feature distributions against the training baseline and raises drift alerts when skew or drift thresholds are exceeded. It is the native service for detecting feature drift on a deployed Endpoint, satisfying the monitoring requirement directly.

Why this answer

Vertex AI Model Monitoring is the purpose-built service that detects training-serving skew and prediction drift by comparing live serving data distributions against a baseline (typically the training dataset). It computes statistics like Jensen-Shannon divergence and can alert when drift exceeds configured thresholds. Explainable AI, Model Evaluation, and Continuous Training serve different purposes and do not provide ongoing drift detection on deployed endpoints.

Exam trap

PMLE often tests the confusion between Model Monitoring (ongoing drift detection on deployed models) and Model Evaluation (one-time performance assessment), catching candidates who pick Evaluation for live drift detection.

How to eliminate wrong answers

Option A is wrong because Explainable AI provides feature attributions (e.g., SHAP values) for individual predictions, not distribution drift monitoring. Option B is wrong because Model Evaluation assesses model performance metrics (accuracy, AUC) on a test set at training time, not live serving data drift. Option C is wrong because Continuous Training is a broader MLOps pattern for automating retraining, not a monitoring service that detects drift—it is the action taken after drift is detected.

475
MCQmedium

You are preparing a Vertex AI custom training job that uses a custom Docker image built on top of the PyTorch pre-built training container. The image must be able to read training data from a Cloud Storage bucket without embedding credentials in the image. You want the job to run on a single NVIDIA T4 GPU. Which approach should you take?

A.Mount the Cloud Storage bucket as a local file system using Cloud Storage FUSE inside the container and rely on the bucket's default IAM permissions for anonymous access.
B.Hard-code a service account JSON key file into the Docker image and set GOOGLE_APPLICATION_CREDENTIALS in the Dockerfile.
C.Specify a user-managed service account with the Storage Object Viewer role when submitting the custom job, and let the container use Application Default Credentials to read from Cloud Storage.
D.Grant the Vertex AI Custom Code Service Agent the Storage Object Viewer role and let the training container use the default credentials from the metadata server.
AnswerC

When you submit a Vertex AI custom training job, you can attach a user-managed service account. The container then obtains credentials from the Compute Engine metadata server through Application Default Credentials. Granting that service account the Storage Object Viewer role allows the training code to read objects from Cloud Storage without embedding keys, and the job can be configured with a single T4 GPU accelerator.

Why this answer

A user-managed service account attached to the custom job provides the identity that the training container uses. Vertex AI injects credentials via the metadata server, so Application Default Credentials in the container can authenticate to Cloud Storage. Granting the least-privilege Storage Object Viewer role satisfies the data access requirement without embedding keys.

The job can also be configured with a single T4 GPU accelerator to meet the compute requirement.

Exam trap

The trap here is confusing the Vertex AI Custom Code Service Agent, which acts on behalf of the Vertex AI service, with the user-managed service account that actually runs inside the training container and accesses Cloud Storage.

476
MCQhard

A company serves a PyTorch model using a custom container on Vertex AI Prediction. They notice that after a few hours, the endpoint returns 502 errors. The logs show 'Out of memory' errors. The container has a memory limit of 4GB, and the model loads a 3GB vocabulary file. What is the most likely cause and best fix?

A.Increase the container memory to 8GB.
B.Load the vocabulary file once at startup and reuse it.
C.Increase the number of replicas to distribute load.
D.Switch to Vertex AI Batch Prediction.
AnswerB

Reloading the 3GB vocabulary per request exhausts the 4GB container limit, causing out-of-memory 502s. Loading it once at startup keeps it resident in memory and reused across requests, eliminating the repeated allocation that triggers the errors.

Why this answer

The 502 errors and 'Out of memory' errors indicate that the container is running out of memory during inference. Since the model loads a 3GB vocabulary file, and the container has only 4GB of memory, loading this file repeatedly for each prediction request (e.g., inside the prediction handler) would quickly exhaust memory. The correct fix is to load the vocabulary file once at container startup and reuse it across all requests, which is a standard best practice for serving models with large static assets.

Exam trap

Google Cloud often tests the misconception that OOM errors are always solved by increasing memory, but the trap here is that the real issue is inefficient resource reuse—loading a large file per request—rather than insufficient total memory.

How to eliminate wrong answers

Option A is wrong because simply increasing memory to 8GB does not address the root cause—the vocabulary file is being loaded repeatedly, which will still cause memory bloat and eventual OOM errors, just at a higher threshold. Option C is wrong because increasing replicas distributes incoming traffic but does not fix the per-container memory leak caused by repeated vocabulary loading; each replica would still suffer the same OOM issue. Option D is wrong because switching to Batch Prediction is for offline, asynchronous processing, not for real-time serving, and does not solve the memory management problem within the container.

477
MCQmedium

You are fine-tuning a pre-trained BERT model from Hugging Face on a custom text classification dataset using Vertex AI Training. You want to speed up training by using mixed precision. What should you do?

A.Modify the model to use half-precision layers
B.Use a custom container with TensorFlow instead of PyTorch
C.Enable mixed precision via Vertex AI hyperparameter tuning
D.Set fp16=True in the TrainingArguments
AnswerD

Setting fp16=True in TrainingArguments enables NVIDIA AMP mixed precision, using float16 for forward and backward passes while keeping master weights in float32. This reduces memory and speeds training on compatible GPUs without manual loss scaling.

Why this answer

Hugging Face's Trainer API exposes mixed precision through the TrainingArguments parameter fp16=True (or bf16=True for bfloat16). Setting this flag enables automatic mixed precision (AMP) via PyTorch's torch.cuda.amp, which casts eligible operations to FP16 while keeping master weights in FP32 for numerical stability. This is the standard, supported way to accelerate fine-tuning on GPU without rewriting the model.

Exam trap

The trap here is confusing framework-level training configuration (fp16=True in TrainingArguments) with infrastructure-level tuning (Vertex AI hyperparameter tuning) or model surgery (half-precision layers), when mixed precision is purely a Trainer API flag.

How to eliminate wrong answers

Option A is wrong because manually converting model layers to half-precision (e.g., .half()) discards FP32 master weights and typically causes NaN losses and degraded accuracy during training. Option B is wrong because switching frameworks from PyTorch to TensorFlow has no bearing on mixed precision and would require rewriting the training code. Option C is wrong because Vertex AI hyperparameter tuning searches over hyperparameter values; it does not enable mixed precision, which is a training configuration flag, not a tunable hyperparameter.

478
Multi-Selectmedium

Your team manages multiple ML models on Vertex AI. You need to implement a centralized monitoring solution to track model performance over time. Which TWO approaches should you consider? (Choose two.)

Select 2 answers
A.Store all prediction logs in BigQuery and analyze using SQL.
B.Use Cloud Source Repositories to track model code versions.
C.Create Cloud Monitoring dashboards and alerts based on Vertex AI metrics.
D.Use Vertex AI Model Monitoring to detect training-serving skew and feature drift for each model.
E.Enable Cloud Billing budgets to track cost per model.
AnswersC, D

Cloud Monitoring natively ingests Vertex AI's built-in metrics, such as prediction drift and training-serving skew, so dashboards and alerts can be centralised across every deployed model without custom instrumentation. This directly satisfies the requirement for a centralised monitoring solution tracking performance over time.

Why this answer

Option C is correct because Cloud Monitoring natively ingests Vertex AI metrics (e.g., prediction counts, latency, error rates, and Model Monitoring anomaly signals) and lets you build centralized dashboards and alerting policies across all deployed models in one place. Option D is correct because Vertex AI Model Monitoring is the purpose-built service that continuously computes training-serving skew and prediction drift against a baseline for each deployed model, emitting metrics and alerts that feed the centralized monitoring view. Option A is not the intended answer because storing prediction logs in BigQuery is a data-warehouse analysis pattern, not a centralized monitoring solution with dashboards and alerts.

Option B is incorrect because Cloud Source Repositories is a Git version-control service for source code, not a runtime performance monitoring tool. Option E is incorrect because Cloud Billing budgets track spend, not model performance metrics such as skew or drift.

Exam trap

The trap here is that candidates may confuse logging (Option A) or cost tracking (Option E) with performance monitoring, or mistakenly think version control (Option B) is part of monitoring, when the question specifically asks for centralized monitoring of model performance over time.

479
Multi-Selecthard

An ML engineer is troubleshooting why a Vertex AI Endpoint is returning high prediction latency. They have enabled request/response logging and see that some requests take >1 second while most are fast. Which THREE actions should they take to diagnose the issue?

Select 3 answers
A.Check CPU and GPU utilisation metrics on the endpoint.
B.Review the request/response logs in BigQuery to identify if large payloads correlate with high latency.
C.Disable Vertex AI Model Monitoring to reduce overhead.
D.Check the p99 latency metric in Cloud Monitoring.
E.Increase the number of replicas to resolve the issue immediately.
AnswersA, B, D

Checking CPU and GPU utilisation metrics directly exposes resource saturation on the endpoint's underlying nodes, which is the most likely cause of sporadic >1-second latency spikes amid otherwise fast responses. Vertex AI publishes these per-replica metrics to Cloud Monitoring, letting the engineer correlate spikes with the slow requests seen in the request/response logs.

Why this answer

Option A is correct because high prediction latency on a Vertex AI Endpoint is frequently caused by resource saturation, so inspecting CPU and GPU utilisation metrics in Cloud Monitoring reveals whether the underlying nodes are maxed out and throttling inference. Option B is correct because request/response logging for Vertex AI Endpoints can be routed to BigQuery, where the engineer can correlate payload size with latency and confirm whether large inputs (e.g., big images or long text) are the outliers exceeding 1 second. Option D is correct because p99 latency in Cloud Monitoring exposes the tail behaviour of the distribution, showing whether the slow requests are a persistent tail problem or isolated spikes, which is exactly the symptom described.

Option C is not appropriate because Model Monitoring runs asynchronously on sampled data and does not add per-request prediction latency, so disabling it would not diagnose the issue. Option E is not appropriate because blindly scaling replicas is a remediation step, not a diagnostic one, and it would not identify the root cause of the >1 second requests.

Exam trap

PMLE often tests the distinction between diagnostic and remediation actions; candidates may choose scaling (E) or disabling monitoring (C) as quick fixes, but the question asks for diagnostic steps, so only actions that gather data (A, B, D) are correct.

480
MCQhard

A financial services firm wants relationship managers to query a natural-language interface such as 'show me customers likely to churn next quarter' and receive results from their BigQuery data warehouse. They want to minimize the code they maintain and keep data in BigQuery. Which approach best fits?

A.Enable Gemini in BigQuery to generate SQL from natural-language questions over the warehouse
B.Export the warehouse to Cloud Storage and train a Vertex AI AutoML Tables model nightly
C.Build a custom chatbot with Dialogflow CX and hand-written SQL for each intent
D.Use BigQuery ML to train an ARIMA_PLUS model on customer activity to predict churn
AnswerA

Gemini in BigQuery translates natural-language prompts into SQL against the firm's tables and schemas, letting managers ask churn-style questions without writing queries. It keeps data in BigQuery and requires no bespoke codebase, matching the low-maintenance, in-warehouse requirement.

Why this answer

Gemini in BigQuery converts natural-language prompts into SQL executed against the firm's existing tables, so relationship managers get ad hoc answers without maintaining query code and data never leaves BigQuery. The alternatives either require hand-written SQL per intent, apply an unsuitable forecasting model, or move data out of the warehouse into a batch training pipeline.

Exam trap

The trap here is equating a conversational agent like Dialogflow CX with natural-language-to-SQL, when the former needs predefined intents and queries.

481
MCQmedium

A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?

A.Deploy the new model to a new endpoint and use a load balancer to distribute traffic between the old and new endpoints.
B.Create a new endpoint for the new model and update the application to call the new endpoint after testing.
C.Use a batch prediction job to test the new model on a large dataset, then replace the existing model on the endpoint.
D.Deploy the new model as a new version on the same endpoint, set it to receive 0% of traffic initially, then gradually increase traffic while monitoring performance.
AnswerD

Deploying a new model version on the same endpoint allows for seamless traffic splitting without downtime. Starting with 0% traffic lets you validate the new model with a small percentage of live traffic or via direct prediction requests. Gradually increasing traffic enables A/B testing and performance monitoring. If issues arise, you can quickly shift traffic back to the previous version.

Why this answer

The optimal approach is to deploy the new model as a version on the existing endpoint and use traffic splitting to gradually shift traffic. This provides zero-downtime updates, allows performance validation with real traffic, and enables quick rollback by adjusting traffic percentages. Other methods either cause downtime, lack gradual rollout, or add unnecessary complexity.

Exam trap

The trap here is thinking that a new endpoint is required for testing, when Vertex AI endpoints support multiple model versions and traffic splitting natively.

482
MCQhard

An organization needs to implement MLOps with standardized pipeline templates across multiple teams. Which Vertex AI feature should they use to create reusable pipeline components?

A.Vertex AI Pipelines
B.Vertex AI Experiments
C.Vertex AI Metadata
D.Vertex AI Workbench
AnswerA

Vertex AI Pipelines lets teams author reusable components and pipeline templates with the Kubeflow Pipelines DSL, then share and parameterise them across projects. This directly satisfies the requirement for standardised, reusable pipeline templates across multiple teams, since components are versioned artefacts rather than one-off scripts.

Why this answer

Vertex AI Pipelines is the managed orchestration service that lets teams define, run, and monitor ML workflows as directed acyclic graphs (DAGs) using Kubeflow Pipelines or TFX. It supports reusable pipeline components and templates, enabling standardization across teams by packaging preprocessing, training, evaluation, and deployment steps into versioned, shareable artifacts.

Exam trap

The trap is confusing Vertex AI Experiments (run tracking) with Vertex AI Pipelines (workflow orchestration) — both are MLOps tools, but only Pipelines provides reusable component templates for standardized workflows.

How to eliminate wrong answers

Option B is wrong because Vertex AI Experiments is designed for tracking and comparing experiment runs (metrics, parameters, artifacts), not for building reusable pipeline components or orchestrating multi-step workflows. Option C is wrong because Vertex AI Metadata is the underlying artifact and lineage tracking service — it records what happened but does not define or execute pipelines. Option D is wrong because Vertex AI Workbench is an interactive notebook environment for development, not a pipeline orchestration or component-reuse system.

483
MCQhard

A financial services company uses Vertex AI Pipelines to train and deploy models for fraud detection. The ML team consists of data scientists who develop models and ML engineers who deploy them. They use a CI/CD pipeline with Cloud Build to build and push Docker images to Artifact Registry, then trigger Vertex AI Pipelines. Recently, the team noticed that a model deployed to production was trained on a dataset that had not been approved by the data governance team. Upon investigation, they found that a data scientist accidentally used an unapproved version of the training data by specifying a Cloud Storage path that was not the latest approved dataset. The company needs to enforce that only approved datasets are used in training jobs. Which approach should they take?

A.Implement a manual approval process where data scientists request dataset paths from the data governance team before each training run.
B.After training, run a validation step that checks if the dataset used matches the latest approved version, and roll back if not.
C.Use a curated dataset registry in BigQuery or Cloud Storage with IAM conditions that allow access only to datasets tagged as 'approved'. Modify the CI/CD pipeline to pass only approved dataset references to the training job.
D.Restrict all Cloud Storage buckets to be read-only for the data scientists, and have ML engineers copy approved datasets to a separate bucket.
AnswerC

This automates governance by restricting training to approved datasets via IAM and pipeline configuration.

Why this answer

It enforces governance at the source by using IAM conditions to restrict access to only approved datasets, preventing unauthorized data from being used in training. This approach integrates with the CI/CD pipeline to automatically pass only approved dataset references, eliminating the risk of human error in specifying Cloud Storage paths.

Exam trap

Google Cloud often tests the distinction between reactive validation (Option B) and proactive enforcement (Option C), where candidates mistakenly choose a post-training check that wastes resources instead of a preventive IAM-based control.

How to eliminate wrong answers

Option A is wrong because a manual approval process is error-prone, slow, and does not scale; it relies on human compliance rather than automated enforcement, leaving the system vulnerable to accidental misuse. Option B is wrong because it is reactive—it detects the issue after training has already occurred, wasting compute resources and potentially exposing the model to unapproved data before rollback. Option D is wrong because it restricts data scientists' access entirely, which hinders their ability to experiment and develop models; it also shifts the burden to ML engineers without addressing the root cause of dataset version control.

484
MCQmedium

A team of ML engineers is collaborating on a project using Vertex AI. They want to ensure that only approved models are deployed to production. Which approach should they use?

A.Store all models in a Cloud Storage bucket and manually control access via IAM permissions.
B.Deploy models directly from training jobs to an endpoint without version tracking.
C.Use Vertex AI Model Registry with version aliases to manage model versions and promote them after approval.
D.Use Cloud Dataflow to transform raw predictions and then store them in BigQuery for analysis.
AnswerC

Vertex AI Model Registry with version aliases lets the team track model versions and control which alias points to an approved artefact, so only vetted models reach production. Promotion after approval enforces the governance gate the scenario requires.

Why this answer

Vertex AI Model Registry provides a centralized repository for managing model versions, with support for version aliases (e.g., 'champion', 'challenger') that allow teams to promote models to production only after approval. This ensures governance and traceability, meeting the requirement that only approved models are deployed.

Exam trap

The trap here is that candidates may confuse storage access control (IAM) with model lifecycle governance, or assume that any data pipeline tool (Dataflow) can manage model approvals, when in fact only a dedicated model registry with version aliases provides the required approval workflow and traceability.

How to eliminate wrong answers

Option A is wrong because storing models in Cloud Storage with manual IAM control lacks version tracking, approval workflows, and integration with Vertex AI's deployment services, making it error-prone and unscalable for production governance. Option B is wrong because deploying directly from training jobs without version tracking bypasses model validation, approval gates, and rollback capabilities, violating the requirement for controlled production deployments. Option D is wrong because Cloud Dataflow is a data processing service for stream/batch pipelines, not a model management or approval mechanism; it is irrelevant to controlling which models are deployed.

485
MCQeasy

A company deploys an online prediction model serving 100 requests per second. They are optimizing for both latency and throughput. Which monitoring strategy should they use?

A.Monitor only the request count and set an alert if it drops below a threshold.
B.Set a single alert on the 99th percentile latency and ignore throughput since it's already high.
C.Monitor the error rate and set an alert if it exceeds 1%.
D.Monitor both the p50 and p99 latency, and the request count. Create a dashboard showing latency vs. throughput at different load levels.
AnswerD

Monitoring p50 and p99 latency alongside request count directly satisfies the latency-and-throughput constraint at 100 requests per second. Percentiles expose tail latency that averages hide, while correlating both metrics against load levels reveals saturation points, letting the team tune the model for throughput without breaching latency targets.

Why this answer

Monitoring both p50 and p99 latency alongside request count provides a comprehensive view of system performance under load. Latency percentiles reveal tail behavior (p99) and typical user experience (p50), while request count tracks throughput. A dashboard correlating latency vs. throughput at different load levels is essential for identifying performance cliffs or degradation before failures occur, aligning with best practices for production ML inference systems.

Exam trap

The trap here is that candidates often focus on a single metric (e.g., error rate or p99 latency) and overlook the need for multi-metric correlation, especially the latency-throughput trade-off, which is a core concept in monitoring ML systems under production load.

How to eliminate wrong answers

Option A is wrong because monitoring only request count and alerting on a drop below threshold ignores latency and error rate, missing critical issues like increased response times or silent failures that degrade user experience without reducing request count. Option B is wrong because setting a single alert on p99 latency and ignoring throughput neglects the trade-off between latency and throughput; high throughput can mask latency spikes, and p99 alone does not capture system capacity limits or performance under varying load. Option C is wrong because monitoring only error rate and alerting on 1% misses latency degradation and throughput drops; a system can have low error rates but high latency (e.g., due to queue buildup) or reduced throughput, both of which violate performance objectives.

486
MCQhard

An ML engineer is using Vertex AI Model Monitoring on a deployed model that predicts customer churn. The model's input features include a mix of numerical and categorical features. The engineer wants to detect changes in the distribution of a specific categorical feature, 'contract_type', which has 20 possible values. Which approach should be used to effectively monitor drift for this feature?

A.Configure drift detection specifically for 'contract_type' and set an appropriate threshold based on its cardinality.
B.Use Vertex AI Explainable AI to compute feature attributions and monitor changes in attributions for 'contract_type'.
C.Convert 'contract_type' into a numerical feature using one-hot encoding and then monitor it as a numerical feature.
D.Enable drift detection for all features and use the default threshold.
AnswerA

Vertex AI Model Monitoring supports feature-level drift detection. For categorical features with multiple values, you can set a specific threshold that accounts for the feature's cardinality. By focusing on 'contract_type' and tuning the threshold, the engineer can effectively detect meaningful distribution shifts without being overwhelmed by noise from other features. This targeted approach is best for monitoring a specific high-cardinality categorical feature.

Why this answer

Vertex AI Model Monitoring allows per-feature drift detection with customizable thresholds. For categorical features with many categories, the default threshold may not be optimal; tuning the threshold based on cardinality helps balance sensitivity and false positives. By explicitly monitoring 'contract_type', the engineer can detect shifts in its distribution, which is crucial for model performance if that feature is important.

Exam trap

The trap here is assuming that default monitoring across all features is sufficient, without considering the need to tune thresholds for high-cardinality categorical features.

487
Multi-Selectmedium

Which TWO of the following are best practices for managing data in a collaborative machine learning environment on Google Cloud?

Select 2 answers
A.Always replicate data across multiple regions to ensure low latency.
B.Implement fine-grained access control using IAM conditions.
C.Use Cloud Data Catalog to discover and annotate datasets.
D.Store all raw data in a single Cloud Storage bucket for easy access.
E.Use data versioning with tools like DVC or Dataflow to track changes.
AnswersC, E

Data Catalog aids in data governance and collaboration.

Why this answer

Cloud Data Catalog provides a managed metadata management service that allows teams to discover, annotate, and manage datasets across Google Cloud. It enables data scientists to search for datasets by tags, descriptions, and schema, which is essential for collaboration and data governance in a multi-user ML environment.

Exam trap

Google Cloud often tests the misconception that 'replication equals performance' or that 'single bucket simplicity is best,' when in reality collaborative ML requires discoverability (Data Catalog) and reproducibility (versioning) over raw storage or access control alone.

488
MCQmedium

A company deploys a model on Vertex AI Prediction with autoscaling enabled. They notice that during a traffic spike, new instances take several minutes to become available, causing high latency. What is the best solution?

A.Disable autoscaling and use a fixed number of replicas
B.Increase the max replicas setting
C.Decrease the machine type to reduce provisioning time
D.Set a higher min replicas to maintain a baseline of warm instances
AnswerD

Autoscaling adds instances reactively, and cold-start provisioning plus model loading takes minutes, so latency spikes before capacity arrives. Keeping a higher minimum replica count maintains pre-warmed instances that absorb the initial burst, directly addressing the stem's slow scale-out constraint.

Why this answer

Setting a higher min replicas ensures that a baseline number of instances are always warm and ready to serve traffic. During a traffic spike, new instances still take time to provision (cold start), but the warm instances handle the initial surge without latency spikes. This directly addresses the observed high latency during spikes.

Exam trap

Google Cloud often tests the misconception that increasing max replicas or decreasing machine type solves cold-start latency, when the real solution is maintaining a warm baseline via min replicas.

How to eliminate wrong answers

Option A is wrong because disabling autoscaling and using a fixed number of replicas eliminates elasticity, leading to either over-provisioning (cost) or under-provisioning (latency) during variable traffic. Option B is wrong because increasing max replicas only raises the ceiling for scaling out; it does not reduce the cold-start provisioning time for new instances during a spike. Option C is wrong because decreasing the machine type reduces compute capacity per instance, which can increase latency under load, and does not meaningfully reduce provisioning time (which is dominated by container image pull and model loading, not machine type).

489
Multi-Selectmedium

A company is using Vertex AI Pipelines to orchestrate a training workflow. They want to implement a CI/CD process where the pipeline is automatically triggered when a new version of the training code is pushed to a GitHub repository. They also want to ensure that the pipeline uses the latest code. Which two actions should they take? (Choose two.)

Select 2 answers
A.Use a Cloud Function to monitor the GitHub repository and trigger the pipeline directly when a push occurs.
B.Configure the Vertex AI Pipeline to pull the training code from GitHub at runtime using a git clone step.
C.Create a Cloud Build trigger that builds a new container image from the GitHub repository and pushes it to Artifact Registry.
D.Set up a Cloud Build trigger that submits the Vertex AI Pipeline job, using the newly built image as a parameter.
E.Store the training code in a Cloud Storage bucket and have the pipeline download it at the start of each run.
AnswersC, D

A Cloud Build trigger can be configured to watch the GitHub repository and automatically build a container image whenever code is pushed. This image contains the latest training code and is pushed to Artifact Registry. The pipeline can then reference this image, ensuring the latest code is used. This is a standard CI step for ML pipelines.

Why this answer

The two correct actions are to create a Cloud Build trigger that builds and pushes a container image from the GitHub repository, and to set up a Cloud Build trigger that submits the Vertex AI Pipeline job using that image. This creates a CI/CD pipeline where code changes automatically trigger a build and then a pipeline run with the latest code, ensuring reproducibility and automation.

Exam trap

The trap here is thinking that the pipeline can directly pull code from GitHub or that a Cloud Function is needed, when Cloud Build triggers are the native integration for CI/CD with Vertex AI Pipelines.

490
Drag & Dropmedium

Drag and drop the steps to create and deploy a custom ML model on Vertex AI using a container in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order to create and deploy a custom ML model on Vertex AI using a container is: First, build and push the container image to Artifact Registry. Next, register the model in Vertex AI Model Registry. Then, deploy the model to an endpoint using Vertex AI Endpoints.

Finally, test the endpoint by sending prediction requests to verify it works.

491
MCQmedium

A team wants to implement continuous training for their ML model. The pipeline should be triggered when new training data arrives in a Cloud Storage bucket. Which combination of services should they use?

A.Cloud Storage → BigQuery → Vertex AI Pipelines
B.Cloud Storage → Cloud Functions → Vertex AI Pipelines
C.Cloud Storage → Cloud Build → Vertex AI Pipelines
D.Cloud Storage → Cloud Scheduler → Vertex AI Pipelines
AnswerB

A Cloud Storage object-change event triggers a Cloud Functions function, which calls the Vertex AI API to launch the training pipeline. This event-driven chain satisfies the stem's requirement to retrain automatically when new training data lands in the bucket.

Why this answer

The correct combination is Cloud Storage → Cloud Functions → Vertex AI Pipelines. Cloud Storage hosts the new training data; a Cloud Function triggered by an object-finalize event (or Pub/Sub notification) detects the new data and invokes a Vertex AI Pipeline to retrain the model. This creates a fully event-driven continuous training pipeline without polling or manual intervention.

Exam trap

PMLE often tests the difference between event-driven triggers (Cloud Functions) and time-based triggers (Cloud Scheduler), catching candidates who pick Scheduler for data-arrival events.

How to eliminate wrong answers

Option A is wrong because BigQuery is a data warehouse, not an event trigger—it does not natively detect new Cloud Storage objects and trigger pipelines. Option C is wrong because Cloud Build is a CI/CD service for building and deploying code/containers, not for reacting to data arrival events. Option D is wrong because Cloud Scheduler is time-based (cron), not event-driven; it cannot trigger on new data arrival, only on a schedule.

492
MCQeasy

A company wants to monitor fairness of a model by evaluating performance metrics across demographic subgroups. They have ground truth labels stored in BigQuery. Which Vertex AI service should they use?

A.Vertex AI Model Monitoring
B.Vertex AI Prediction
C.Vertex AI Explainability
D.Vertex AI Model Evaluation
AnswerD

Vertex AI Model Evaluation computes performance metrics, including fairness and bias slices, by comparing model predictions against ground truth labels. Pointing it at the BigQuery label data lets the team evaluate metrics across demographic subgroups, satisfying the fairness monitoring requirement.

Why this answer

Vertex AI Model Evaluation is the service designed to assess model performance using ground truth labels, including slicing metrics across demographic subgroups to detect fairness issues. It computes metrics like precision, recall, and AUC per slice, directly supporting bias and fairness analysis. This makes it the correct choice when labels are available in BigQuery.

Exam trap

The trap is conflating Model Monitoring with Model Evaluation — candidates assume 'monitoring fairness' means Model Monitoring, but fairness against ground truth requires Model Evaluation's labeled metrics.

How to eliminate wrong answers

Option A (Vertex AI Model Monitoring) is wrong because it detects drift and skew in production data without requiring ground truth labels — it watches feature and prediction distributions, not accuracy against labels. Option B (Vertex AI Prediction) is wrong because it only serves predictions from a deployed model; it does not evaluate performance. Option C (Vertex AI Explainability) is wrong because it provides feature attributions (e.g., SHAP values) to explain why a model made a prediction, not fairness metrics against ground truth.

493
MCQmedium

A team uses Vertex AI Workbench notebooks for collaborative model development. They want to ensure that code changes are version-controlled, that multiple data scientists can work on the same notebook without conflicts, and that the environment is reproducible across team members. Which approach should they take?

A.Use a shared JupyterLab instance launched on a single VM; data scientists connect simultaneously.
B.Use Vertex AI Workbench managed notebooks with Git integration and a custom container image for environment reproducibility.
C.Store notebooks in Cloud Storage and share the bucket; each user edits their own copy.
D.Use Vertex AI Pipelines to run all code as pipelines; data scientists only view results in notebooks.
AnswerB

Git integration in managed notebooks provides branch-based version control and conflict resolution for concurrent editors, while a custom container image pins library versions so every data scientist gets an identical runtime. This satisfies the stem's three constraints: version-controlled changes, conflict-free collaboration, and reproducible environments across the team.

Why this answer

Vertex AI Workbench managed notebooks support Git integration for version control and custom container images for reproducible environments. This combination allows multiple data scientists to collaborate on notebooks with version history and ensures the runtime environment is consistent across team members.

Exam trap

The trap is thinking that sharing a VM or Cloud Storage bucket enables collaboration — the exam expects you to recognize that Git integration and custom containers are required for version control and reproducibility.

How to eliminate wrong answers

Option A is wrong because a shared JupyterLab instance on a single VM does not provide version control, conflict resolution, or environment reproducibility, and simultaneous editing can cause conflicts. Option C is wrong because storing notebooks in Cloud Storage and sharing copies lacks Git-based version control and does not prevent conflicts or ensure reproducibility. Option D is wrong because Vertex AI Pipelines is for orchestrating ML workflows, not for interactive collaborative notebook development with Git integration.

494
MCQmedium

An ML team wants to implement data versioning for large datasets stored in Google Cloud Storage. They need to track changes over time and reproduce previous data states. Which tool is most appropriate?

A.Cloud Storage Object Versioning
B.BigQuery table snapshots
C.Git LFS
D.DVC
AnswerD

DVC versions large datasets held in Google Cloud Storage by storing content hashes in small metafiles committed to Git, leaving the data in place. This delivers change tracking and reproducible checkouts of previous data states, exactly the capability the team requires.

Why this answer

DVC (Data Version Control) is purpose-built for versioning large datasets and ML models alongside Git. It stores small metadata files in Git while pushing the actual large data to remote storage like Google Cloud Storage, enabling teams to track dataset changes over time and reproduce exact previous data states via commits or tags.

Exam trap

The trap is assuming native GCS Object Versioning is sufficient for ML data versioning — candidates overlook that object-level versioning lacks dataset-level snapshots, lineage, and Git-integrated reproducibility that DVC provides.

How to eliminate wrong answers

Option A is wrong because Cloud Storage Object Versioning only retains prior versions of individual objects — it does not provide dataset-level versioning, lineage, or the ability to check out a coherent snapshot of an entire dataset at a point in time. Option B is wrong because BigQuery table snapshots version BigQuery tables, not datasets stored in GCS, and the question specifies data in Cloud Storage. Option C is wrong because Git LFS is designed for large binary files in Git repos but is not optimized for ML dataset workflows, lacks data-specific features like pipeline stage caching, and struggles with very large datasets.

495
MCQeasy

You need to run a batch prediction job on Vertex AI for a large dataset stored in BigQuery. The model expects CSV input. Which input format should you specify for the batch prediction job?

A.JSON Lines files on Cloud Storage
B.CSV files on Cloud Storage
C.TFRecord files on Cloud Storage
D.BigQuery table
AnswerB

For batch prediction, Vertex AI supports CSV files on Cloud Storage as an input format. Since the model expects CSV, providing CSV files on Cloud Storage directly matches the model's input requirements and avoids any format conversion or export step.

Why this answer

Vertex AI batch prediction supports multiple input formats, but the format must match what the model expects. Since the model expects CSV, specifying CSV files on Cloud Storage is the correct choice. This avoids any need for data conversion and ensures the batch prediction job can parse the input correctly.

Exam trap

The trap here is assuming that any supported input format works regardless of the model's expected input, when the format must match the model's signature.

496
MCQmedium

A team has a prototype image classification model trained on a small dataset using TensorFlow Keras on a single GPU. They need to train on a larger dataset (1 million images) using a distributed strategy on Vertex AI with 8 GPUs. They implement a MirroredStrategy for data parallelism. During the first few epochs, the training speed does not improve significantly compared to a single GPU, and GPU utilization is low. The data is stored as JPEG files in Cloud Storage, and the input pipeline uses tf.data with map to decode images. What is the most likely cause?

A.The batch size per GPU is too large.
B.The MirroredStrategy is not properly configured.
C.The data loading from Cloud Storage is a bottleneck.
D.The model is too small for distributed training.
AnswerC

Decoding JPEGs inside tf.data.map runs on CPU and reads from Cloud Storage over the network, starving eight GPUs. MirroredStrategy replicates compute but not input throughput, so low GPU utilisation persists until parallel interleave and prefetch are added.

Why this answer

Reading and decoding JPEG images from Cloud Storage can be I/O-bound, causing low GPU utilization. Even though MirroredStrategy is used, if the input pipeline cannot supply data fast enough, GPUs will spend time waiting for data. Option A is incorrect because a large batch size per GPU typically increases memory usage and computational load, not decreases utilization; it might even improve utilization if the model fits.

Option B is incorrect because MirroredStrategy is a straightforward configuration for synchronous data parallelism and does not require complex tuning; it is likely configured correctly. Option D is incorrect because the model size does not directly cause low GPU utilization; even a small model can saturate GPUs if the data pipeline is efficient. The primary bottleneck here is the data loading from Cloud Storage.

497
MCQhard

A company uses Vertex AI Predictions with a custom container that invokes an external API for feature enrichment. The prediction response time is highly variable. The engineer wants to monitor the external API's contribution to latency. What should the engineer do?

A.Instrument the prediction container to emit custom metrics for the time spent in each prediction step, including the external API call.
B.Add a timeout setting to the endpoint's request to limit the external API call duration.
C.Monitor the Vertex AI endpoint latency metric and correlate with system metrics like CPU and memory.
D.Use Cloud Trace to trace the prediction request end-to-end, including the external API call.
AnswerA

Instrumenting the container to emit custom metrics isolates the external API call's duration from total prediction latency, satisfying the requirement to monitor that specific contribution. Vertex AI's built-in metrics only capture overall request latency, so per-step timing must be emitted by the container itself.

Why this answer

Instrumenting the custom container to emit custom metrics (e.g., using OpenTelemetry or a Prometheus client library) allows the engineer to directly measure the time spent in each prediction step, isolating the external API call's contribution to latency. This provides granular, real-time visibility into the specific bottleneck, which is essential when the response time is highly variable and the external API is a known dependency.

Exam trap

Google Cloud often tests the distinction between monitoring (custom metrics) and tracing (Cloud Trace) — the trap here is that candidates assume Cloud Trace automatically captures all downstream calls, but it requires explicit instrumentation of the external API call to record its duration, whereas custom metrics can be emitted directly from the container code without needing distributed tracing context.

How to eliminate wrong answers

Option B is wrong because adding a timeout setting to the endpoint's request limits the duration of the external API call but does not provide any monitoring data; it only caps the latency, potentially causing failures without diagnosing the root cause. Option C is wrong because monitoring the Vertex AI endpoint latency metric and correlating with CPU/memory only gives aggregate performance data, not the specific contribution of the external API call, making it impossible to isolate the external API's impact. Option D is wrong because Cloud Trace can trace the request end-to-end, but it requires the custom container to be instrumented with trace context propagation; without explicit instrumentation of the external API call, Cloud Trace will not capture the time spent in that external call, leaving the same gap in visibility.

498
MCQmedium

A company wants to build a product recommendation engine for their e-commerce website. They have historical purchase data and user interaction logs. They want a managed service that can quickly generate personalized recommendations without building custom models. Which service should they use?

A.Dataflow with TensorFlow
B.BigQuery ML with MATRIX_FACTORIZATION
C.AutoML Tables
D.Recommendations AI
AnswerD

Recommendations AI is a managed Google Cloud service that trains on historical purchase and interaction data to serve personalised product recommendations, satisfying the requirement to avoid building custom models. It handles model training, tuning and serving, so the team gains personalisation without ML engineering effort.

Why this answer

Recommendations AI is a fully managed Google Cloud service purpose-built for personalized product recommendations, ingesting purchase and interaction data and serving recommendations via API without custom model development. It handles training, tuning, and serving, which matches the requirement for a managed service that quickly generates personalized recommendations.

Exam trap

PMLE often tests the confusion between general ML services (AutoML, BigQuery ML) and purpose-built recommendation services, where only the latter meets a 'no custom models' requirement.

How to eliminate wrong answers

Option A is wrong because Dataflow with TensorFlow requires building and operating a custom recommendation model, which contradicts the 'without building custom models' requirement. Option B is wrong because BigQuery ML with MATRIX_FACTORIZATION requires writing SQL to train and tune a model and does not provide a turnkey recommendation serving API. Option C is wrong because AutoML Tables is a general tabular ML service, not a recommendation-specific product, and would still require significant feature engineering and serving infrastructure.

499
MCQmedium

A financial services company wants to detect fraudulent transactions in real-time. They have a trained XGBoost model that runs on a single Compute Engine instance. The current solution processes about 100 transactions per second, but they need to scale to 10,000 transactions per second. Which approach should they take?

A.Increase the VM to a machine type with more vCPUs and memory
B.Deploy the model to Vertex AI Prediction with autoscaling enabled
C.Use Dataflow to process transactions in micro-batches every second
D.Rewrite the model as a Cloud Function triggered by Pub/Sub messages
AnswerB

Vertex AI Prediction autoscaling distributes inference across managed replicas, so throughput scales horizontally beyond a single Compute Engine instance's ceiling. This satisfies the 10,000 transactions per second requirement while preserving the trained XGBoost model, which Vertex AI serves natively as a pre-built container.

Why this answer

Vertex AI Prediction with autoscaling is the correct choice because it is purpose-built for serving ML models at scale, automatically adjusting the number of compute nodes based on incoming request traffic. This allows the company to seamlessly handle the increase from 100 to 10,000 transactions per second without manual intervention, while XGBoost is natively supported as a framework.

Exam trap

Google Cloud often tests the misconception that vertical scaling (bigger VM) is sufficient for large throughput increases, when in reality horizontal scaling with a managed service like Vertex AI is required for elasticity and high availability.

How to eliminate wrong answers

Option A is wrong because simply scaling up a single VM (vertical scaling) has hardware limits and cannot reliably handle a 100x increase in throughput; it also introduces a single point of failure and lacks automatic scaling. Option C is wrong because Dataflow is designed for batch and stream processing pipelines, not for low-latency real-time model serving; processing in micro-batches every second would add unacceptable latency for fraud detection. Option D is wrong because Cloud Functions have a maximum timeout of 9 minutes and are not designed for sustained high-throughput inference workloads; they are better suited for lightweight, event-driven tasks, not for serving a complex XGBoost model at 10,000 TPS.

500
MCQhard

An ML engineer is designing a Vertex AI Pipeline that includes a hyperparameter tuning step. The tuning step runs multiple trials and outputs the best model. The engineer wants to ensure that the pipeline can resume from the tuning step if it fails later, without re-running all the tuning trials. What should the engineer do?

A.Configure the tuning step to use a Vertex AI HyperparameterTuningJob with a persistent resource pool.
B.Store the tuning results in a Cloud Storage bucket and add a conditional step that checks for existing results before running tuning.
C.Enable caching for the tuning step by setting the component's 'enable_caching' attribute to True.
D.Set the pipeline's 'resume' option to True when submitting the pipeline.
AnswerC

Vertex AI Pipelines supports caching of component executions. When caching is enabled, the pipeline stores the outputs of a component based on its inputs and code. If the pipeline is re-run and the tuning step's inputs have not changed, the cached outputs (including the best model) are reused, avoiding re-execution of all trials. This is the correct way to resume without re-running tuning.

Why this answer

The correct method is to enable caching for the tuning component. Vertex AI Pipelines caching stores the outputs of a component based on a hash of its inputs, code, and environment. When the pipeline is re-run, if the tuning step's inputs are unchanged, the cached outputs are used, and the tuning trials are not re-executed.

This saves time and cost, and allows the pipeline to resume from the tuning step.

Exam trap

The trap here is assuming that Vertex AI Pipelines has a built-in 'resume' feature or that manual result storage is needed, when caching is the intended mechanism.

501
Multi-Selecteasy

Which TWO actions are best practices when scaling a prototype ML model to production in Google Cloud?

Select 2 answers
A.Store and manage features in a feature store like Vertex AI Feature Store.
B.Test the model only on a small sample of the production data to save costs.
C.Set up monitoring and logging for model performance and data drift.
D.Manually scale inference instances based on historical traffic patterns.
E.Use one-hot encoding for all categorical features without considering cardinality.
AnswersA, C

Vertex AI Feature Store provides a centralised repository that serves consistent feature values for both training and online prediction, eliminating training-serving skew. It also enables feature reuse across models and point-in-time correctness, which is essential when moving a prototype into a production pipeline.

Why this answer

Option A is correct because storing and managing features in Vertex AI Feature Store provides a centralized, consistent source of feature values for both training and serving, which prevents training/serving skew and enables feature reuse and low-latency online serving at production scale. Option C is correct because production ML requires continuous monitoring and logging of model performance metrics and data drift (for example via Vertex AI Model Monitoring), so degradation, skew, and drift can be detected and the model retrained or rolled back. Option B is not a best practice because validating only on a small sample of production data risks missing edge cases and distribution shifts; production readiness requires representative validation and testing.

Option D is not a best practice because manual scaling is error-prone and slow; production inference should use autoscaling based on real-time metrics such as CPU/GPU utilization or request load. Option E is not a best practice because applying one-hot encoding to all categorical features regardless of cardinality can create extremely high-dimensional sparse vectors, causing memory and training problems; high-cardinality features should use embeddings or hashing instead.

Exam trap

Google Cloud often tests the misconception that cost-saving shortcuts like limited testing or manual scaling are acceptable in production, when in fact reliability and monitoring are non-negotiable for ML systems at scale.

502
MCQmedium

A team is building a fraud detection model that requires joining real-time transaction features with historical user features. They need to ensure that the training data does not use future information (data leakage). Which Vertex AI Feature Store capability should they use?

A.Online store serving with Bigtable
B.Feature store time travel
C.Point-in-time correct join
D.Feature monitoring for drift
AnswerC

Point-in-time correct join retrieves feature values as they existed at each training example's timestamp, preventing future information from leaking into training. This satisfies the stem's requirement that training data avoid data leakage when joining real-time and historical features.

Why this answer

Point-in-time correct joins in Vertex AI Feature Store ensure that when training examples are generated, each row uses only feature values that were valid as of the event timestamp of the label — preventing future data from leaking into training. This is the specific capability designed to avoid label leakage in time-series or event-driven ML. Time travel and online serving do not by themselves guarantee temporal correctness of the join.

Exam trap

PMLE often tests the confusion between time travel (retrieving historical values) and point-in-time correct joins (aligning features to label timestamps to prevent leakage) — candidates pick time travel thinking it solves leakage.

How to eliminate wrong answers

Option A is wrong because online store serving with Bigtable is about low-latency retrieval of the latest feature values for real-time inference — it does not enforce temporal alignment between features and labels during training. Option B is wrong because feature store time travel lets you retrieve historical feature values as of a past timestamp, but it does not automatically align each training row's features to the label's event time; you still need a point-in-time join to prevent leakage. Option D is wrong because feature monitoring for drift detects distribution changes in served features over time — it is an observability capability, not a training-data correctness mechanism.

503
Multi-Selecthard

A team uses Vertex AI Pipelines and wants to track lineage of artifacts and executions. Which three resources should they use? (Choose three.)

Select 3 answers
A.Artifacts
B.Vertex AI Experiments
C.Vertex AI Metadata
D.Model Registry
E.Executions
AnswersA, C, E

Artifacts are the versioned, typed outputs and inputs that Vertex AI ML Metadata records, such as datasets, models and metrics. Naming them satisfies the lineage requirement because each Artifact node links to the Executions and Contexts that produced or consumed it, forming the traceable graph.

Why this answer

Vertex AI Metadata is the core lineage service that stores and connects metadata about ML resources, so option C is correct because it provides the underlying metadata store for tracking lineage. Within that metadata store, Artifacts (option A) represent the inputs and outputs of pipeline steps—such as datasets, models, and metrics—and are the nodes whose lineage is tracked. Executions (option E) represent a single run of a pipeline step or component and record the events that consume and produce artifacts, which is exactly what links artifacts together into a lineage graph.

Vertex AI Experiments (option B) is for tracking and comparing experiment runs and metrics, not for artifact/execution lineage, and Model Registry (option D) is for managing model versions and deployment, not for general lineage tracking.

Exam trap

PMLE often tests the distinction between Metadata (lineage: artifacts/executions) and Experiments (runs/metrics) — candidates pick Experiments because it sounds like the tracking service.

504
MCQmedium

An ML engineer is using Vertex AI Vizier to tune hyperparameters for a PyTorch model. They want to maximise the chance of finding the global optimum within a fixed trial budget of 50 trials. Which algorithm should they select?

A.Random search
B.Bayesian optimisation
C.Grid search
D.Evolutionary algorithm
AnswerB

Bayesian optimisation builds a probabilistic surrogate model of the objective and uses an acquisition function to pick each next trial, balancing exploration against exploitation. Over a fixed 50-trial budget this converges on the global optimum faster than grid or random search, directly satisfying the budget constraint.

Why this answer

Bayesian optimisation (option B) is the correct choice because it builds a probabilistic surrogate model of the objective function and uses an acquisition function to balance exploration and exploitation, making it highly sample-efficient. With only 50 trials, Bayesian optimisation maximises the probability of finding the global optimum by focusing trials on the most promising hyperparameter regions, unlike random or grid search which waste trials on unpromising areas.

Exam trap

The trap here is that candidates often choose random search (option A) because they recall it is better than grid search for high-dimensional spaces, but they overlook that Bayesian optimisation is strictly more sample-efficient and is the default recommendation in Vertex AI Vizier for maximising global optimum discovery under a fixed trial budget.

How to eliminate wrong answers

Option A is wrong because random search, while better than grid search in high-dimensional spaces, does not use past trial results to guide future trials, so it wastes trials on suboptimal regions and has a lower probability of finding the global optimum within a fixed budget of 50 trials. Option C is wrong because grid search exhaustively evaluates a fixed set of points, which scales exponentially with the number of hyperparameters and is extremely inefficient for more than a few parameters, often missing the global optimum entirely within a limited budget. Option D is wrong because evolutionary algorithms (e.g., genetic algorithms) are population-based and require many generations to converge, typically needing hundreds or thousands of trials to be effective, making them impractical for a tight budget of 50 trials.

505
MCQhard

A real-time recommendation model deployed on Vertex AI Endpoints is experiencing increased latency, especially during peak hours. The model is hosted on a single machine with 4 CPUs. Which set of actions should you take to diagnose and resolve the issue?

A.Increase the machine type to with 32 CPUs and disable autoscaling.
B.Switch the endpoint to use GPUs and enable batch requests.
C.Enable autoscaling on the endpoint and analyze request patterns to set min/max instances.
D.Change the serving framework to use TensorFlow Serving with gRPC.
AnswerC

Autoscaling adds replicas so peak-hour traffic spreads across more machines, removing the single 4-CPU bottleneck; analysing request patterns then sets min/max instances to match demand. This addresses the latency constraint by scaling horizontally rather than resizing one machine.

Why this answer

Enabling autoscaling on a Vertex AI Endpoint allows the deployment to dynamically adjust the number of serving instances based on real-time traffic, directly addressing peak-hour latency. Analyzing request patterns to set appropriate min/max instances ensures that the endpoint scales proactively without over-provisioning, which is the standard diagnostic and resolution approach for latency issues caused by insufficient capacity under variable load.

Exam trap

Google Cloud often tests the misconception that scaling up (vertical scaling) or changing frameworks is the first step to fix latency, when the correct approach is to first diagnose capacity constraints and then scale out horizontally with autoscaling.

How to eliminate wrong answers

Option A is wrong because simply increasing the machine type to 32 CPUs without autoscaling does not resolve peak-hour latency; it only increases static capacity, leading to over-provisioning during low traffic and still failing under sudden spikes if the single instance is overwhelmed. Option B is wrong because switching to GPUs is not a direct fix for latency caused by CPU-bound serving; GPUs benefit compute-heavy models (e.g., deep learning) but add overhead for small models, and enabling batch requests increases latency for real-time predictions as it waits to accumulate requests. Option D is wrong because changing the serving framework to TensorFlow Serving with gRPC does not address the root cause of insufficient compute capacity; it may improve throughput per instance but cannot compensate for a single machine being overloaded during peak hours.

506
MCQhard

Your team is training a very large transformer model that does not fit on a single GPU. They are using Vertex AI custom training with PyTorch. Which distributed training approach should they use?

A.Data parallelism using PyTorch DistributedDataParallel (DDP)
B.Horovod with allreduce
C.Model parallelism using pipeline parallelism
D.Vertex AI distributed training with TF_CONFIG
AnswerC

Pipeline parallelism splits the transformer's layers across GPUs, so each device holds only a subset of parameters and activations. This addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

Why this answer

When a transformer model is too large to fit on a single GPU, model parallelism (specifically pipeline parallelism) is required because it splits the model's layers across multiple devices, with each device holding a subset of the model's parameters. Data parallelism (DDP) replicates the entire model on each GPU, which fails if the model exceeds a single GPU's memory. Pipeline parallelism allows training very large models by partitioning the model into stages and passing activations and gradients sequentially between devices.

Exam trap

Google Cloud often tests the distinction between data parallelism (which replicates the model) and model parallelism (which splits the model), and the trap here is that candidates assume any distributed training framework (like DDP or Horovod) can handle oversized models, ignoring the fundamental memory constraint that data parallelism cannot overcome.

How to eliminate wrong answers

Option A is wrong because PyTorch DistributedDataParallel (DDP) implements data parallelism, which requires the entire model to fit on each GPU; if the model is too large for one GPU, DDP cannot be used. Option B is wrong because Horovod with allreduce is also a data-parallel approach that replicates the model on every worker, suffering the same memory limitation as DDP. Option D is wrong because Vertex AI distributed training with TF_CONFIG is a configuration mechanism for TensorFlow-based distributed training (using MirroredStrategy or MultiWorkerMirroredStrategy), not a PyTorch-native approach, and it still relies on data parallelism unless combined with model parallelism; the question specifies PyTorch, making this option technically incompatible.

507
MCQmedium

An MLOps team needs to automatically retrain a model when new training data becomes available. They use Vertex AI Pipelines. What is the recommended way to trigger the pipeline?

A.Use Model Evaluation to decide
B.Set up a trigger in Vertex AI Pipelines
C.Cloud Functions triggered by Cloud Storage events
D.Cloud Scheduler on a daily basis
AnswerC

Cloud Storage object-finalise events routed through Eventarc invoke a Cloud Function, which then submits the Vertex AI Pipeline job. This satisfies the stem's requirement to retrain automatically when new training data lands, since the trigger fires on data arrival rather than on a schedule or manual invocation.

Why this answer

Vertex AI Pipelines does not natively support event-driven triggers. The recommended pattern is to use Cloud Functions, which can be triggered by Cloud Storage events (e.g., object finalize/create) when new training data is uploaded. The Cloud Function then programmatically submits the pipeline run via the Vertex AI Pipelines client library or REST API, enabling an automated retraining workflow.

Exam trap

The trap here is that candidates assume Vertex AI Pipelines has a built-in trigger mechanism (Option B) because many CI/CD tools do, but Google Cloud's recommended pattern relies on external event-driven services like Cloud Functions.

How to eliminate wrong answers

Option A is wrong because Model Evaluation is a post-training assessment step, not a trigger mechanism; it cannot initiate pipeline execution. Option B is wrong because Vertex AI Pipelines itself does not provide a built-in trigger; triggers must be implemented externally via Cloud Functions, Cloud Scheduler, or similar services. Option D is wrong because Cloud Scheduler on a daily basis is a time-based trigger, not an event-driven one; it would retrain on a fixed schedule regardless of whether new data has arrived, leading to unnecessary runs or missed retraining opportunities.

508
Multi-Selecteasy

Which TWO statements about Vertex AI Feature Store are correct? (Choose 2)

Select 2 answers
A.Feature Store automatically applies feature engineering transformations.
B.Feature Store can only store numerical features.
C.Feature Store can only be used with Vertex AI models.
D.Feature Store provides a centralized repository for feature data.
E.Feature Store supports both online and offline serving.
AnswersD, E

Vertex AI Feature Store acts as a centralised repository, letting teams register, version and share feature data across projects and models instead of duplicating feature engineering per pipeline, which is the defining architectural property this statement asserts.

Why this answer

Option D is correct because Vertex AI Feature Store acts as a centralized repository where feature data is registered, versioned, and managed as feature groups and features, so multiple teams and models can share a single consistent source of features. Option E is correct because Feature Store supports both online serving, which returns low-latency feature values for real-time predictions, and offline serving, which reads historical feature values in bulk for training and batch scoring. Options A, B, and C are incorrect: Feature Store does not automatically perform feature engineering transformations (you define and ingest the transformed values yourself), it is not limited to numerical features (it supports types such as strings, booleans, and arrays as well), and it is not restricted to Vertex AI models since features can be served to any application or model via the online/offline APIs.

Exam trap

Google Cloud often tests the misconception that Vertex AI Feature Store is tightly coupled to Vertex AI models or that it performs automatic feature engineering, when in fact it is a decoupled storage and serving layer that supports any ML framework and requires explicit feature engineering steps.

509
MCQeasy

An ML team is using Vertex AI Online Prediction and wants to receive alerts when the 99th percentile latency exceeds 500ms for more than 5 minutes. What is the best practice to set up this alert in Cloud Monitoring?

A.Create a custom metric from the prediction container that emits latency percentiles, then set an alert on that metric.
B.Use the 'aiplatform.googleapis.com/prediction/online_prediction_latencies' metric with a metric threshold condition set to 500ms and a percentile aligner of 99.
C.Use a log-based metric to parse latency from Cloud Logging and alert when the average exceeds 500ms.
D.Export prediction latency logs to BigQuery and run a scheduled query to check the 99th percentile, then trigger a Cloud Function to send an alert.
AnswerB

The online_prediction_latencies metric exposes serving latency histograms; applying a 99th percentile aligner isolates tail latency, and a 500ms threshold condition with a five-minute duration triggers only on sustained breaches, matching the stem's alerting requirement precisely.

Why this answer

Cloud Monitoring provides a pre-built metric, `aiplatform.googleapis.com/prediction/online_prediction_latencies`, which directly captures prediction latency. By applying a percentile aligner of 99 and a metric threshold condition of 500ms, you can alert when the 99th percentile latency exceeds 500ms for the specified duration, without needing custom instrumentation or external processing.

Exam trap

Google Cloud often tests the misconception that you must create custom metrics or use log-based solutions for percentile-based alerting, when in fact Cloud Monitoring's distribution metrics and percentile aligners handle this natively.

How to eliminate wrong answers

Option A is wrong because creating a custom metric from the prediction container is unnecessary and adds complexity; Vertex AI already emits the required latency metric natively, and custom metrics would require additional code and maintenance. Option C is wrong because using a log-based metric to parse latency from Cloud Logging and alerting on the average (not the 99th percentile) does not meet the requirement to monitor the 99th percentile latency; log-based metrics also introduce latency and parsing overhead. Option D is wrong because exporting logs to BigQuery and running scheduled queries is an overly complex, non-real-time approach that violates the best practice of using built-in monitoring capabilities; it also introduces additional cost and delay compared to native Cloud Monitoring alerts.

510
MCQhard

A data engineering team needs to compute rolling window features (7-day average, 30-day sum) from a high-volume stream of e-commerce events stored in BigQuery. They must output the features to Vertex AI Feature Store for online serving. Which approach is MOST cost-effective and scalable?

A.Use Cloud Composer (Airflow) with a daily DAG to run SQL queries on BigQuery and export results
B.Schedule a query in BigQuery using scheduled queries and export results to Feature Store
C.Use Cloud Functions triggered by Pub/Sub to compute features on the fly
D.Use Dataflow with Apache Beam, reading from BigQuery, computing windowed aggregations, and writing to Vertex AI Feature Store
AnswerD

Dataflow with Apache Beam provides managed, autoscaling stream processing that reads from BigQuery, applies windowed aggregations for the 7-day and 30-day features, and writes directly to Vertex AI Feature Store. This avoids custom serving infrastructure while meeting the online-serving and scalability constraints.

Why this answer

Dataflow (Apache Beam) is ideal for processing large-scale batch and streaming data. It can read from BigQuery, perform windowing computations, and write to Feature Store's online store. Cloud Functions have timeouts, Cloud Composer is not optimal for streaming, and BigQuery scheduled queries are not designed for streaming-feature computation.

511
MCQmedium

You have a champion model serving 100% traffic on a Vertex AI endpoint. You want to deploy a challenger model and gradually shift 10% of traffic to it for A/B testing. What is the correct approach?

A.Use Cloud Run to deploy both models and use Cloud Endpoints for traffic splitting.
B.Deploy the challenger on the same endpoint and use the traffic split parameter to allocate 10% traffic to it.
C.Deploy the challenger on a separate endpoint and use Cloud Armor to split traffic.
D.Create a new endpoint for the challenger and route 10% of requests via a load balancer.
AnswerB

Deploying the challenger to the same endpoint lets Vertex AI distribute prediction requests across both models using its built-in traffic split, which accepts percentage allocations per deployed model ID. Setting the challenger's split to 10 satisfies the gradual A/B requirement without a second endpoint, DNS changes, or client-side routing logic.

Why this answer

Vertex AI endpoints support traffic splitting by deploying multiple model versions and assigning traffic percentages. You deploy the challenger as a new deployed model on the same endpoint and set traffic split: champion 90%, challenger 10%.

512
Multi-Selectmedium

Which THREE are best practices for implementing CI/CD for ML pipelines on Google Cloud? (Choose THREE.)

Select 3 answers
A.Maintain separate environments for dev, staging, and production
B.Track all experiments and artifacts using Vertex ML Metadata
C.Use Cloud Build to automate testing, building, and deployment of pipeline components
D.Design pipelines with low-code components to reduce development time
E.Write unit tests for every training job
AnswersA, B, C

Separate dev, staging, and production environments isolate pipeline changes so untested components never reach production data or endpoints. This satisfies CI/CD best practice by enabling promotion gates, reproducible validation, and rollback, preventing a faulty model or pipeline from affecting live traffic.

Why this answer

Option A is correct because maintaining separate dev, staging, and production environments isolates changes and lets you validate ML pipeline updates on representative data before promoting them to production, which is a core CI/CD practice on Google Cloud. Option B is correct because Vertex ML Metadata records experiments, parameters, metrics, and artifacts, giving the lineage and reproducibility needed to compare runs and roll back or promote models reliably. Option C is correct because Cloud Build is Google Cloud's managed CI/CD service and can automate testing, building, and deploying pipeline components (for example, via cloudbuild.yaml and Vertex AI Pipelines), which is exactly the automation CI/CD requires.

Option D is not a CI/CD best practice per se; low-code components may speed development but do not address continuous integration, delivery, or reproducibility. Option E is too narrow and absolute: unit tests for training jobs are useful, but writing them for every training job is not one of the three defining best practices for CI/CD on Google Cloud.

Exam trap

Google Cloud often tests the distinction between general software CI/CD practices and ML-specific CI/CD needs, trapping candidates who over-apply traditional unit testing or assume low-code tools are always best practices for production ML pipelines.

513
MCQhard

A company uses BigQuery ML to train a boosted tree classifier on a large dataset. After training, they want to understand which features most influence predictions. Which BigQuery ML function should they use?

A.ML.EXPLAIN_PREDICT
B.ML.FEATURE_IMPORTANCE
C.ML.EVALUATE
D.ML.PREDICT
AnswerB

ML.FEATURE_IMPORTANCE returns a per-feature score showing how strongly each input influenced the trained boosted tree model's predictions. It is the BigQuery ML function designed for interpreting model behaviour, satisfying the requirement to identify the most influential features.

Why this answer

ML.FEATURE_IMPORTANCE is the BigQuery ML function specifically designed to return the relative importance of each input feature for a trained model, including boosted tree classifiers (Boosted Tree, XGBoost, Random Forest). It computes importance scores using the model's internal split-gain or weight-based metrics, giving a ranked list of features that most influence predictions. This is the direct, purpose-built answer for global feature attribution on a trained model.

Exam trap

PMLE often tests the confusion between global feature importance (ML.FEATURE_IMPORTANCE) and local per-prediction explanations (ML.EXPLAIN_PREDICT) — candidates pick EXPLAIN_PREDICT because it sounds more 'explanatory,' but it returns Shapley values per row, not a global ranking.

How to eliminate wrong answers

Option A is wrong because ML.EXPLAIN_PREDICT returns per-row local explanations (Shapley values) for individual predictions, not a global ranked feature-importance list — it answers 'why this prediction' rather than 'which features matter overall.' Option C is wrong because ML.EVALUATE returns model quality metrics (accuracy, precision, recall, AUC, log loss) on a dataset, not feature attribution. Option D is wrong because ML.PREDICT generates predictions on new data; it provides no insight into feature influence.

514
MCQeasy

A data science team has deployed a model on Vertex AI and wants to automatically detect when the distribution of a specific feature shifts significantly from the training data. Which service should they use?

A.Cloud Data Loss Prevention
B.Vertex AI Model Monitoring
C.Vertex AI Explainable AI
D.Cloud Composer
AnswerB

Vertex AI Model Monitoring computes feature distribution statistics on live traffic and compares them with the training baseline, automatically flagging significant shifts. This satisfies the requirement to detect distribution changes for a specific feature without building custom drift-detection pipelines.

Why this answer

Vertex AI Model Monitoring is the correct service because it is specifically designed to detect feature distribution drift (skew) between training and serving data for deployed models. It continuously monitors the input features and alerts when statistical metrics like the Jensen-Shannon divergence or the L-infinity distance exceed a configured threshold, enabling proactive model retraining.

Exam trap

Google Cloud often tests the distinction between monitoring model performance (e.g., accuracy, latency) versus monitoring data distribution drift, and candidates may confuse Vertex AI Model Monitoring with Explainable AI because both involve model analysis, but only Model Monitoring tracks shifts over time.

How to eliminate wrong answers

Option A is wrong because Cloud Data Loss Prevention (DLP) is used for inspecting, classifying, and de-identifying sensitive data (e.g., PII, credit card numbers), not for monitoring feature distributions or model drift. Option C is wrong because Vertex AI Explainable AI provides feature attributions and explanations for model predictions (e.g., Shapley values, integrated gradients), but does not monitor distribution shifts over time. Option D is wrong because Cloud Composer is a managed Apache Airflow service for orchestrating workflows and pipelines, not a dedicated tool for detecting feature drift in deployed models.

515
MCQmedium

A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

A.Package the model with Triton in a custom container and deploy it to Vertex AI
B.Vertex AI automatically uses Triton for all PyTorch models
C.Deploy the model to GKE and use Vertex AI as a frontend
D.Use a prebuilt Vertex AI PyTorch container that includes Triton
AnswerA

Vertex AI lets you supply a custom prediction container, so bundling the PyTorch model with Triton and its model repository into that image gives Triton full control of inference, enabling its batching and optimisation features on Vertex AI's managed endpoint infrastructure.

Why this answer

Vertex AI supports custom containers for prediction, so the standard integration pattern is to build a container that bundles the model with NVIDIA Triton Inference Server and deploy it as a Vertex AI Model with a custom prediction container. This lets Triton handle dynamic batching, model ensembles, and multi-framework serving while Vertex AI manages endpoints, autoscaling, and monitoring.

Exam trap

PMLE often tests the misconception that managed services like Vertex AI automatically apply third-party optimizers (Triton, TensorRT) to any framework — in reality, integration requires the candidate to package the optimizer inside a custom container.

How to eliminate wrong answers

Option B is wrong because Vertex AI does not automatically wrap PyTorch models with Triton — the default prebuilt PyTorch containers use TorchServe or a simple Flask-based handler, not Triton. Option C is wrong because deploying to GKE and using Vertex AI as a 'frontend' is not a supported integration pattern; Vertex AI endpoints run on managed infrastructure, not on customer GKE clusters as a frontend proxy. Option D is wrong because Google's prebuilt Vertex AI PyTorch containers do not include Triton Inference Server — Triton must be added by the customer in a custom container.

516
MCQmedium

You are deploying a custom PyTorch model to a Vertex AI Endpoint for real-time inference. The model artifact is stored in a Cloud Storage bucket. Your security team requires that the model be served from a container that runs as a non-root user and has no network access except to the Vertex AI prediction service. Which deployment configuration should you use?

A.Deploy the model using a custom container that runs as root but configure the endpoint to use a private VPC with no external IP.
B.Use the pre-built PyTorch container provided by Vertex AI and set the endpoint to use a private VPC with no external IP.
C.Deploy the model using a custom container that sets the USER instruction to a non-root user in the Dockerfile, and configure the endpoint to use a private VPC with no external IP.
D.Deploy the model using a custom container that sets the USER instruction to a non-root user, but allow the endpoint to have an external IP for easier debugging.
AnswerC

A custom container allows you to control the user and network settings. Setting USER to a non-root user in the Dockerfile satisfies the non-root requirement, and deploying the endpoint in a private VPC with no external IP restricts network access to only the Vertex AI prediction service. This meets both security constraints.

Why this answer

A custom container is necessary to control the user context, and deploying it in a private VPC with no external IP ensures network isolation. The pre-built container runs as root, and allowing an external IP breaks the network restriction. Thus, the configuration that combines a non-root custom container with a private VPC is the only one that satisfies both security requirements.

Exam trap

The trap here is assuming that Vertex AI's pre-built containers automatically run as non-root or that network isolation alone is sufficient without controlling the container user.

517
MCQmedium

An MLOps team is implementing a CI/CD pipeline for a TensorFlow model on Vertex AI. The model training job takes 2 hours and produces a SavedModel. The team wants to automatically trigger a new pipeline run whenever a change is pushed to the 'main' branch of their source repository. The pipeline should include training, evaluation, and if metrics exceed a threshold, deploy the model to a Vertex AI endpoint. Which trigger configuration should they use?

A.Use Eventarc to listen for Cloud Source Repository push events and invoke a Cloud Run service that starts the pipeline.
B.Use an Artifact Registry trigger to detect new model images and then start the pipeline.
C.Set up a Cloud Scheduler job that runs every 2 hours and triggers a Vertex AI Pipeline run.
D.Configure a Cloud Build trigger that watches the 'main' branch of Cloud Source Repositories; in the build config, use steps to run the pipeline via the Vertex AI API.
AnswerD

A Cloud Build trigger watching the 'main' branch fires on each push, and its build config calls the Vertex AI API to run the training, evaluation and conditional deployment pipeline. This satisfies the automatic-on-commit trigger requirement.

Why this answer

Cloud Build triggers can be configured to watch a specific branch (e.g., 'main') in Cloud Source Repositories and automatically execute a build configuration. Within that build config, you can use the `gcloud` or `curl` steps to invoke the Vertex AI Pipeline API, which starts the training, evaluation, and conditional deployment workflow. This directly matches the requirement for a branch-based push trigger that orchestrates the full ML pipeline.

Exam trap

Google Cloud often tests the distinction between event-driven triggers (Cloud Build for source code changes) and artifact-based triggers (Artifact Registry for new images), leading candidates to confuse the two when the requirement is to start a pipeline from a code push.

How to eliminate wrong answers

Option A is wrong because Eventarc is designed for event-driven, asynchronous invocations (e.g., from Cloud Storage or Pub/Sub), but it does not natively integrate with Cloud Source Repositories push events; Cloud Build triggers are the correct service for repository push events. Option B is wrong because an Artifact Registry trigger would fire only after a new model image is pushed, but the requirement is to trigger on a source code change (push to 'main'), not on a new artifact. Option C is wrong because a Cloud Scheduler job running every 2 hours is a time-based schedule, not a push-triggered event; it would not respond to code changes and would run even when no changes occur, wasting resources.

518
MCQmedium

You have a trained XGBoost model that you want to deploy on Vertex AI for online prediction. The model expects input features in a specific order and requires a custom preprocessing step that normalizes numerical features using statistics computed during training. You need to ensure that the same preprocessing is applied at serving time. What should you do?

A.Use Vertex AI's built-in preprocessing feature by specifying a preprocessing function in the model's metadata when uploading.
B.Export the model as a SavedModel and include the preprocessing logic in the model's serving signature using TensorFlow's preprocessing layers.
C.Preprocess the input data on the client side before sending requests to the Vertex AI endpoint, using the same statistics.
D.Deploy the model using a custom container that includes the preprocessing code and the trained model artifacts, and implement the preprocessing in the container's prediction handler.
AnswerD

A custom container gives you full control over the serving logic. You can load the model and the training statistics (e.g., mean and standard deviation) and apply the exact same preprocessing in the prediction handler before passing data to the model. This ensures consistency between training and serving, which is critical for model performance.

Why this answer

For custom preprocessing with non-TensorFlow models like XGBoost, a custom container is the most reliable way to ensure that the exact same preprocessing is applied at serving time. You can embed the training statistics and logic in the container, maintaining consistency and simplifying client requests.

Exam trap

The trap here is assuming that Vertex AI provides built-in preprocessing for custom models, when in fact you must implement it yourself, often via a custom container.

519
MCQmedium

You are training a large tabular model on Vertex AI using a custom training job. The dataset is stored in BigQuery and is several terabytes in size. Training reads the same data for many epochs, and reading directly from BigQuery each epoch is slow and expensive. You want to maximize training throughput while keeping the data accessible to the training container. What should you do?

A.Enable Vertex AI Pipelines caching so that the BigQuery read step is not re-executed on subsequent training runs.
B.Export the BigQuery table to Cloud Storage in TFRecord or Parquet format, then read the exported files from the training container using the appropriate data loader.
C.Increase the number of worker replicas in the training job so that each worker reads a smaller portion of the BigQuery table in parallel.
D.Keep reading directly from BigQuery each epoch and rely on the BigQuery Storage Read API to stream rows to the training container.
AnswerB

Exporting to Cloud Storage once and reading the files repeatedly avoids repeated BigQuery scans and network round trips, and Cloud Storage offers high aggregate throughput for large sequential reads. This is the standard pattern for scaling prototype training on Vertex AI when the source is BigQuery and the job needs multiple epochs over multi-terabyte data.

Why this answer

The most efficient approach is to materialize the BigQuery data in Cloud Storage once and then have the training container read those files for every epoch. This removes repeated BigQuery scans, reduces cost and latency, and lets the training job use high-throughput file readers. The other choices either keep the expensive repeated reads or add complexity without fixing the core bottleneck.

Exam trap

The trap here is assuming that a newer BigQuery API or more workers will make repeated full-table reads cheap, when the real win is exporting once and reusing the data locally.

520
MCQeasy

An ML team wants to use Vertex AI Hyperparameter Tuning to tune a custom training job. They have a budget of 50 trials and want to use an algorithm that balances exploration and exploitation. Which algorithm should they choose?

A.Random search
B.Grid search
C.Bayesian optimization (Vizier default)
D.Manual search
AnswerC

Bayesian optimization (Vizier's default) models the objective probabilistically, using prior trial results to pick promising configurations while still sampling uncertain regions — directly balancing exploration and exploitation. It converges within the 50-trial budget far more efficiently than grid or random search, satisfying the stated constraint.

Why this answer

Bayesian optimization (the default algorithm in Vertex AI Vizier) is the correct choice because it explicitly balances exploration and exploitation by building a probabilistic model of the objective function and using an acquisition function to select the next hyperparameter configuration. With a budget of 50 trials, this algorithm efficiently converges to optimal regions while still exploring uncertain areas, making it ideal for tuning custom training jobs where each trial is computationally expensive.

Exam trap

A common pitfall is assuming that random search is the best default for balancing exploration and exploitation. However, random search lacks any exploitation mechanism, making Bayesian optimization the correct choice for efficient tuning within a constrained budget, as emphasized in Google PMLE.

How to eliminate wrong answers

Option A is wrong because random search does not balance exploration and exploitation; it samples hyperparameters uniformly at random without using past trial results to guide future selections, which wastes budget on suboptimal regions. Option B is wrong because grid search exhaustively evaluates a fixed set of hyperparameter combinations, which is computationally inefficient for a budget of 50 trials and does not incorporate any exploitation mechanism. Option D is wrong because manual search relies on human intuition and ad-hoc adjustments, which is not an automated algorithm and cannot systematically balance exploration and exploitation within a defined trial budget.

521
MCQeasy

You want to use a pre-trained model from TensorFlow Hub for image classification, but you need to adapt it to classify your own custom categories with a small dataset. Which Vertex AI approach is most appropriate?

A.Write a custom training script that loads the pre-trained model and fine-tunes it on your dataset
B.Deploy the pre-trained model as-is via Vertex AI JumpStart
C.Build a custom container with the pre-trained model and deploy to Vertex AI Endpoints
D.Use Vertex AI AutoML for image classification
AnswerA

Fine-tuning loads the TensorFlow Hub pre-trained model's weights into a custom training script and continues training on the small labelled dataset, adapting output categories. This suits limited data far better than training from scratch, which would require far more examples.

Why this answer

Fine-tuning a pre-trained model on your custom dataset is the most appropriate approach when you have a small dataset and need to adapt the model to new categories. This leverages transfer learning, where the pre-trained weights are used as a starting point and updated with your data.

Exam trap

The trap is confusing deployment with adaptation. Candidates may think that deploying a pre-trained model via JumpStart or a custom container will somehow adapt it to new categories, but adaptation requires training/fine-tuning.

How to eliminate wrong answers

Option B is wrong because deploying the pre-trained model as-is would not classify your custom categories; it would only recognize the original classes. Option C is wrong because building a custom container with the pre-trained model and deploying it does not adapt the model to your categories; it just serves the original model. Option D is wrong because AutoML for image classification builds a model from scratch (though it may use some transfer learning internally), but it does not allow you to start from a specific pre-trained model from TensorFlow Hub.

522
MCQmedium

You want to reduce training costs by using preemptible VMs on Vertex AI for a fault-tolerant distributed training job that uses checkpointing. Which machine type should you choose in the worker pool configuration?

A.Use spot VMs by setting 'spot' to true in the machine spec
B.Use custom machine types with preemptible flag
C.Use standard VMs and rely on Vertex AI auto-restart
D.Use TPU VMs because they are cheaper
AnswerA

Setting spot to true in the machine spec provisions spot (preemptible) VMs, which cost substantially less than standard VMs. Checkpointing makes the fault-tolerant job resilient to preemption, satisfying the requirement to reduce training costs while tolerating interruptions.

Why this answer

Vertex AI supports spot VMs (the successor to preemptible VMs) by setting the 'spot' field to true in the machine spec of the worker pool. Spot VMs offer up to 60-91% discounts and are suitable for fault-tolerant jobs with checkpointing, since they can be preempted with 30 seconds' notice. This is the documented, supported configuration for cost-optimized training.

Exam trap

The trap is using the legacy GCE term 'preemptible' or assuming auto-restart alone provides cost savings, when Vertex AI requires the explicit 'spot: true' field in the machine spec and checkpointing for fault tolerance.

How to eliminate wrong answers

Option B is wrong because 'preemptible' is the legacy GCE term; Vertex AI custom training uses the 'spot' boolean in the machine spec, and custom machine types are about vCPU/memory sizing, not preemption. Option C is wrong because standard VMs are on-demand and do not provide the cost savings of spot VMs; auto-restart does not change the billing model. Option D is wrong because TPU VMs are a different accelerator family and are not inherently cheaper than spot GPU/CPU VMs; choosing TPUs requires code compatibility with XLA and is not a drop-in cost optimization.

523
MCQhard

A company wants to use ML to predict customer churn. They have user activity logs in Cloud Storage, account data in BigQuery, and want an automated pipeline. Which pipeline architecture on Google Cloud should they use?

A.Load both data sources into AutoML Tables and train directly
B.Export logs from Cloud Storage to Cloud Dataproc for preprocessing, then train
C.Use Cloud Functions to preprocess data, then train on AI Platform
D.Use BigQuery to join logs and account data, train on Vertex AI, deploy to an endpoint
AnswerD

BigQuery joins the Cloud Storage activity logs with account data using external tables or load jobs, Vertex AI trains the churn model on that consolidated dataset, and endpoint deployment serves predictions. This satisfies the automated pipeline requirement while keeping data processing within managed Google Cloud services.

Why this answer

It leverages BigQuery's ability to join structured account data with semi-structured logs (via federated queries or external tables), then uses Vertex AI for end-to-end ML training and deployment. This architecture minimizes data movement, keeps the pipeline serverless, and directly addresses the requirement for an automated pipeline with both data sources.

Exam trap

Google Cloud often tests the misconception that AutoML Tables can handle multi-source data natively, when in fact it requires a single pre-joined dataset, and that Cloud Functions are suitable for heavy preprocessing workloads despite their strict resource limits.

How to eliminate wrong answers

Option A is wrong because AutoML Tables requires data to be in a single table format (CSV/JSON) and cannot directly ingest data from two separate sources without prior joining; it also lacks native pipeline automation. Option B is wrong because Cloud Dataproc (managed Spark/Hadoop) is overkill for simple preprocessing and introduces unnecessary cluster management overhead; BigQuery can perform the join and preprocessing more efficiently without spinning up ephemeral clusters. Option C is wrong because Cloud Functions have a 9-minute timeout and 2GB memory limit, making them unsuitable for preprocessing large-scale log data; Vertex AI is the correct training platform, but the preprocessing should be done in BigQuery, not Cloud Functions.

524
MCQmedium

A retail team must run nightly batch predictions over 20 TB of Parquet data stored in Cloud Storage using a custom PyTorch model registered in Vertex AI Model Registry. They want the job to finish within a fixed maintenance window and prefer not to manage the underlying compute. Which configuration should they use?

A.Schedule the model on a Vertex AI pipeline with a Dataflow step that calls the endpoint for each record, using streaming inserts to write results back to Cloud Storage.
B.Deploy the model to a Vertex AI endpoint with autoscaling and write a client script that submits all rows as individual online prediction requests.
C.Create a Vertex AI CustomJob that runs a PyTorch training script with the input path as a parameter, since CustomJob automatically performs batch inference when given a Parquet input.
D.Create a Vertex AI BatchPredictionJob with the model from Model Registry, set the input to the Cloud Storage Parquet path, and specify a machine type plus a starting and maximum replica count for the distributed workers.
AnswerD

Batch prediction reads directly from Cloud Storage, shards the input across workers, and scales horizontally according to the replica count you configure, all on managed infrastructure. Setting a maximum replica count lets the job finish within the maintenance window without the team provisioning or patching any compute themselves.

Why this answer

Batch prediction is purpose-built for large offline scoring: it reads files directly from Cloud Storage, distributes work across managed replicas, and writes sharded output without any endpoint. Specifying machine type and replica bounds gives the team control over throughput so the job completes inside the maintenance window while Vertex AI handles provisioning.

Exam trap

The trap here is assuming that a CustomJob or an online endpoint is needed for large-scale inference when the managed batch prediction service already covers it.

525
MCQhard

After setting up model monitoring on Vertex AI for a classification model, the engineer sees a high number of anomaly alerts for the "age" feature. Upon investigation, the age distribution in recent predictions is similar to training data. What might be the cause?

A.The feature importance of age has changed
B.The monitoring baseline was incorrectly set
C.The monitoring threshold for age is too low
D.The model is overfitting to age
AnswerC

Alerts fire when monitored values breach the configured threshold. Since the recent age distribution matches training data, no genuine drift exists; an excessively low threshold triggers alerts on normal statistical variation, so raising it eliminates the false positives.

Why this answer

The high number of anomaly alerts despite the age distribution being similar to training data indicates that the monitoring threshold for the 'age' feature is set too low. In Vertex AI Model Monitoring, anomaly detection compares recent prediction distributions against a baseline using statistical tests (e.g., the Kolmogorov-Smirnov test for numerical features). If the threshold is too sensitive, even minor, statistically insignificant deviations can trigger alerts, leading to false positives even when the distribution is essentially unchanged.

Exam trap

The trap here is that candidates confuse 'anomaly alerts' with 'model performance degradation' or 'data drift,' but the question specifically states the distribution is similar, so the root cause is a misconfigured sensitivity threshold, not a genuine distribution shift.

How to eliminate wrong answers

Option A is wrong because feature importance measures the contribution of a feature to model predictions, not the distribution of the feature values themselves; a change in feature importance would not directly cause distribution-based anomaly alerts. Option B is wrong because if the monitoring baseline were incorrectly set (e.g., using a non-representative sample), the age distribution in recent predictions would likely differ from the training data, but the question states the distribution is similar, so the baseline is not the issue. Option D is wrong because overfitting to age would manifest as poor generalization on unseen data, not as anomaly alerts on the feature distribution; overfitting does not inherently trigger monitoring alerts unless the distribution shifts.

Page 6

Page 7 of 11

Page 8

All pages