Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 151–225

775 questions total · 11pages · All types, answers revealed

Page 2

Page 3 of 11

Page 4
151
MCQhard

Your organization uses Vertex AI Feature Store to serve features for a real-time fraud detection model. Multiple teams contribute features, and you need to ensure that feature values are consistent between training and serving. Which practice should you implement to prevent training-serving skew?

A.Implement a separate feature engineering pipeline for training and another for serving to optimize each for its specific needs.
B.Export features from the featurestore to BigQuery for training, and use the featurestore online serving for predictions.
C.Use a single featurestore for both training and serving, and ingest features via the same pipeline.
D.Use Vertex AI Feature Store's built-in monitoring to detect skew, and retrain the model when skew is detected.
AnswerC

Using a single featurestore and a unified ingestion pipeline ensures that the same transformation logic and data sources are used for both training and serving, eliminating discrepancies. This practice directly addresses training-serving skew by maintaining consistency in feature computation and storage, which is essential for model reliability.

Why this answer

Training-serving skew arises when feature values differ between training and serving due to disparate data processing. Using a single featurestore and a unified ingestion pipeline ensures that features are computed and stored consistently, so the model sees identical feature distributions in both phases. This is a fundamental best practice in ML engineering.

Exam trap

The trap here is thinking that monitoring and retraining can solve training-serving skew, but they only detect and react to it; the only way to prevent it is to ensure identical feature computation for training and serving.

152
Multi-Selectmedium

A company wants to monitor features in Vertex AI Feature Store for drift over time. Which two services should they use? (Choose two.)

Select 2 answers
A.Vertex AI Feature Store monitoring
B.Cloud Logging
C.Vertex AI Model Monitoring
D.Vertex AI Experiments
E.Cloud Monitoring
AnswersA, E

Vertex AI Feature Store monitoring computes drift and skew metrics natively against feature values, detecting distribution shifts over time without custom code. It satisfies the requirement to monitor features for drift directly within the feature store, providing scheduled drift detection and alerting on the stored feature data.

Why this answer

Option A, Vertex AI Feature Store monitoring, is correct because it is the native capability that computes feature-level drift and skew statistics for features served from a Vertex AI Feature Store, generating metrics and alerts when distributions shift over time. Option E, Cloud Monitoring, is correct because the drift metrics produced by Feature Store monitoring are exported to Cloud Monitoring, where they can be visualized on dashboards and used to create alerting policies. Option B, Cloud Logging, is not the right choice because it captures log entries rather than computing or charting feature drift metrics.

Option C, Vertex AI Model Monitoring, targets model prediction drift/skew on deployed endpoints, not feature drift within a Feature Store. Option D, Vertex AI Experiments, is for tracking training runs, parameters, and metrics, not for ongoing drift monitoring of stored features.

Exam trap

PMLE often tests the distinction between Feature Store monitoring (feature-level drift) and Model Monitoring (prediction-level drift/skew), so candidates who conflate the two pick Vertex AI Model Monitoring.

153
MCQhard

A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?

A.Store the path in Data Catalog
B.Use Cloud Pub/Sub
C.Use PipelineParam to pass the output path
D.Write the output to a fixed Cloud Storage path and hardcode it in the pipeline
AnswerC

PipelineParam passes the Dataflow output path as a runtime parameter between pipeline components, satisfying the requirement to hand the preprocessed data location to the training step. Vertex AI Pipelines resolves this value at execution, decoupling the training component from hard-coded paths and enabling reuse across runs.

Why this answer

PipelineParam is the native mechanism in Vertex AI Pipelines (Kubeflow Pipelines SDK) to pass runtime outputs—such as a Cloud Storage path—between components. It creates a dependency graph that ensures the training step receives the exact output path from the preprocessing step, enabling dynamic, reproducible pipelines without hardcoding.

Exam trap

The trap here is that candidates confuse metadata services (Data Catalog) or messaging systems (Pub/Sub) with pipeline parameter passing, overlooking that Vertex AI Pipelines uses Kubeflow Pipelines' built-in component I/O for deterministic, graph-based data flow.

How to eliminate wrong answers

Option A is wrong because Data Catalog is a metadata management service for discovering and tagging assets, not designed to pass runtime pipeline parameters between steps; it would introduce unnecessary latency and coupling. Option B is wrong because Cloud Pub/Sub is an asynchronous messaging service for event-driven architectures, not a direct parameter-passing mechanism within a single pipeline execution; it would add complexity and potential ordering issues. Option D is wrong because hardcoding a fixed Cloud Storage path defeats pipeline reproducibility and scalability—if the preprocessing step changes its output location (e.g., due to timestamped folders), the training step would fail or use stale data.

154
MCQmedium

A data scientist deployed a TensorFlow model for sentiment analysis to Vertex AI Prediction. The model expects input key 'text' but the client sends requests with key 'review_text'. Which step should the data scientist take to resolve the error without retraining the model?

A.Use a Cloud Function to strip the 'review_text' key and replace it with 'text'
B.Retrain the model with input key 'review_text'
C.Create a new Vertex AI Endpoint with an alias mapping 'review_text' to 'text'
D.Modify the client code to send requests with input key 'text'
AnswerD

Vertex AI Prediction validates incoming instances against the model's signature, which expects the key 'text'. Renaming the client payload key from 'review_text' to 'text' aligns the request with that signature, resolving the error without retraining or altering the model.

Why this answer

The most straightforward and reliable solution is to modify the client code to send the request with the expected input key 'text'. This avoids any additional infrastructure, latency, or complexity, and does not require retraining the model or altering the deployed endpoint. Vertex AI Prediction serves the model as-is, so aligning the client's request format with the model's expected input is the simplest and most maintainable fix.

Exam trap

Google Cloud often tests the misconception that you need to add infrastructure (like Cloud Functions) or modify the model to handle input key mismatches, when the correct answer is to adjust the client code to match the model's expected input schema.

How to eliminate wrong answers

Option A is wrong because introducing a Cloud Function adds an unnecessary hop, increases latency, and creates an extra point of failure; it also violates the principle of keeping the architecture simple when a direct client-side fix exists. Option B is wrong because retraining the model is an expensive and time-consuming process that is not needed when the only issue is a key name mismatch in the request payload. Option C is wrong because Vertex AI Endpoints do not support alias mappings for input keys; the endpoint simply forwards the request payload to the model, and the model's input signature is fixed at deployment time.

155
MCQmedium

An ML team wants to share feature definitions across multiple projects to reduce training-serving skew and ensure consistency. They currently store features in Cloud Storage and manually coordinate updates, leading to errors. Which Google Cloud service should they use to centrally manage and serve features for both training and online inference?

A.Cloud Data Catalog
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Cloud Storage with versioning
AnswerC

Vertex AI Feature Store provides a central registry where feature definitions are authored once and served consistently to both training and online inference, removing manual Cloud Storage coordination errors. It also supplies point-in-time correctness, which prevents training-serving skew.

Why this answer

Vertex AI Feature Store is purpose-built to centrally define, store, and serve ML features for both training (batch) and online inference (low-latency) with a consistent feature definition, which directly eliminates training-serving skew. It provides a single source of truth so teams across projects can share features without manual coordination.

Exam trap

PMLE often tests whether candidates confuse metadata (Data Catalog), model (Model Registry), and feature (Feature Store) services — the trap is picking a storage or catalog service when the requirement is centralized feature serving with consistency.

How to eliminate wrong answers

Option A is wrong because Cloud Data Catalog is a metadata management service for discovery and governance, not a feature serving system — it cannot serve features for online inference. Option B is wrong because Vertex AI Model Registry stores and versions models, not features; it has no feature-serving capability. Option D is wrong because Cloud Storage with versioning is just object storage — it can hold feature data but provides no online serving, no point-in-time correctness, and no centralized feature definition, so it does not solve the coordination problem.

156
MCQhard

A company runs a Vertex AI endpoint that serves a model for real-time predictions. The endpoint uses a custom container that loads a 10 GB model into memory. During a traffic spike, the autoscaler adds new replicas, but each new replica takes 8 minutes to become ready because it must download the model from Cloud Storage. The team wants to reduce scale-up time. Which approach is most effective?

A.Increase the minReplicaCount to a value that covers peak traffic so that new replicas are not needed.
B.Bake the model artifacts into the custom container image and push it to Artifact Registry, then deploy the endpoint with that image.
C.Store the model artifacts in a regional Cloud Storage bucket and enable parallel downloads in the custom container.
D.Enable Vertex AI Model Monitoring to detect when the model is slow to load and trigger a preemptive scale-up.
AnswerB

Baking the model into the container image eliminates the runtime download from Cloud Storage. The container registry can serve the image layers efficiently, and the model is available as soon as the container starts. This significantly reduces startup time, often from minutes to seconds, and is the recommended pattern for large models. It also avoids dependency on external storage during scale-up.

Why this answer

The slow scale-up is caused by downloading a 10 GB model from Cloud Storage at container startup. By embedding the model into the container image and storing it in Artifact Registry, the model is available locally when the container starts. This removes the network transfer from the critical path and dramatically reduces the time for a new replica to become ready, allowing the autoscaler to respond faster to traffic spikes.

Exam trap

The trap here is focusing on autoscaler settings or monitoring instead of addressing the root cause: the model is fetched at runtime rather than being part of the image.

157
MCQmedium

Your team manages a production ML pipeline on Google Cloud that trains a fraud detection model every 6 hours using new transaction data. The pipeline steps are: (1) Cloud Function triggered by new files in Cloud Storage to validate data, (2) Dataflow job for feature engineering, (3) Vertex AI CustomJob for training, (4) Cloud Function to deploy the model to a Vertex AI endpoint after evaluation. You notice that the pipeline sometimes fails during the Dataflow job step with an error: 'Workflow failed. Causes: The job encountered a system error. Please try again later.' The error occurs sporadically, and retrying the pipeline manually usually succeeds. The team needs a reliable automated solution. What should you do?

A.Schedule the pipeline to run less frequently to reduce load on the Dataflow service.
B.Use Cloud Tasks to queue the Dataflow job and retry on failure.
C.Increase the number of Dataflow workers and use flexRS to handle transient errors.
D.Orchestrate the pipeline using Cloud Composer with retry policies on the Dataflow operator.
AnswerD

Cloud Composer's Dataflow operator supports configurable retry policies, so transient system errors during the Dataflow step are retried automatically without manual intervention. This directly satisfies the reliability constraint: sporadic failures that succeed on retry become self-healing, while dependency sequencing across all four pipeline steps is preserved.

Why this answer

Cloud Composer (Apache Airflow) provides native retry policies on its Dataflow operators, enabling automatic retries of the Dataflow job when it fails due to transient system errors. This addresses the sporadic failure pattern without manual intervention, ensuring the pipeline runs reliably every 6 hours.

Exam trap

The trap here is that candidates confuse scaling solutions (Option C) with fault-tolerance mechanisms, or they choose a generic queuing service (Option B) instead of a dedicated orchestrator with built-in retry policies for pipeline steps.

How to eliminate wrong answers

Option A is wrong because reducing pipeline frequency does not resolve transient system errors in Dataflow; it only delays processing and may cause data staleness. Option B is wrong because Cloud Tasks is a generic task queue that lacks native integration with Dataflow job lifecycle management and retry logic for pipeline-specific errors. Option C is wrong because increasing workers and using FlexRS improves resource availability but does not handle transient system errors that are unrelated to worker count or preemptibility; FlexRS is for cost savings on preemptible VMs, not for retry logic.

158
MCQmedium

A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?

A.Cloud Bigtable
B.Cloud Memorystore for Redis
C.Cloud Storage
D.Cloud Firestore
AnswerB

Memorystore for Redis provides sub-millisecond key-value lookups, letting the custom container return cached predictions for identical requests before invoking the model. This satisfies the latency and cost reduction goal, since repeated inference is skipped entirely.

Why this answer

Cloud Memorystore for Redis is an in-memory data store with sub-millisecond latency, making it the ideal GCP service for caching prediction results keyed by request hash. It supports TTL-based expiration and high-throughput reads, which directly reduce latency and repeated model inference costs. This is the canonical GCP caching layer for Vertex AI prediction workloads.

Exam trap

PMLE often tests the confusion between caching (in-memory, sub-millisecond, Redis/Memorystore) and persistent storage (Bigtable, Firestore, Cloud Storage) — candidates must match the latency and access-pattern requirements to the correct service class.

How to eliminate wrong answers

Option A (Cloud Bigtable) is wrong because it is a wide-column NoSQL store optimized for large-scale analytical and time-series workloads with millisecond latency — not sub-millisecond in-memory caching. Option C (Cloud Storage) is wrong because it is object storage with high latency (tens to hundreds of milliseconds) and is unsuitable for low-latency cache lookups. Option D (Cloud Firestore) is wrong because it is a document database with single-digit millisecond latency at best and is designed for mobile/web app data, not high-throughput prediction caching.

159
MCQmedium

An ML engineer is authoring a Vertex AI Pipelines component that runs a custom Python script. The component must accept a GCS path to training data and output a model artifact. The engineer wants the component interface to be strongly typed and to automatically generate the component specification from the Python function. Which approach should the engineer use?

A.Define the component using the @component decorator from the google.cloud.aiplatform.v1alpha1 package, specifying the input and output types as function annotations.
B.Use the google.cloud.aiplatform.CustomContainerTrainingJob class to package the script and run it as a pipeline step.
C.Write a Dockerfile that installs the required dependencies and exposes the script as an entrypoint, then build and push the image to Artifact Registry.
D.Define the component using the @dsl.component decorator from the kfp package, annotating the function parameters with types like str and Output[Model].
AnswerD

The kfp.dsl.component decorator (Kubeflow Pipelines SDK v2) generates a component specification from the Python function's type annotations. It supports InputPath, OutputPath, and artifact types like Model, ensuring strong typing. This is the standard method for authoring lightweight Python components in Vertex AI Pipelines.

Why this answer

The Kubeflow Pipelines SDK v2 @dsl.component decorator introspects Python type annotations to build a component specification, including input and output artifacts. This provides strong typing and automatic generation, which is exactly what the engineer needs. Other approaches either require manual YAML definition or are intended for different use cases like custom training jobs.

Exam trap

The trap here is assuming that any decorator or containerization automatically generates a typed component interface, when only the KFP v2 @dsl.component decorator does so from annotations.

160
MCQeasy

A data scientist has trained a TensorFlow model locally and wants to deploy it to Vertex AI for online predictions. The model accepts a single input tensor of shape (1, 224, 224, 3) and outputs a probability distribution over 10 classes. The data scientist wants to minimize deployment effort and ensure the model is served with low latency. What is the simplest way to deploy this model on Vertex AI?

A.Convert the model to ONNX format and use a custom container with ONNX Runtime for serving.
B.Use Vertex AI Batch Prediction with a pre-built TensorFlow container, and then set up a Cloud Function to serve online requests.
C.Export the model as a SavedModel and upload it to Vertex AI Model Registry, then deploy to an endpoint with a pre-built TensorFlow Serving container.
D.Package the model in a Docker container with a Flask app that loads the model and exposes a REST API, then deploy as a custom container on Vertex AI.
AnswerC

Vertex AI supports deploying TensorFlow SavedModels directly using a pre-built TensorFlow Serving container. This requires no custom code, and the serving container handles the model signature. It is the simplest and most efficient way to deploy a standard TensorFlow model for online predictions with low latency.

Why this answer

The simplest and most efficient way to deploy a standard TensorFlow model on Vertex AI for online predictions is to export it as a SavedModel and use the pre-built TensorFlow Serving container. This requires no custom serving code, leverages Vertex AI's managed infrastructure, and provides low-latency predictions. It also integrates with Vertex AI Model Registry for versioning and monitoring.

Exam trap

The trap here is overcomplicating the deployment by considering custom containers or format conversions when a pre-built container for TensorFlow is readily available.

161
Multi-Selecteasy

An ML team is converting a prototype model to a production pipeline using Vertex AI. They want to ensure model versioning and lineage. Which two practices should they adopt? (Select TWO)

Select 2 answers
A.Use Vertex AI Model Registry to manage model versions.
B.Only keep the latest model version to save storage.
C.Store model artifacts in Cloud Storage with unique versioned directories.
D.Train models directly in production without tracking.
E.Use a separate GCP project for each model version.
AnswersA, C

Vertex AI Model Registry provides centralised versioning and lineage tracking, directly satisfying the stem's requirement to manage model versions across the production pipeline. Each registered model version links to its training artefacts, enabling reproducible deployments and audit trails.

Why this answer

Option A is correct because Vertex AI Model Registry is the managed service specifically designed to track model versions, store metadata, and provide lineage information for models deployed in Vertex AI, which directly satisfies the requirement for model versioning and lineage. Option C is correct because storing model artifacts in Cloud Storage under unique versioned directories (e.g., gs://bucket/model/v1/, v2/) preserves immutable, traceable artifacts and supports lineage by linking each registered model version to its exact training output. Option B is incorrect because discarding older versions destroys the version history and lineage the team needs for reproducibility, rollback, and auditing.

Option D is incorrect because training directly in production without tracking eliminates any record of data, parameters, or artifacts, defeating versioning and lineage entirely. Option E is incorrect because creating a separate GCP project per model version is an unnecessary, costly isolation pattern; Vertex AI Model Registry and versioned Cloud Storage paths handle versioning within a single project.

Exam trap

PMLE often tests the difference between model versioning (Model Registry) and artifact storage (Cloud Storage), so candidates pick storage-only or project-per-version answers instead of the combined best practice.

162
MCQmedium

A team of data scientists and ML engineers is collaborating on a project using Vertex AI Workbench. They need to share notebooks and code, but want to avoid conflicts and maintain a history of changes. Which approach should they use?

A.Email notebook files to each other and manually merge changes.
B.Store notebooks in a shared Cloud Storage bucket and access them simultaneously.
C.Use Vertex AI Experiments to share notebook outputs.
D.Use a git repository (e.g., Cloud Source Repositories) to manage code and notebooks.
AnswerD

A git repository provides version control: commits preserve a full change history, and branching or merging resolves concurrent edits, preventing the conflicts that shared notebook storage causes. Cloud Source Repositories hosts this centrally, letting data scientists and ML engineers collaborate on Vertex AI Workbench notebooks safely.

Why this answer

Using a git repository (e.g., Cloud Source Repositories) provides version control, branching, and a full history of changes, which is essential for collaborative development. This approach avoids conflicts by allowing team members to work on separate branches and merge changes systematically, unlike shared storage or manual methods that lack conflict resolution and audit trails.

Exam trap

The trap here is that candidates confuse collaboration tools (like shared storage or experiment tracking) with version control, assuming that any shared access or logging mechanism can replace the structured history and conflict resolution of a git-based workflow.

How to eliminate wrong answers

Option A is wrong because emailing notebook files and manually merging changes is error-prone, lacks any version history or conflict detection, and does not scale for team collaboration. Option B is wrong because storing notebooks in a shared Cloud Storage bucket and accessing them simultaneously can lead to write conflicts, data corruption, and no built-in version history or merge capabilities. Option C is wrong because Vertex AI Experiments is designed for tracking and comparing model training runs and their metrics, not for managing source code or notebook version control.

163
MCQeasy

The exhibit shows a Vertex AI PipelineJob submission command. The pipeline fails because the component cannot find the input data. What is the most likely cause?

A.The pipeline root path is incorrect
B.The pipeline name is misspelled
C.The input data path is not accessible by the Vertex AI Pipelines service account
D.The region does not support the component
AnswerC

Vertex AI Pipelines components run under a service account whose IAM permissions and network access govern Cloud Storage reads. If that account lacks access to the input path, the component cannot fetch data, satisfying the inaccessible-input constraint described.

Why this answer

The most likely cause of the pipeline failing to find input data is that the Vertex AI Pipelines service account lacks the necessary permissions to access the specified input data path. Vertex AI Pipelines uses the Compute Engine default service account (or a custom service account) to read data from Cloud Storage or other sources; if this account does not have the `storage.objectViewer` role (or equivalent) on the bucket or object, the component will fail with a permission-denied error, even if the path is syntactically correct.

Exam trap

Google Cloud often tests the misconception that a misspelled pipeline name or incorrect pipeline root path is the cause of runtime data access failures, when in fact the service account's IAM permissions on the data source are the critical factor.

How to eliminate wrong answers

Option A is wrong because an incorrect pipeline root path would cause a failure to store pipeline artifacts or metadata, not a failure to find input data; the input data path is specified separately in the component's parameters. Option B is wrong because a misspelled pipeline name would cause the pipeline submission to fail at the API validation stage (e.g., an invalid name error), not during runtime when the component tries to access input data. Option D is wrong because the region not supporting the component would result in a resource or API availability error at submission time, not a runtime data access failure.

164
Multi-Selectmedium

You are configuring Vertex AI Model Monitoring for a deployed model on a Vertex AI Endpoint. The model uses a mix of numerical and categorical features. You want to ensure that the monitoring job effectively detects drift while minimizing false alerts. Which two actions should you take? (Choose two.)

Select 2 answers
A.Set the sampling rate to 1.0 to analyze every prediction request for drift.
B.Enable monitoring for all features, including those with low importance, to ensure comprehensive coverage.
C.Configure separate drift thresholds for each feature based on its historical variability and importance to the model.
D.Set the monitoring frequency to be aligned with the expected rate of change in the underlying data distribution.
E.Use the default drift threshold provided by Vertex AI Model Monitoring for all features to simplify configuration.
AnswersC, D

Using feature-specific thresholds accounts for differences in natural variability and importance. A feature like age might fluctuate more than a binary flag, so a single global threshold could cause false alerts for age or miss drift in the flag. Tailoring thresholds improves detection accuracy and reduces noise, making alerts more meaningful.

Why this answer

Aligning monitoring frequency with the data's rate of change and setting feature-specific thresholds based on variability and importance are key to effective drift detection. These actions reduce false alerts while ensuring critical drifts are caught, balancing sensitivity and operational efficiency.

Exam trap

The trap here is to use a one-size-fits-all approach with default thresholds or to monitor all features indiscriminately, which can lead to alert fatigue or missed critical drifts.

165
MCQmedium

Your team has deployed a tabular model to a Vertex AI Endpoint and enabled Vertex AI Model Monitoring with training-serving skew detection. You configured the monitoring job to run hourly and store statistics in a Cloud Storage bucket. After the first run, you notice that the job did not produce any drift metrics. You want to determine the cause of the missing metrics. What should you do first?

A.Increase the monitoring job's sampling rate to 1.0 to ensure that all prediction requests are analyzed.
B.Re-deploy the model to the endpoint with a larger machine type to handle the monitoring workload.
C.Verify that the Cloud Storage bucket has the correct IAM permissions for the Vertex AI service account to write objects.
D.Check that the endpoint's deployed model has a reference dataset specified and that the monitoring job's input schema matches the endpoint's prediction schema.
AnswerD

Monitoring requires a reference dataset (the baseline) and a matching schema to compute skew. If the reference dataset is missing or the schema does not align with the endpoint's inputs, the job cannot compute metrics and may silently produce no output. Verifying these configurations is the correct first step to diagnose the issue.

Why this answer

Vertex AI Model Monitoring requires a reference dataset and a matching schema to compute training-serving skew. Without these, the job cannot generate metrics. Checking the configuration of the reference dataset and schema alignment is the logical first step to resolve missing metrics.

Exam trap

The trap here is assuming that missing metrics are caused by insufficient sampling or permissions, while the most common cause is a missing or misconfigured reference dataset.

166
Multi-Selecthard

A media company wants a low-code pipeline that ingests uploaded video files, detects scenes and on-screen text, and stores structured metadata for search. They prefer managed services and minimal custom code. Which TWO Google Cloud capabilities should they combine? (Choose two.)

Select 2 answers
A.Cloud Vision API for detecting labels and text in sampled video frames
B.Vertex AI Vision for building a custom object-tracking pipeline with trained detectors
C.Cloud Data Fusion for building a visual ETL pipeline for video ingestion
D.Cloud Storage and Pub/Sub to trigger processing when new videos arrive
E.Video Intelligence API for shot change detection and text detection
AnswersD, E

A Cloud Storage bucket receiving uploads can publish object notifications to Pub/Sub, which then triggers the video analysis step. This serverless eventing is the standard low-code way to automate ingestion so each new video is processed and its metadata stored without manual intervention.

Why this answer

The Video Intelligence API supplies managed shot change and text detection on video, producing the scene and on-screen text metadata needed for search. Pairing it with Cloud Storage object notifications to Pub/Sub creates an event-driven trigger so each upload is analyzed automatically. Together they deliver a managed, low-code pipeline without custom frame extraction or model training.

Exam trap

The trap here is substituting frame-by-frame Cloud Vision API calls for the video-native Video Intelligence API, which ignores the extra code required to sample frames and align timestamps.

167
MCQmedium

A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?

A.Switch the endpoint to a private endpoint and increase the machine type of each replica.
B.Enable request-response logging and reduce the model's input feature count.
C.Configure a dedicated autoscaling metric with a lower utilization target and define a scale-up schedule aligned to the evening peak.
D.Set a higher maxReplicaCount and rely on the default autoscaling metrics.
AnswerC

Vertex AI Model Deployment autoscaling supports both a target utilization metric and scheduling options. Lowering the target utilization makes the autoscaler add replicas earlier, while a schedule pre-warms capacity exactly when the predictable evening surge begins. Together they cut the reactive ramp-up delay without paying for peak capacity around the clock, which is precisely the SLO gap described.

Why this answer

The latency breach is a capacity-timing problem, not a capacity-ceiling problem, so the fix must make replicas available before demand arrives. Vertex AI Model Deployment autoscaling lets you choose the metric the autoscaler tracks and set a target utilization, and it also supports scheduled scaling. Lowering the target makes scale-out begin earlier, and a schedule aligned with the predictable evening peak pre-warms replicas, eliminating the reactive ramp-up lag while keeping off-peak cost low.

Exam trap

The trap here is assuming that a larger maximum replica count automatically reduces scale-up latency, when the delay actually comes from reactive scaling that only starts after load is already observed.

168
MCQhard

A team uses Vertex AI Feature Store with an online store for low-latency serving. They need to support frequent updates to features (e.g., every minute) and require high write throughput (thousands of writes per second). Which online store type should they choose?

A.Optimized online store
B.Firestore online store
C.Bigtable online store
D.Cloud SQL online store
AnswerC

Bigtable online store supports high write throughput and frequent feature updates, scaling to thousands of writes per second with low-latency reads. This satisfies the stated requirement for minute-level updates and high-throughput ingestion that the default online store cannot match.

Why this answer

Bigtable online store is designed for high-throughput, low-latency serving with frequent updates, making it suitable for thousands of writes per second and minute-level feature refreshes. It leverages Bigtable's scalable NoSQL architecture, which handles high write loads efficiently. Optimized online store is for low-latency but may not sustain such high write throughput, while Firestore and Cloud SQL have lower write limits.

Exam trap

The trap is assuming that any low-latency online store can handle high write throughput; candidates must distinguish between read-optimized and write-optimized stores, with Bigtable being the only one designed for massive write scalability.

How to eliminate wrong answers

Option A is wrong because the optimized online store, while low-latency, is not optimized for high write throughput and may throttle under thousands of writes per second. Option B is wrong because Firestore online store is designed for smaller-scale, lower-throughput applications and has write limits that cannot handle thousands of writes per second. Option D is wrong because Cloud SQL is a relational database with lower write scalability and is not intended for high-throughput feature serving.

169
MCQeasy

An ML team wants to monitor feature drift in their production model. Which Vertex AI Feature Store capability should they use?

A.Feature views
B.Online store
C.Point-in-time retrieval
D.Feature monitoring (drift detection)
AnswerD

Feature monitoring in Vertex AI Feature Store continuously computes drift metrics by comparing production feature distributions against a baseline snapshot, directly satisfying the requirement to monitor feature drift. It detects skew and drift per feature, emitting alerts without retraining, which is precisely the capability the stem requests.

Why this answer

Feature monitoring (drift detection) is the Vertex AI Feature Store capability that computes drift and skew metrics for feature values against a baseline and emits them for alerting. It is purpose-built to detect when production feature distributions diverge from training or reference distributions. This directly answers the requirement to monitor feature drift.

Exam trap

PMLE often tests the confusion between feature views (serving constructs) and feature monitoring (drift detection), so candidates pick feature views when the question asks specifically about drift.

How to eliminate wrong answers

Option A is wrong because feature views are the logical groupings that serve features for online/offline use, not the drift-detection mechanism itself. Option B is wrong because the online store is the low-latency serving layer for real-time feature retrieval, not a monitoring component. Option C is wrong because point-in-time retrieval ensures training data uses historically correct feature values to avoid leakage, which is a correctness feature, not drift monitoring.

170
MCQmedium

Refer to the exhibit. A team configured Vertex AI Model Monitoring with skew detection for feature "income" with a threshold of 0.2. However, they have not received any alerts even though they suspect data drift. What is the most likely reason?

A.The monitoring is not enabled for the endpoint
B.The 'income' feature is not present in the serving data
C.The actual skew is below the threshold
D.The drift detection threshold is set higher
AnswerB

Skew detection compares training and serving feature distributions, so it only evaluates features present in both. If 'income' is absent from serving data, no comparison occurs and no alert fires regardless of the 0.2 threshold, explaining the silence despite suspected drift.

Why this answer

Vertex AI Model Monitoring skew detection compares the training data distribution against the serving (prediction) data distribution for each monitored feature. If the 'income' feature is missing from the serving requests — for example, the endpoint's input schema does not include it, or it is named differently — the monitor has no serving-side data to compare and cannot compute skew, so no alert fires even if drift is suspected. The feature must be present in the serving payload for skew to be evaluated.

Exam trap

PMLE often tests the assumption that configuring a monitor guarantees alerts — candidates overlook that skew detection requires the feature to be present in serving data, so a missing or renamed feature silently disables detection for that feature.

How to eliminate wrong answers

Option A is wrong because if monitoring were not enabled for the endpoint, the team would see no monitoring metrics or configuration at all, and the question states monitoring was configured with a threshold. Option C is wrong because the team suspects drift and the question asks for the most likely reason no alerts fired — if skew were genuinely below threshold, that would be expected behavior, but the scenario implies a configuration issue. Option D is wrong because the threshold is stated as 0.2 and the question does not indicate it was changed; raising the threshold would only make alerts less likely, but the root cause here is missing serving data, not threshold tuning.

171
MCQeasy

An ML team uses Vertex AI Pipelines and wants to automatically generate model cards documenting model purpose, evaluation results, and intended use. Which approach should they take?

A.Manually create a Google Doc and share it with the team.
B.Use Cloud Data Catalog to annotate the model artifact.
C.Write a custom Kubeflow Pipelines component that creates a BigQuery table with model metadata.
D.Use Vertex AI Model Registry to generate model cards automatically from model metadata.
AnswerD

Vertex AI Model Registry generates model cards from registered model metadata, including purpose, evaluation results and intended use, when the model is uploaded with those details. This automates documentation from the pipeline's existing metadata rather than requiring manual authoring.

Why this answer

Vertex AI Model Registry automatically generates model cards from model metadata, including purpose, evaluation results, and intended use, when models are registered and their metadata is populated. This provides a native, integrated way to document models without manual effort or custom components. It aligns with Vertex AI's governance and documentation features.

Exam trap

PMLE often tests the distinction between metadata storage and model card generation — candidates pick Cloud Data Catalog or custom BigQuery tables because they store metadata, but only Vertex AI Model Registry generates the structured model card.

How to eliminate wrong answers

Option A is wrong because manually creating a Google Doc is not automated, not integrated with Vertex AI Pipelines, and does not generate model cards from model metadata. Option B is wrong because Cloud Data Catalog is a metadata management service for discovery and governance, not a model card generator; annotating a model artifact does not produce a model card. Option C is wrong because writing a custom Kubeflow component to create a BigQuery table stores metadata but does not generate the structured model card format that Vertex AI Model Registry provides.

172
MCQhard

A team trained a TensorFlow model locally and wants to deploy it to BigQuery ML for predictions without retraining. They have exported the SavedModel to Cloud Storage. Which statement is correct?

A.They need to convert the model to a BigQuery ML native format first.
B.They can create a model using CREATE MODEL with model_type='tensorflow' and the path to the SavedModel.
C.They must first retrain the model using ML.TRAIN on BigQuery.
D.They can use ML.PREDICT directly on the SavedModel in Cloud Storage.
AnswerB

CREATE MODEL with model_type='tensorflow' imports an existing SavedModel directly from Cloud Storage, so BigQuery ML serves predictions without retraining. This satisfies the stem's constraint: the locally trained TensorFlow artefact is deployed as-is, with no data or training step required in BigQuery.

Why this answer

BigQuery ML supports importing TensorFlow models in SavedModel format directly using CREATE MODEL with model_type='tensorflow' and the path to the SavedModel in Cloud Storage. This allows you to use the model for prediction in BigQuery without retraining or converting to a native format. The model is imported as a BigQuery ML model and can be used with ML.PREDICT.

Exam trap

The trap is assuming that BigQuery ML requires models to be trained within BigQuery or converted to a proprietary format. In fact, it supports direct import of TensorFlow SavedModels.

How to eliminate wrong answers

Option A is wrong because no conversion to a native BigQuery ML format is needed; TensorFlow SavedModels are supported directly. Option C is wrong because retraining is not required; the whole point is to use the pre-trained model. Option D is wrong because ML.PREDICT cannot be used directly on a SavedModel in Cloud Storage; you must first create a BigQuery ML model that references the SavedModel.

173
Multi-Selectmedium

A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)

Select 2 answers
A.Store the index in Cloud Storage and query via Python
B.Enable streaming updates for the index
C.Use a brute-force index for exact results
D.Deploy the index to a Vertex AI Matching Engine endpoint
E.Use batch updates only
AnswersB, D

Streaming updates allow new embeddings to be inserted or removed incrementally while the index remains queryable, avoiding the full rebuild and redeploy that batch updates force. This satisfies the frequent-update requirement without interrupting the low-latency online serving the stem specifies.

Why this answer

Option B is correct because Vertex AI Matching Engine supports streaming updates, which allow the index to be modified incrementally as new data arrives without requiring a full rebuild, directly satisfying the requirement for frequent updates in a real-time recommendation system. Option D is correct because deploying the index to a Vertex AI Matching Engine endpoint is the standard mechanism for serving low-latency, real-time nearest-neighbor queries at scale, which is exactly what the scenario demands. Option A is not appropriate because storing the index in Cloud Storage and querying it via Python does not provide the managed, low-latency serving infrastructure of Matching Engine.

Option C is not appropriate because a brute-force index performs exhaustive comparisons, which is far too slow for low-latency real-time serving at scale. Option E is not appropriate because batch-only updates cannot keep pace with frequent updates and would introduce staleness in a real-time recommendation system.

Exam trap

The trap here is that in Google's Vertex AI Matching Engine, candidates often confuse batch updates with streaming updates, assuming that batch updates can be made frequent enough to approximate real-time, but they fail to recognize that batch updates require full index rebuilds, which introduce significant latency and downtime for serving.

174
MCQmedium

A data science team deploys a custom container on Vertex AI Prediction for a PyTorch model. After deployment, the model returns predictions that are consistently off by a constant factor. The model performed correctly during local testing. What is the most likely cause?

A.The model is loaded in evaluation mode, but the training mode was used in testing.
B.The serving input function in the container is not applying the same normalization as during training.
C.The container is using a different PyTorch version than the training environment.
D.There is a bug in the custom container's prediction route.
AnswerB

The container's serving input function must replicate training-time preprocessing exactly. If it omits the same normalisation, inputs reach the PyTorch model on a different scale, producing predictions offset by a constant factor — matching the symptom while local testing, which used preprocessed data, passed.

Why this answer

A consistent constant-factor error in predictions, with correct local behavior, points to a preprocessing mismatch: the serving input function in the container is not applying the same normalization (e.g., scaling or mean-centering) that was applied during training. Without identical normalization, inputs are systematically shifted, producing outputs off by a constant factor.

Exam trap

The trap is assuming a constant-factor error indicates a code bug or version mismatch, when it almost always signals a preprocessing/normalization mismatch between training and serving.

How to eliminate wrong answers

Option A is wrong because evaluation vs. training mode affects dropout and batch normalization behavior, which would cause erratic or degraded predictions, not a consistent constant-factor offset. Option C is wrong because a PyTorch version mismatch typically causes errors, warnings, or subtle numerical differences — not a uniform constant-factor error. Option D is wrong because a bug in the prediction route would more likely cause crashes, malformed responses, or variable errors, not a systematic constant scaling.

175
MCQhard

A data scientist wants to perform A/B testing between two model versions deployed on the same Vertex AI endpoint. They need to route 10% of traffic to the challenger model. Which approach should they use?

A.Use Vertex AI Experiments to compare models offline, then deploy the winner
B.Deploy the challenger model to a separate endpoint and use a load balancer to split traffic
C.Update the champion model with a new version and use model version aliases
D.Deploy both models to the same endpoint and set traffic_split to 90 for champion and 10 for challenger
AnswerD

A single Vertex AI endpoint can host multiple model versions as DeployModel resources, and the endpoint's traffic_split dictionary assigns integer percentage weights per deployed model ID. Setting 90/10 routes exactly one tenth of prediction requests to the challenger, giving the required A/B comparison without a second endpoint.

Why this answer

Vertex AI endpoints support traffic splitting directly, allowing you to route a percentage of requests to different model versions deployed on the same endpoint. By setting `traffic_split` to 90 for the champion and 10 for the challenger, the data scientist can perform online A/B testing without additional infrastructure. This is the simplest and most cost-effective approach, as it avoids managing separate endpoints or load balancers.

Exam trap

A common trap in Google exams is the misconception that separate endpoints or load balancers are required for A/B testing, when in fact Vertex AI endpoints provide built-in traffic splitting for this exact purpose.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is an offline evaluation tool for comparing model performance on historical data, not for live traffic splitting. Option B is wrong because deploying the challenger to a separate endpoint and using a load balancer adds unnecessary complexity and cost; Vertex AI endpoints natively support traffic splitting across model versions. Option C is wrong because updating the champion model with a new version and using model version aliases does not provide granular traffic splitting; aliases are for version management, not for routing a specific percentage of live traffic.

176
Multi-Selectmedium

A company wants to set up end-to-end monitoring for a Vertex AI model. Which three components should they include?

Select 3 answers
A.Feature store backup status
B.Model performance metrics
C.Data drift and concept drift detection
D.Prediction latency
E.Model training cost
AnswersB, C, D

Model performance metrics provide continuous visibility into prediction quality, drift and skew once the Vertex AI model is deployed, satisfying the end-to-end monitoring requirement. Vertex AI Model Monitoring tracks feature and prediction drift against training baselines, alerting when live traffic diverges, so the company detects degradation rather than only infrastructure health.

Why this answer

For end-to-end monitoring of a Vertex AI model, the company needs to track how well the deployed model is performing, so B (Model performance metrics) is correct because Vertex AI Model Monitoring surfaces metrics like accuracy, precision, and recall against ground-truth or skewed prediction distributions. C (Data drift and concept drift detection) is correct because Vertex AI Model Monitoring specifically detects training-serving skew and drift in feature distributions (data drift) as well as changes in the relationship between features and labels (concept drift), which are core to maintaining model quality over time. D (Prediction latency) is correct because operational monitoring of an endpoint must include request/response latency to ensure the model meets serving SLOs, and Vertex AI exposes these metrics via Cloud Monitoring.

The remaining options do not belong: A (Feature store backup status) concerns data durability of a feature store, not model monitoring, and E (Model training cost) is a billing/training concern rather than an end-to-end runtime monitoring component.

Exam trap

The trap here is that candidates often confuse operational or cost-related metrics (like backup status or training cost) with the three core pillars of model monitoring: performance metrics, drift detection, and latency tracking.

177
MCQhard

An organization runs a batch prediction job on Vertex AI for a large dataset (10 TB). The job is configured to use a cluster of 100 n1-standard-16 machines. Midway through, the job fails with 'Out of memory' errors. What is the most effective mitigation strategy?

A.Split the input data into smaller chunks and run multiple jobs.
B.Enable model parallelism within the prediction script.
C.Increase the number of machines to distribute data more.
D.Use a machine type with more memory per instance.
AnswerD

The 'Out of memory' failure indicates each worker exhausts RAM while processing its shard, not that the cluster is too small. Moving to a machine type with more memory per instance directly addresses the per-instance memory ceiling.

Why this answer

The 'Out of memory' error indicates that individual worker nodes are running out of RAM when processing their assigned data shards. Using a machine type with more memory per instance (e.g., n1-highmem-16) directly addresses the root cause by providing each node with sufficient memory to hold the model and its intermediate computations, without changing the data distribution or parallelism strategy.

Exam trap

The trap here is that candidates confuse scaling horizontally (adding more machines) with scaling vertically (increasing per-machine resources), assuming that distributing data further will fix memory exhaustion when the bottleneck is per-node RAM capacity, not data volume per node.

How to eliminate wrong answers

Option A is wrong because splitting the input data into smaller chunks and running multiple jobs does not increase the memory available per machine; it only reduces the data per job, but the same memory constraint per node will still cause OOM errors if the model or batch size per node remains unchanged. Option B is wrong because model parallelism splits the model across devices, which is typically used for very large models that cannot fit on a single GPU/TPU, not for batch prediction jobs where the model is already loaded and the issue is data processing memory. Option C is wrong because increasing the number of machines distributes the data across more nodes, but each node still has the same 16 GB of RAM (n1-standard-16), so the per-node memory pressure remains identical and OOM errors will persist.

178
Multi-Selecthard

Your team is deploying a large model on edge devices and needs to reduce its size by 80% while maintaining reasonable accuracy. Which THREE techniques should they consider? (Choose 3.)

Select 3 answers
A.Quantisation to INT8
B.Transfer learning from a larger model
C.Knowledge distillation
D.Increasing model capacity with more layers
E.Pruning of redundant connections
AnswersA, C, E

Reduces model size by reducing precision of weights.

Why this answer

Quantisation to INT8 reduces the precision of model weights and activations from 32-bit floating point to 8-bit integers, cutting memory usage by approximately 75% (4x compression). This directly addresses the 80% size reduction target while often preserving accuracy within 1-2% through careful calibration and scaling, making it a primary technique for edge deployment.

Exam trap

Google Cloud often tests the misconception that transfer learning reduces model size, when in fact it only transfers learned features and does not compress the model; candidates may confuse it with knowledge distillation.

179
MCQmedium

A data scientist wants to train a PyTorch model on Vertex AI using a pre-built container for GPU training. She needs to use 4 NVIDIA A100 GPUs on a single machine. Which machine configuration should she select?

A.n1-highmem-16 with 4 NVIDIA V100 GPUs
B.n1-standard-16 with 4 NVIDIA T4 GPUs
C.a2-highgpu-4g (4 A100 GPUs)
D.a2-megagpu-16g (16 A100 GPUs)
AnswerC

The a2-highgpu-4g machine type provides exactly four NVIDIA A100 GPUs on a single node, matching the stem's requirement for four A100s on one machine. This satisfies the GPU count and single-machine constraint, letting the pre-built PyTorch container train without custom configuration or distributed multi-node setup.

Why this answer

The a2-highgpu-4g machine type is purpose-built for GPU workloads and provides exactly 4 NVIDIA A100 GPUs on a single machine, matching the requirement precisely. Vertex AI supports this machine type for custom training with pre-built containers, so the data scientist can select it directly in the training configuration. Choosing it avoids over-provisioning or mismatching GPU counts.

Exam trap

PMLE often tests machine-type specificity — candidates pick a generic n1 machine with the right GPU count, missing that A100 GPUs require the A2 machine family, not N1.

How to eliminate wrong answers

Option A is wrong because n1-highmem-16 with 4 V100 GPUs uses older V100 GPUs, not the requested A100s, and n1 machines are general-purpose, not GPU-optimized. Option B is wrong because n1-standard-16 with 4 T4 GPUs provides T4 GPUs, which are inference-oriented and far less powerful than A100s for training. Option D is wrong because a2-megagpu-16g provides 16 A100 GPUs, four times the requested count, wasting cost and exceeding the stated requirement.

180
MCQeasy

Which Vertex AI service is designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale?

A.Vertex AI AutoML
B.Vertex AI Workbench
C.Vertex AI Prediction
D.Vertex AI Matching Engine (Vector Search)
AnswerD

Matching Engine (Vector Search) is the Vertex AI service purpose-built for creating and serving ANN indexes, delivering scalable similarity search across billions of embeddings. It directly satisfies the stem's requirement for approximate nearest neighbour indexing at scale, unlike general prediction or training services.

Why this answer

Vertex AI Matching Engine (Vector Search) is specifically designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale. It allows you to create indexes from embeddings and perform low-latency similarity queries. This service is optimized for large-scale vector search use cases.

Exam trap

PMLE often tests the distinction between Vertex AI services, and candidates may confuse Matching Engine with Prediction or AutoML, not realizing it's specifically for vector similarity search.

How to eliminate wrong answers

Option A is wrong because Vertex AI AutoML is for building custom machine learning models without coding, not for ANN indexes. Option B is wrong because Vertex AI Workbench is a notebook environment for data science, not a similarity search service. Option C is wrong because Vertex AI Prediction is for deploying models for online or batch predictions, not for managing ANN indexes.

181
MCQmedium

Refer to the exhibit. A machine learning engineer deployed a model on Vertex AI using this configuration. When testing the endpoint, the engineer receives a 400 error with the message: 'Invalid argument: Explanation metadata missing required field: `outputs`.' What is the most likely cause?

A.The explanation metadata outputs field is missing the required 'displayName' attribute.
B.The explanation metadata needs a 'baseline' configuration for the input.
C.The explanation metadata inputs field should be wrapped inside a 'visualization' block.
D.The explainability method chosen is not supported for the model type.
AnswerA

Incorrect. The error is about missing outputs field, not just missing displayName.

Why this answer

The error message 'Invalid argument: Explanation metadata missing required field: `outputs`' indicates that the `outputs` field is entirely absent from the explanation metadata configuration. None of the provided options correctly identifies this issue. Option A incorrectly attributes the error to a missing 'displayName' attribute within an existing outputs field, which is not the cause.

Exam trap

Candidates often focus on the subfields like displayName, but the error indicates the top-level outputs field is missing.

How to eliminate wrong answers

Option B is wrong because a `baseline` configuration is required for the input, not the output; the error specifically points to the missing `outputs` field, not the input baseline. Option C is wrong because the `visualization` block is used for image-specific explanations (e.g., integrated gradients with visualization), not for wrapping the inputs field; the error is about the `outputs` field, not the inputs. Option D is wrong because the error message does not mention an unsupported explainability method; it explicitly states that the `outputs` field is missing, which is a metadata configuration issue, not a method compatibility problem.

182
MCQeasy

A data analyst wants to create a classification model directly in BigQuery using SQL. Which feature should they use?

A.BigQuery ML
B.Vertex AI
C.Dataflow
D.Cloud ML Engine
AnswerA

BigQuery ML lets analysts build and train models using SQL directly inside BigQuery, satisfying the requirement to create a classification model without leaving SQL or exporting data. It supports models such as logistic regression and boosted trees via CREATE MODEL statements, so no separate machine-learning tooling or Python is needed.

Why this answer

BigQuery ML (BQML) enables users to create and execute machine learning models directly in BigQuery using standard SQL syntax, without needing to export data or manage separate ML infrastructure. For a data analyst who wants to build a classification model entirely within BigQuery, BQML provides the CREATE MODEL statement with classification algorithms like logistic regression or XGBoost, making it the correct and most direct feature.

Exam trap

Google Cloud often tests the distinction between services that run inside BigQuery (BQML) versus external ML platforms (Vertex AI), trapping candidates who think any ML service qualifies without checking if it operates directly via SQL in BigQuery.

How to eliminate wrong answers

Option B is wrong because Vertex AI is a full MLOps platform for training, deploying, and managing models, but it requires data to be exported from BigQuery and does not allow model creation directly in SQL within BigQuery. Option C is wrong because Dataflow is a stream and batch data processing service (based on Apache Beam) used for ETL and data pipelines, not for creating classification models. Option D is wrong because Cloud ML Engine (now part of Vertex AI) is a managed service for training and serving custom ML models, but it does not support SQL-based model creation inside BigQuery.

183
MCQhard

A data science team deploys a large language model (LLM) on Vertex AI Prediction using an NVIDIA A100 GPU. The end-to-end latency is acceptable, but the cost is high due to low GPU utilization. The model is stateless and requests are independent. Which strategy would most effectively reduce cost per prediction?

A.Migrate the model to Cloud TPU using TensorFlow to benefit from higher throughput.
B.Use a smaller GPU, such as NVIDIA T4, and increase the number of replicas to maintain throughput.
C.Reduce the number of min replicas to 0 and scale from 0 on each request.
D.Implement dynamic batching in the serving container to aggregate multiple requests into a single inference call.
AnswerD

Dynamic batching groups concurrent independent requests into one GPU inference call, raising utilisation per A100 second. Since the model is stateless and requests are independent, batching introduces no correctness risk while directly cutting cost per prediction.

Why this answer

Dynamic batching aggregates multiple independent inference requests into a single forward pass on the GPU, dramatically increasing GPU utilization and throughput per GPU-second. Since the model is stateless and requests are independent, batching is safe and does not affect correctness. Higher utilization directly translates to lower cost per prediction because the same A100 does far more work per unit time.

Exam trap

The trap here is assuming that hardware changes (smaller GPU, TPU, autoscaling to zero) solve a utilization problem, when the actual fix is software-level request aggregation; PMLE often tests whether you identify batching as the canonical cost-optimization lever for stateless inference.

How to eliminate wrong answers

Option A is wrong because migrating to Cloud TPU requires re-engineering the model in TensorFlow/JAX and does not address the root cause of low GPU utilization; TPUs also don't automatically fix batching inefficiency. Option B is wrong because using a smaller GPU (T4) with more replicas increases the number of underutilized GPUs rather than improving per-GPU efficiency, and T4 has less memory/compute for large LLMs. Option C is wrong because scaling min replicas to 0 introduces cold-start latency and does not improve utilization during active serving; it only saves cost when idle, which is not the stated problem.

184
Multi-Selecthard

An ML team is using Vertex AI Pipelines to orchestrate a training workflow. They need to pass a large dataset (500 GB) between two components. The first component preprocesses the data and writes the output to Cloud Storage. The second component trains a model using that preprocessed data. The team wants to minimize pipeline execution time and cost. Which two strategies should they use? (Choose two.)

Select 2 answers
A.Mount a Cloud Storage bucket as a Persistent Disk to both components so they can share the data via a common file system.
B.Serialize the preprocessed data into a single TFRecord file and pass it as a pipeline parameter to the second component.
C.Have the first component write the preprocessed data to a Cloud Storage location and output the URI as a string parameter to the second component.
D.Use a Dataset artifact to represent the preprocessed data and pass it as an output of the first component and input to the second.
E.Use a Model artifact to pass the preprocessed data to the training component, and have the training component load the data from the artifact's URI.
AnswersC, D

Passing the Cloud Storage URI as a string parameter is lightweight and allows the second component to read the data directly from Cloud Storage. This avoids data duplication and keeps the pipeline spec small. It is a common pattern when a dedicated artifact type is not necessary, though using a Dataset artifact is more semantically rich.

Why this answer

For large data, the pipeline should pass references (URIs) rather than the data itself. Using a Dataset artifact or a string parameter to convey the Cloud Storage location lets each component access the data independently, minimizing data movement and pipeline overhead. Serializing or sharing via disks is not feasible at this scale and would increase cost and time.

Exam trap

The trap here is thinking that large data can be passed directly between components as parameters or that shared storage can be mounted, when in fact only lightweight references should be passed.

185
Multi-Selecthard

A team wants to implement CI/CD for their ML pipeline using Cloud Build. They want to automatically compile and deploy the pipeline when code is pushed to the main branch. Which three steps should they include in the Cloud Build configuration? (Choose three.)

Select 3 answers
A.Create or update the pipeline in Vertex AI using the compiled file
B.Upload the compiled pipeline to Cloud Storage
C.Run the pipeline immediately after deployment
D.Install KFP SDK and compile the pipeline
E.Configure Cloud Scheduler to trigger on push
AnswersA, B, D

Creating or updating the pipeline in Vertex AI registers the compiled pipeline definition, making it runnable and schedulable. This satisfies the deployment half of the CI/CD requirement, since compilation alone leaves the pipeline unregistered and unable to execute on Vertex AI.

Why this answer

Option D is correct because the Cloud Build configuration must first set up the environment by installing the Kubeflow Pipelines (KFP) SDK and then compile the pipeline definition into a compiled YAML/JSON artifact, which is the essential build step for a CI/CD ML pipeline. Option B is correct because the compiled pipeline artifact needs to be uploaded to a Cloud Storage bucket so it can be referenced and used by Vertex AI when creating or updating the pipeline. Option A is correct because the final deployment step is to create or update the pipeline in Vertex AI using the compiled file, which is the actual CD action that makes the pipeline available in Vertex AI Pipelines.

Option C is not correct because running the pipeline immediately after deployment is not a required CI/CD configuration step; deployment and execution are separate concerns, and the scenario only asks to compile and deploy on push. Option E is not correct because Cloud Scheduler is used for time-based triggering, whereas the scenario requires triggering on code push to the main branch, which is handled by Cloud Build triggers, not Cloud Scheduler.

Exam trap

A common trap is confusing build-time actions (compilation, upload, registration) with runtime actions (execution, scheduling). Candidates often mistakenly include immediate pipeline execution as a CI/CD step instead of focusing on deploying and registering the pipeline artifact.

186
Multi-Selecthard

An engineer is designing a distributed training job on Vertex AI for a TensorFlow model that uses the MultiWorkerMirroredStrategy. They need to ensure proper communication between workers. Which environment variable must be set correctly for each worker?

Select 1 answer
A.CLUSTER_SPEC
B.TF_CPP_MIN_LOG_LEVEL
C.TF_CONFIG_JSON
D.TF_DISTRIBUTED_STRATEGY
E.TF_CONFIG
AnswersE

TF_CONFIG is the environment variable that carries the cluster specification and task details to each worker, letting MultiWorkerMirroredStrategy identify the chief and workers and establish collective communication. Without it correctly set per worker, distributed training on Vertex AI cannot coordinate.

Why this answer

In TensorFlow distributed training with MultiWorkerMirroredStrategy, the only required environment variable is `TF_CONFIG`. It provides the cluster topology and task identity, enabling gRPC communication between workers. The distribution strategy is defined in code, not via an environment variable. `TF_DISTRIBUTED_STRATEGY` is not a standard TensorFlow environment variable.

Exam trap

The exam may confuse candidates with plausible but incorrect environment variable names like TF_CONFIG_JSON or TF_DISTRIBUTED_STRATEGY, but only TF_CONFIG is required.

187
MCQhard

Two teams train models in separate Vertex AI projects but must share the same curated feature set. The platform team wants a single authoritative definition of each feature so that online serving and offline training always return consistent values, while each team keeps its own model training pipeline. Which approach should the platform team implement?

A.Publish the features as a BigQuery authorized view and instruct teams to call it from their pipelines.
B.Create the features once in a Vertex AI Feature Group with a BigQuery source and register them in a shared Feature Online Store, then grant both teams read access.
C.Have each team copy the feature engineering SQL into its own project and schedule it independently.
D.Export the feature table to a shared Cloud Storage bucket as Parquet and let both teams read the files.
AnswerB

Vertex AI Feature Registry and Feature Groups let the platform team define each feature once over a BigQuery source, and a Feature Online Store serves the same definitions online. Granting both teams read access means their separate pipelines consume identical feature definitions, which is exactly the consistency and single-source-of-truth requirement described.

Why this answer

A shared, authoritative feature definition is what prevents training-serving skew across teams. Vertex AI Feature Groups in Feature Registry define features once over a BigQuery source, and a Feature Online Store serves those same definitions at low latency, so both teams' pipelines consume identical feature semantics while retaining their own training pipelines. Copies, static Parquet exports, and authorized views all lack the managed online serving and registry linkage.

Exam trap

The trap here is equating shared data access with shared feature definitions, when access control mechanisms like authorized views do not provide a managed online store or consistent training-serving values.

188
MCQmedium

You are deploying a scikit-learn model to Vertex AI for online predictions. The model requires a custom preprocessing step that transforms raw JSON input into a feature vector before calling predict. You want to minimize latency and avoid managing infrastructure. What should you do?

A.Export the scikit-learn model to a TensorFlow SavedModel and deploy it with the TensorFlow pre-built container, using a preprocessing layer in the model.
B.Deploy the scikit-learn model using the pre-built scikit-learn container and implement preprocessing in a Cloud Function that calls the endpoint.
C.Create a custom container that includes the scikit-learn model and a Flask app that performs preprocessing and calls predict, then deploy it to a Vertex AI endpoint.
D.Use Vertex AI Batch Prediction with a pre-built scikit-learn container and a custom preprocessing script.
AnswerC

Vertex AI supports custom containers for prediction, allowing you to package the model, preprocessing logic, and a web server. By building a container with a Flask app that implements the preprocessing and prediction logic, you can deploy it to a Vertex AI endpoint. This approach minimizes latency because preprocessing and prediction happen in the same process, and Vertex AI manages the infrastructure, including scaling and health checks.

Why this answer

Vertex AI custom containers for prediction allow you to bundle the scikit-learn model, custom preprocessing code, and an HTTP server such as Flask. This keeps preprocessing and inference in the same process, reducing latency, and Vertex AI handles scaling and infrastructure management. Pre-built containers do not support arbitrary preprocessing, and batch prediction or model conversion would not meet the low-latency online requirement.

Exam trap

The trap here is assuming that a pre-built scikit-learn container can execute custom preprocessing, when in fact it only serves the model's predict method and requires raw inputs to be already preprocessed.

189
MCQeasy

A machine learning engineer wants to monitor model performance on Vertex AI for a regression model. Which metric is most appropriate to track the average prediction error?

A.F1 score
B.Precision
C.Accuracy
D.RMSE
AnswerD

RMSE measures the square root of mean squared prediction error, expressing average deviation in the target's original units. For a regression model, this directly quantifies average prediction error, unlike classification metrics such as accuracy or AUC.

Why this answer

RMSE (Root Mean Squared Error) is the most appropriate metric for tracking average prediction error in a regression model because it measures the standard deviation of residuals (prediction errors) in the same units as the target variable. On Vertex AI, RMSE is a built-in evaluation metric for regression models, directly quantifying how far predictions deviate from actual values on average.

Exam trap

Google Cloud often tests the distinction between classification and regression metrics, and the trap here is that candidates mistakenly apply classification metrics like F1, precision, or accuracy to a regression problem, not recognizing that RMSE is the standard for continuous prediction error.

How to eliminate wrong answers

Option A is wrong because F1 score is a classification metric that combines precision and recall, not applicable to regression tasks. Option B is wrong because precision measures the proportion of true positive predictions among all positive predictions, used only in classification contexts. Option C is wrong because accuracy is the ratio of correct predictions to total predictions, suitable for classification but meaningless for continuous-valued regression outputs.

190
MCQmedium

A company deploys a custom TensorFlow model to Vertex AI Endpoint for online predictions. After deployment, prediction latency is consistently high (over 500ms) even under low traffic. The model is CPU-only and the default machine type (n1-standard-2) is used. Which action will most likely reduce prediction latency?

A.Increase the max_replica_count to 10 to allow more parallel requests.
B.Change the machine type to n1-highcpu-16 with a GPU accelerator.
C.Set min_replica_count to 3 to ensure always-on capacity.
D.Increase the batch size in the prediction request.
AnswerB

Adding a GPU accelerator offloads TensorFlow's matrix operations from the CPU, directly addressing the CPU-only constraint causing the 500ms latency. The n1-highcpu-16 shape supplies proportionally more vCPUs and memory for preprocessing and batching, so inference completes faster even at low traffic.

Why this answer

Changing the machine type to n1-highcpu-16 with a GPU accelerator provides significantly more compute resources for the custom TensorFlow model. The n1-highcpu-16 offers 16 vCPUs (vs. 2 in n1-standard-2), which reduces CPU-bound inference time, and adding a GPU accelerates matrix operations common in TensorFlow models, directly reducing latency per request. Option A is wrong because increasing max_replica_count allows more parallel requests but does not improve the processing time of a single request.

Option C is wrong because setting min_replica_count ensures always-on capacity to avoid cold starts, but does not reduce steady-state latency. Option D is wrong because increasing batch size in the prediction request increases throughput by processing multiple inputs together, but does not reduce latency for a single prediction—it may actually increase the time to return a result for a given request.

191
Multi-Selecthard

A company trains a model using Vertex AI Training and then deploys it to Vertex AI Prediction. They notice that prediction requests fail with 'InvalidArgument: input tensor shape mismatch'. Which THREE are possible causes?

Select 3 answers
A.The model was exported in a different format than supported
B.The batch size in the request is too large
C.The input data types do not match the expected types (e.g., float vs int)
D.The input data has a different number of features than the model expects
E.The serving function does not include the same preprocessing as training
AnswersC, D, E

Tensor shape validation includes dtype checking, so passing float32 where the model signature expects int32 raises InvalidArgument before inference. The declared input schema fixes each tensor's dtype, and a mismatch is reported as a shape error rather than a type error.

Why this answer

Option C is correct because Vertex AI Prediction validates request tensors against the model's signature, and a dtype mismatch such as sending int32 where the signature expects float32 raises InvalidArgument due to incompatible tensor types. Option D is correct because the input tensor's feature dimension must match the model's expected input shape; supplying a different number of features produces the shape mismatch error. Option E is correct because if the serving function omits the preprocessing applied during training (for example, normalization, tokenization, or reshaping), the raw request tensor will not match the shape the model's signature expects, triggering the same error.

Option A is not the cause here because an unsupported export format would typically fail at model upload or deployment with a format/import error rather than at prediction time with a tensor shape mismatch. Option B is not the cause because an oversized batch size generally results in resource or memory errors, not an InvalidArgument tensor shape mismatch.

Exam trap

Google Cloud often tests the misconception that 'shape mismatch' only refers to the number of features or dimensions, when in fact it also encompasses data type mismatches and preprocessing inconsistencies that alter the tensor structure before it reaches the model.

192
MCQeasy

A data scientist wants to automatically generate model documentation that includes model purpose, training data, evaluation results, and intended use. Which tool should they use?

A.Vertex AI Workbench
B.Vertex AI Experiment
C.Cloud Datalab
D.Model Cards in Vertex AI
AnswerD

Model Cards in Vertex AI automatically capture and display model purpose, training data, evaluation metrics, and intended use, directly satisfying the documentation requirement. Unlike generic metadata or manual reports, Model Cards generate this structured governance artefact from the model's own evaluation results, giving the data scientist the exact fields the stem demands.

Why this answer

Model Cards in Vertex AI is a feature specifically designed to generate structured documentation for models, including purpose, training data, evaluation results, and intended use. It provides a standardized format for model transparency and governance. The other tools are for development, experimentation, or notebooks, not documentation.

Exam trap

The trap is confusing experiment tracking with model documentation; candidates might pick Vertex AI Experiments because it deals with models, but Model Cards is the dedicated tool for documentation.

How to eliminate wrong answers

Option A is wrong because Vertex AI Workbench is a Jupyter notebook environment for development, not a documentation tool. Option B is wrong because Vertex AI Experiment is for tracking and comparing ML experiments, not for generating model documentation. Option C is wrong because Cloud Datalab is an interactive data exploration tool, not a model documentation service.

193
MCQhard

Refer to the exhibit. The team wants to automatically deploy the best-performing model version to production. They have set up a Cloud Function triggered by Model Registry events. Which alias should they use in the function to get the latest champion?

A.'champion'
B.''
C.'experiment'
D.'latest'
AnswerA

The 'champion' alias conventionally indicates the best-performing production version.

Why this answer

The 'champion' alias is specifically reserved in Vertex AI Model Registry to denote the best-performing model version in production. By configuring the Cloud Function to trigger on the assignment of the 'champion' alias, the team ensures that only the model version promoted as the production champion is automatically deployed, aligning with MLOps best practices for staged model promotion.

Exam trap

Google Cloud often tests the distinction between 'champion' (a production alias) and 'latest' (a version number concept), leading candidates to incorrectly choose 'latest' because they confuse chronological recency with performance-based promotion.

How to eliminate wrong answers

Option B is wrong because an empty string is not a valid alias in MLflow; aliases must be non-empty strings, and using an empty string would cause the function to fail or match no events. Option C is wrong because 'experiment' is not a predefined alias in MLflow Model Registry; it refers to an MLflow Experiment, not a model version alias, and would not trigger on model promotion events. Option D is wrong because 'latest' is not a standard alias in MLflow; while MLflow can retrieve the latest model version by version number, the 'latest' alias does not exist, and using it would not capture the champion promotion event.

194
MCQmedium

A data analyst wants to train a binary classification model in BigQuery ML on a dataset of 10 million rows with 50 features. They need to evaluate the model's performance on a held-out test set. Which sequence of SQL statements should they run?

A.CREATE MODEL then ML.FEATURE_IMPORTANCE
B.ML.TRAIN then ML.EVALUATE
C.CREATE MODEL then ML.PREDICT
D.CREATE MODEL then ML.EVALUATE
AnswerD

BigQuery ML trains the model with CREATE MODEL, which handles the 10-million-row, 50-feature dataset natively in SQL. ML.EVALUATE then scores that trained model against the held-out test set, returning precision, recall, AUC and related metrics, satisfying the evaluation requirement without exporting data.

Why this answer

To train and evaluate a model in BigQuery ML, you use CREATE MODEL to train the model, and then ML.EVALUATE to assess its performance on a held-out dataset. CREATE MODEL automatically splits the data into training and evaluation sets if specified, or you can use a separate table for evaluation. ML.EVALUATE returns metrics like accuracy, precision, recall, etc.

This sequence is standard for model development in BigQuery ML.

Exam trap

The trap is confusing ML.EVALUATE with ML.PREDICT or thinking that ML.TRAIN exists. Candidates might also think they need to use ML.FEATURE_IMPORTANCE for evaluation.

How to eliminate wrong answers

Option A is wrong because ML.FEATURE_IMPORTANCE is used to understand feature importance, not to evaluate model performance. Option B is wrong because ML.TRAIN is not a valid BigQuery ML function; training is done via CREATE MODEL. Option C is wrong because ML.PREDICT is for generating predictions, not for evaluating performance on a test set.

195
Matchingmedium

Match each model evaluation metric to its use case.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Measure of false positives in classification

Measure of false negatives in classification

Harmonic mean of precision and recall

Root mean squared error for regression

Cross-entropy loss for probabilistic classification

Why these pairings

Accuracy is correctly matched with balanced classes and equal error cost; Precision with minimizing false positives; Recall with minimizing false negatives; F1 Score with balancing precision and recall for imbalanced classes. Option E incorrectly pairs Accuracy with 'when false positives are costly' which is actually the domain of Precision. The main trap is confusing accuracy with precision when dealing with asymmetric costs.

196
MCQeasy

You need to serve a TensorFlow model that has a cold start latency of 20 seconds. The model is used for a real-time application with unpredictable traffic, but occasional bursts require immediate responses. What is the best deployment strategy to minimize both cold start impact and cost?

A.Set min_replica_count to 1 to keep at least one instance always warm.
B.Use a larger machine type to reduce cold start time.
C.Set min_replica_count to 0 and rely on autoscaling to handle bursts.
D.Enable serving on Cloud Run for faster cold start.
AnswerA

Setting min_replica_count to 1 keeps one replica permanently loaded, so the 20-second TensorFlow cold start is paid only at deployment rather than on each burst. Bursts are absorbed by scaling additional replicas, and idle capacity stays minimal, balancing latency against cost.

Why this answer

Setting min_replica_count to 1 keeps one model instance always loaded, eliminating the 20-second cold start for the first request and ensuring immediate responses during bursts. Because only one replica is kept warm, cost stays low compared to maintaining multiple always-on replicas, and autoscaling can add more when traffic spikes. This balances latency and cost for unpredictable real-time workloads.

Exam trap

The trap is assuming a bigger machine or serverless platform eliminates cold start, when the real lever is keeping at least one replica warm while letting autoscaling handle bursts.

How to eliminate wrong answers

Option B is wrong because a larger machine type may reduce load time slightly but does not eliminate the cold start; the model still must be loaded into memory on each new instance. Option C is wrong because min_replica_count=0 means every burst triggers cold starts, directly violating the immediate-response requirement. Option D is wrong because Cloud Run is a container platform, not a native Vertex AI model-serving option, and it does not inherently solve a 20-second model load time.

197
Multi-Selectmedium

Which THREE considerations are important when setting up a shared feature store in Vertex AI Feature Store for multiple teams?

Select 3 answers
A.Enable feature monitoring for data quality and freshness
B.Use separate BigQuery tables for each team's features
C.Implement data governance policies for feature creation and access
D.Create a feature sharing policy to enable cross-team discovery
E.Allow each team to build independent ingestion pipelines
AnswersA, C, D

Feature monitoring detects drift and staleness across shared feature views, satisfying the multi-team requirement for trustworthy, reusable features. In Vertex AI Feature Store, monitoring alerts each consuming team when data quality or freshness degrades, preventing silent model failures without duplicating validation effort per team.

Why this answer

Option A is correct because enabling feature monitoring in Vertex AI Feature Store lets you track data quality and freshness metrics (such as drift and staleness) so that multiple consuming teams can trust the shared features. Option C is correct because a shared feature store requires data governance policies that define who can create, modify, and access features, ensuring consistent standards and compliance across teams. Option D is correct because a feature sharing policy enables cross-team discovery and reuse of features, which is the core purpose of a centralized shared feature store.

Option B is not appropriate because using separate BigQuery tables per team fragments the store and undermines the centralized sharing model. Option E is not appropriate because independent ingestion pipelines per team lead to duplicated, inconsistent feature definitions rather than a governed shared store.

Exam trap

Google Cloud often tests the misconception that a shared feature store requires separate physical storage per team (Option B) or fully independent ingestion (Option E), when in reality the value lies in centralization with controlled access and standardized pipelines.

198
MCQhard

A healthcare organization wants to build a model to predict patient readmission risk using structured electronic health record (EHR) data. They need to train a model using SQL in BigQuery, but they also want to leverage AutoML's ability to automatically search for the best architecture. Which approach should they take?

A.Use a pre-built Vision API model via BigQuery ML remote model
B.Use BigQuery ML with the AUTOML_CLASSIFIER model type
C.Use AutoML Tables with Vertex AI and export predictions
D.Use BigQuery ML with a DNN_CLASSIFIER and manual hyperparameter tuning
AnswerB

BigQuery ML's AUTOML_CLASSIFIER model type trains directly on BigQuery data using SQL while running AutoML architecture search and hyperparameter tuning behind the scenes. This satisfies both the SQL-in-BigQuery requirement and the demand for automated model selection on structured EHR data.

Why this answer

BigQuery ML's AUTOML_CLASSIFIER model type automatically performs architecture search and hyperparameter tuning, making it ideal for users who want to leverage AutoML capabilities directly within SQL on structured EHR data. This approach avoids manual model selection while staying entirely within BigQuery's SQL interface, which is the stated requirement.

Exam trap

The trap here is that candidates confuse AutoML Tables (a separate Vertex AI service) with BigQuery ML's built-in AUTO model type, assuming they must export data to use AutoML, when in fact BigQuery ML provides AutoML capabilities directly within SQL.

How to eliminate wrong answers

Option A is wrong because Vision API is designed for image analysis, not structured EHR data, and BigQuery ML remote models require a pre-built API endpoint, not AutoML architecture search. Option C is wrong because AutoML Tables (now Vertex AI Tabular) is a separate service that requires exporting data out of BigQuery and does not allow training via SQL in BigQuery. Option D is wrong because DNN_CLASSIFIER with manual hyperparameter tuning contradicts the requirement to 'automatically search for the best architecture' — it requires explicit user-specified parameters and does not perform automated architecture search.

199
Multi-Selectmedium

Which TWO practices are important when scaling a prototype ML model to production on Google Cloud? (Choose two.)

Select 2 answers
A.Set up model monitoring for data drift and concept drift
B.Manually engineer features for each training iteration
C.Run the model on a single high-memory Compute Engine VM
D.Use proprietary libraries to maximize performance regardless of lock-in
E.Implement CI/CD pipelines for model training and deployment
AnswersA, E

Production traffic inevitably diverges from training data, so monitoring for data drift and concept drift detects degradation before it harms users. This satisfies the scaling requirement by catching silent model decay that static offline evaluation cannot reveal once the prototype serves live requests.

Why this answer

Option A is correct because production ML models degrade as input data distributions shift (data drift) and as the relationship between features and labels changes (concept drift), so Vertex AI Model Monitoring must track these and trigger retraining or alerts. Option E is correct because scaling to production requires reproducible, automated MLOps: CI/CD pipelines (e.g., Cloud Build, Vertex AI Pipelines) automate training, evaluation, and deployment, ensuring consistency and fast iteration. Option B is wrong because manual feature engineering per iteration does not scale and should be replaced by automated feature pipelines (e.g., Vertex AI Feature Store).

Option C is wrong because a single high-memory Compute Engine VM is a non-scalable, non-resilient deployment; production serving should use managed, horizontally scalable endpoints like Vertex AI Prediction. Option D is wrong because proprietary, lock-in-heavy libraries hinder portability, maintainability, and integration with Google Cloud's managed ML services.

Exam trap

Google Cloud often tests the misconception that production ML can rely on manual processes or single-instance deployments, whereas the correct approach emphasizes automation, monitoring, and scalability through managed services.

200
Multi-Selectmedium

A company needs to reduce inference latency for their online prediction service on Vertex AI. Which two actions would help? (Choose 2)

Select 2 answers
A.Increase the maximum number of replicas
B.Deploy the model on a GPU-enabled machine
C.Enable model quantization via Vertex AI Model Optimization
D.Use a smaller machine type with less memory
E.Enable autoscaling with a lower target CPU utilization
AnswersB, C

GPUs accelerate the matrix operations underpinning neural network inference, cutting per-request compute time. This addresses the stem's latency constraint directly, since the bottleneck in online prediction is typically model computation rather than network transfer or storage access.

Why this answer

Option B is correct because deploying the model on a GPU-enabled machine type gives the inference workload hardware acceleration, which speeds up the matrix/tensor computations typical of ML models and thereby reduces per-request prediction latency. Option C is correct because enabling model quantization via Vertex AI Model Optimization reduces the numerical precision of weights and activations (for example to INT8), shrinking model size and memory bandwidth requirements and enabling faster inference, which lowers latency. Option A is not correct because increasing the maximum number of replicas improves throughput and availability under load, but it does not reduce the latency of an individual prediction.

Option D is not correct because using a smaller machine type with less memory would likely increase latency or cause out-of-memory failures rather than reduce it. Option E is not correct because autoscaling with a lower target CPU utilization only triggers scaling sooner to add capacity; it addresses throughput/load handling, not the intrinsic latency of a single inference.

Exam trap

Google often tests the distinction between scaling for throughput (replicas, autoscaling) versus reducing per-request latency (hardware acceleration, model optimization), leading candidates to confuse horizontal scaling with performance optimization.

201
MCQhard

A global retailer has deployed a real-time product recommendation model on Vertex AI Endpoints. The model is a large neural network that runs on a single node with 8 vCPUs and 30 GB memory. Over the past week, the p99 latency has increased from 200ms to 2 seconds, and the error rate has risen to 5%. Cloud Monitoring shows that the endpoint's CPU utilization is consistently near 100%, and memory is at 80%. The ML engineer suspects the model is too large for the node, but model size has not changed. Logs show no increase in request volume (steady at 50 QPS). There are no recent model updates. The engineer has tried to increase the node to 16 vCPUs, but latency decreased only slightly. What is the most likely root cause and the best first step to resolve it?

A.Profile the inference code to identify inefficient operations, such as unnecessary copies or suboptimal batch processing, and optimize the model serving logic.
B.Add more nodes to the endpoint by enabling autoscaling to distribute the load.
C.Retrain the model with a smaller architecture to reduce inference time.
D.Move the model to a machine type with more CPU cores and a GPU to accelerate inference.
AnswerA

Steady QPS, unchanged model size and only marginal gains from extra vCPUs point to CPU-bound serving code rather than capacity. Profiling the inference path exposes inefficient operations such as redundant tensor copies or poor batching, which optimisation resolves.

Why this answer

The p99 latency spike and high CPU utilization despite unchanged model size and request volume indicate a software bottleneck, not a hardware one. Profiling the inference code (Option A) can reveal inefficient operations like unnecessary data copies or suboptimal batch processing that degrade performance on the existing node. Since increasing vCPUs barely helped, the root cause is likely within the serving logic, not the compute capacity.

Exam trap

Google Cloud often tests the misconception that latency and CPU issues are always solved by scaling up hardware, when in fact software inefficiencies in the serving stack are a frequent root cause in ML deployments.

How to eliminate wrong answers

Option B is wrong because adding nodes via autoscaling would not address the root cause of high CPU utilization per node; it would only distribute the load, but each node would still suffer from the same inefficiency, and the steady 50 QPS suggests no need for more nodes. Option C is wrong because retraining with a smaller architecture is a long-term solution that ignores the immediate issue of serving inefficiency; the model size hasn't changed, and the problem is runtime performance, not model accuracy. Option D is wrong because moving to a GPU or more CPU cores treats the symptom (high CPU) rather than the cause; the minimal improvement from doubling vCPUs suggests the bottleneck is in software, not hardware, and a GPU would not fix inefficient code paths.

202
MCQhard

A media company wants to automatically moderate user-uploaded videos by detecting explicit content (e.g., violence, adult material). They need a solution that integrates with their video processing pipeline and scales to millions of videos. Which approach should they take?

A.Use Video Intelligence API with explicit content detection
B.Use AutoML Video to train a custom explicit content detection model
C.Use Natural Language API on video transcripts
D.Use Vision API to analyze each video frame
AnswerA

The Video Intelligence API provides explicit content detection purpose-built for video, analysing frames and audio for violence and adult material. It integrates into automated pipelines and scales elastically, satisfying the stem's requirement to moderate millions of videos without building custom models.

Why this answer

Video Intelligence API provides explicit content detection specifically designed to identify violence, adult material, and other explicit content in videos. It is a managed service that scales automatically and can be integrated into video processing pipelines via its API. This is the most direct and scalable solution for moderating user-uploaded videos.

Exam trap

The trap is overcomplicating the solution by considering custom model training (AutoML) or using image analysis frame-by-frame. The exam expects you to know that Video Intelligence API has built-in explicit content detection.

How to eliminate wrong answers

Option B is wrong because AutoML Video requires training a custom model, which is time-consuming and unnecessary when a pre-trained model for explicit content detection already exists. Option C is wrong because Natural Language API analyzes text, not video content; it would only work on transcripts and miss visual explicit content. Option D is wrong because Vision API analyzes individual images, not videos; processing each frame would be inefficient and costly, and it lacks temporal context.

203
Multi-Selecthard

You are using Vertex AI Model Monitoring to detect prediction drift on a deployed model that serves online predictions. You want to ensure that the monitoring job can correctly compute drift metrics for numerical features. Which two configurations are required for the monitoring job to compute drift for a numerical feature? (Choose two.)

Select 2 answers
A.Specify the feature's distribution as a numerical distribution and provide the training dataset as the reference.
B.Enable explainability on the endpoint to generate feature attributions for drift analysis.
C.Configure the monitoring job to use a default threshold for drift detection, such as 0.3 for the Jensen-Shannon distance.
D.Ensure that the feature's data type is correctly specified in the model's input schema.
E.Set the monitoring job's sampling rate to at least 0.5 to ensure sufficient data for statistical tests.
AnswersA, D

For numerical features, Vertex AI Model Monitoring requires a reference dataset to establish the baseline distribution. The monitoring job compares the live distribution to this baseline to compute drift. Without a reference, drift cannot be calculated. Thus, specifying the numerical distribution and providing the training dataset are essential.

Why this answer

To compute drift for numerical features, Vertex AI Model Monitoring needs a reference dataset to establish the baseline distribution and a correct input schema to parse the feature as numerical. These two configurations are essential; sampling rate, explainability, and thresholds are not required for the computation itself.

Exam trap

The trap here is focusing on sampling rate or thresholds as prerequisites, while the fundamental requirements are the reference dataset and correct schema.

204
MCQhard

A company has a TensorFlow model trained outside of Google Cloud and wants to use it for online predictions on Vertex AI. They have saved the model in SavedModel format. What is the most efficient way to deploy this model?

A.Import the model into BigQuery ML using CREATE MODEL with model_type='TENSORFLOW'
B.Use Vertex AI AutoML Tables to retrain the model
C.Use Cloud Functions to run the model for each prediction request
D.Upload the saved model to Vertex AI and create an endpoint for online predictions
AnswerD

Vertex AI accepts SavedModel artefacts directly, so uploading the existing model and deploying it to an endpoint avoids retraining or conversion. This is the most efficient route to online predictions for a TensorFlow model trained outside Google Cloud.

Why this answer

The most efficient way to deploy a TensorFlow SavedModel for online predictions on Vertex AI is to upload the SavedModel to Vertex AI and create an endpoint. Vertex AI supports importing custom models in SavedModel format and deploying them to endpoints for low-latency online predictions. This leverages Vertex AI's managed infrastructure for scaling and serving.

Exam trap

The trap is confusing deployment targets: BigQuery ML is for SQL predictions, Cloud Functions for lightweight tasks, and AutoML for training. The correct choice is Vertex AI for custom model serving.

How to eliminate wrong answers

Option A is wrong because importing into BigQuery ML is for SQL-based predictions, not for online serving; it also may not support all TensorFlow ops. Option B is wrong because AutoML Tables is for training models from tabular data, not for deploying existing TensorFlow models. Option C is wrong because Cloud Functions is not designed for serving ML models at scale and would have cold starts and limitations.

205
MCQmedium

An ML team is using Vertex AI to train a deep learning model on a large dataset. To reduce costs, they want to use preemptible VMs for training jobs. However, training must complete within a bounded time. Which strategy should they use?

A.Use Cloud TPU instead of GPU; TPUs are not preemptible.
B.Use Vertex AI Training without spot VMs, because preemptible VMs are not supported for training.
C.Use Vertex AI Training with spot VMs and ensure the training code saves checkpoints periodically to Cloud Storage.
D.Use a single powerful non-preemptible VM to avoid interruptions.
AnswerC

Spot VMs suit this scenario because Vertex AI automatically restarts preempted training jobs, satisfying the bounded-time constraint. Periodic checkpointing to Cloud Storage preserves progress across interruptions, so restarts resume from the last saved state rather than beginning again. This combination absorbs preemption while keeping costs low.

Why this answer

Vertex AI Training supports spot VMs (preemptible instances) for cost savings, and periodic checkpointing to Cloud Storage ensures that training can resume from the last saved state if a VM is preempted, allowing the job to complete within a bounded time despite interruptions.

Exam trap

A common misconception is that preemptible VMs are not supported in Vertex AI Training, but they are fully supported as spot VMs. The key to bounded-time completion is checkpointing to Cloud Storage for resumability.

How to eliminate wrong answers

Option A is wrong because Cloud TPUs are not inherently non-preemptible; they can also be preempted, and using TPUs does not address the cost-reduction goal with preemptible VMs. Option B is wrong because Vertex AI Training does support spot VMs (preemptible VMs) for training jobs, so the claim that they are not supported is incorrect. Option D is wrong because using a single powerful non-preemptible VM increases costs significantly and does not leverage the cost savings of preemptible instances, while still being susceptible to other failures without checkpointing.

206
MCQmedium

Your team trains models in a shared Vertex AI project. A data engineer accidentally overwrites a BigQuery training table that three production pipelines depend on, and nobody can tell which pipeline used which version of the data. You need to make dataset versions immutable and traceable so that any training run can be reproduced. What should you do?

A.Copy the table nightly into a Cloud Storage bucket using a scheduled Dataproc job.
B.Enable BigQuery table snapshots and record the snapshot ID in each pipeline's run metadata.
C.Enable BigQuery time travel and query the table with a FOR SYSTEM_TIME AS OF clause during retraining.
D.Grant the data engineer only bigquery.dataViewer on the table so they can no longer modify it.
AnswerB

BigQuery table snapshots create immutable, point-in-time copies of a table that persist independently of later writes, so a training job can always be reproduced from the exact snapshot it consumed. Recording the snapshot ID in pipeline metadata ties the artifact to the run, giving the traceability the team currently lacks without duplicating data or blocking the engineer's writes.

Why this answer

Immutable dataset versions are needed so any training run can be reproduced exactly, even after the source table changes. BigQuery table snapshots provide durable, read-only point-in-time copies, and storing the snapshot identifier alongside the pipeline run metadata creates the audit trail the team is missing. Access control, nightly copies, and time travel all fail to combine immutability with explicit per-run traceability.

Exam trap

The trap here is assuming that restricting write permissions or copying data on a schedule provides reproducibility, when only an immutable version identifier recorded per run actually ties a training job to its exact input data.

207
MCQmedium

A logistics company wants to classify shipping documents into categories (invoice, packing slip, bill of lading) using a custom model with minimal code. They have labeled training images. Which Google Cloud service is most appropriate?

A.Vertex AI AutoML Tables
B.AutoML Vision for image classification
C.Document AI custom extractor
D.Cloud Vision API with label detection
AnswerB

AutoML Vision trains a custom image classification model from labelled images through a graphical interface, requiring minimal code. It fits the three document categories and the labelled training images, unlike pre-built Vision API which cannot learn bespoke classes.

Why this answer

AutoML Vision for image classification is designed to train custom image classification models with minimal code using labeled images. It automatically handles data preprocessing, model selection, and hyperparameter tuning, making it ideal for classifying shipping documents from images. The other services are either for tabular data, document extraction, or pre-trained label detection without custom training.

Exam trap

The trap is confusing Document AI with AutoML Vision; candidates might think Document AI can classify documents, but it is primarily for extraction, while AutoML Vision is for custom image classification.

How to eliminate wrong answers

Option A is wrong because Vertex AI AutoML Tables is for tabular data, not images. Option C is wrong because Document AI custom extractor is for extracting structured data from documents, not for classifying document types. Option D is wrong because Cloud Vision API with label detection uses pre-trained models to detect general labels, but it cannot be customized to classify specific document categories like invoice, packing slip, or bill of lading.

208
MCQmedium

An ML engineer needs to run batch predictions on tens of petabytes of data using a trained model. The data is stored in Cloud Storage. Which service should they choose?

A.Cloud Dataflow with the model as a side input
B.Cloud Dataproc running Spark ML
C.Cloud Run with multiple revisions
D.Vertex AI Batch Prediction
AnswerD

Vertex AI Batch Prediction handles petabyte-scale batch inference directly from Cloud Storage, satisfying the tens-of-petabytes constraint. It streams data without loading it all into memory, distributing the job across managed compute, unlike online prediction endpoints, which suit low-latency single requests rather than massive offline scoring.

Why this answer

Vertex AI Batch Prediction is the correct choice because it is a managed service specifically designed for high-throughput, large-scale batch inference on data stored in Cloud Storage. It automatically handles sharding, scaling, and resource management for tens of petabytes, without requiring the engineer to manage infrastructure or write custom distributed processing code.

Exam trap

Google Cloud often tests the distinction between batch inference and data processing pipelines, so the trap here is that candidates confuse Cloud Dataflow (a data processing tool) with a batch prediction service, not realizing that Vertex AI Batch Prediction is the dedicated service for running models on large static datasets.

How to eliminate wrong answers

Option A is wrong because Cloud Dataflow with the model as a side input is optimized for stream and batch data processing pipelines, not for running a trained model's predictions on petabytes of static data; side inputs are not designed for large model inference and would cause severe performance bottlenecks and memory issues. Option B is wrong because Cloud Dataproc running Spark ML requires the engineer to manually manage clusters, configure Spark jobs for inference, and handle scaling, which adds operational overhead and is less efficient than a purpose-built batch prediction service for petabyte-scale data. Option C is wrong because Cloud Run is a serverless container platform for request-driven, low-latency applications, not for batch processing of tens of petabytes; it has a maximum request timeout of 60 minutes and cannot handle the volume or duration required.

209
Multi-Selecthard

A team is designing a ML pipeline that includes training, evaluation, and conditional deployment. They want to use Vertex AI Pipelines. Which THREE concepts should they use? (Choose three.)

Select 3 answers
A.Artifact types (e.g., Model, Metrics) for passing outputs
B.Manual approval via Cloud Console
C.Cloud SQL for storing intermediate results
D.Pre-built Google Cloud Pipeline Components for training and evaluation
E.dsl.If for conditional execution
AnswersA, D, E

Artifact types such as Model and Metrics carry typed outputs between pipeline steps, letting the evaluation component consume the trained model and emit metrics that the conditional deployment step reads. This satisfies the stem's need to pass outputs across training, evaluation and conditional deployment.

Why this answer

Option A is correct because Vertex AI Pipelines is built on ML Metadata, and typed artifacts such as Model, Metrics, Dataset, and Artifact let components pass structured outputs between training and evaluation steps so downstream steps and lineage tracking work correctly. Option D is correct because pre-built Google Cloud Pipeline Components (e.g., CustomTrainingJobOp, ModelEvaluationOp) provide ready-made, versioned steps for training and evaluation, reducing boilerplate and integrating natively with Vertex AI services. Option E is correct because conditional deployment requires branching logic in the pipeline graph, which is expressed with the Kubeflow Pipelines DSL construct dsl.If (or dsl.Condition) to run a deployment component only when evaluation metrics meet a threshold.

Option B is not appropriate because manual approval via the Cloud Console is not a Vertex AI Pipelines concept for conditional deployment; gating is done programmatically in the pipeline DAG. Option C is not appropriate because intermediate results in Vertex AI Pipelines are passed as artifacts and metadata, not stored in Cloud SQL, which is a relational database service unrelated to pipeline data flow.

Exam trap

PMLE often tests whether candidates confuse pipeline orchestration concepts with general GCP services, so distractors like Cloud SQL or manual approval must be recognized as non-pipeline constructs.

210
Multi-Selecteasy

A company wants to use Vertex AI JumpStart to deploy a pre-trained image classification model and later fine-tune it on their own data. Which TWO statements are true about Vertex AI JumpStart?

Select 2 answers
A.JumpStart requires users to build custom Docker containers for all models
B.JumpStart only supports text-based models
C.JumpStart allows you to fine-tune foundation models like Gemma
D.JumpStart only supports tabular data models
E.JumpStart provides one-click deployment of pre-trained models and ML solutions
AnswersC, E

JumpStart supports fine-tuning of foundation models such as Gemma.

Why this answer

Vertex AI JumpStart supports fine-tuning of foundation models like Gemma, allowing users to adapt pre-trained models to their specific datasets. This capability is built into JumpStart's managed environment, which handles the underlying infrastructure for training and deployment.

Exam trap

In the Google PMLE exam, candidates often mistakenly think that JumpStart only supports a narrow set of model types (e.g., text-only or tabular-only), when in fact it supports a broad range including image, text, and tabular models, and provides one-click deployment and fine-tuning capabilities.

211
MCQmedium

An ML engineer is building a Vertex AI pipeline that must run a custom training component for each of 12 hyperparameter combinations. The component is defined as a custom Python function (Lightweight Python component). The engineer wants each combination to run as a separate parallel task so the pipeline completes faster, and wants the pipeline to fail fast if any single trial fails. Which approach should the engineer take?

A.Wrap the training component in a single Vertex AI CustomJob with a hyperparameter tuning job specification and let the service manage trials.
B.Create 12 separate pipelines, one per hyperparameter combination, and schedule them with Cloud Scheduler at the same time.
C.Use a ParallelFor loop over the hyperparameter list and set the component's retry policy to 0.
D.Define a sequential for-loop in the pipeline function that calls the training component 12 times in order.
AnswerC

A ParallelFor loop in Vertex AI Pipelines (using dsl.ParallelFor) unrolls the loop into independent parallel tasks, one per hyperparameter combination, which is exactly the fan-out pattern needed for 12 trials. Setting the retry policy to 0 ensures that when any trial fails, the pipeline does not silently retry and mask the failure, so it fails fast as required.

Why this answer

Vertex AI Pipelines supports dsl.ParallelFor to fan out a list of values into parallel component tasks in one DAG, which is the idiomatic way to run 12 hyperparameter combinations concurrently. Setting retries to 0 on the component ensures a failed trial stops the pipeline rather than being retried and hidden. The other approaches either collapse trials into a single job, serialize them, or split them across pipelines, none of which meet the parallel fan-out and fail-fast requirements.

Exam trap

The trap here is assuming that a Vertex AI Hyperparameter Tuning CustomJob is the same as a pipeline ParallelFor fan-out, when the former is a single managed job and the latter is a pipeline-level DAG construct.

212
Multi-Selecthard

A machine learning team is building a feature engineering pipeline using Dataflow. They need to compute features from streaming data and store them in Vertex AI Feature Store for online serving. The features must be updated within 5 seconds of the event. Which TWO services should they combine? (Select 2)

Select 2 answers
A.Cloud Dataflow for stream processing and feature computation
B.Cloud Pub/Sub for event ingestion
C.Cloud Storage for feature store
D.Cloud Functions for feature transformation
E.BigQuery for feature storage
AnswersA, B

Dataflow can compute features in near real-time and write to Feature Store.

Why this answer

Cloud Dataflow is correct because it provides unified stream and batch processing with exactly-once semantics, enabling low-latency feature computation from streaming data. It integrates natively with Vertex AI Feature Store for online serving, ensuring features are updated within the required 5-second SLA.

Exam trap

The exam often tests the distinction between general-purpose storage services (Cloud Storage, BigQuery) and the dedicated online feature store (Vertex AI Feature Store) required for real-time ML serving, leading candidates to pick a storage option instead of the correct streaming ingestion (Pub/Sub) and processing (Dataflow) pair.

213
MCQeasy

A data scientist wants to share a trained model with the team for review before deployment. The model is stored in Vertex AI Model Registry. What is the recommended way to grant the team read access to the model?

A.Grant the IAM role 'roles/aiplatform.admin' to the team members.
B.Export the model as a local file and share it via a shared drive.
C.Grant the IAM role 'roles/aiplatform.viewer' to the team members on the project.
D.Add the team members to the Cloud Storage bucket ACL with 'READER' access.
AnswerC

Granting `roles/aiplatform.viewer` at project level gives team members read-only access to all Vertex AI resources, including registered models in Model Registry, satisfying the review-before-deployment requirement. It is the least-privilege predefined role that permits viewing models without granting deploy or edit permissions.

Why this answer

The 'roles/aiplatform.viewer' IAM role grants read-only access to Vertex AI resources, including models in the Model Registry. Option A is incorrect because 'roles/aiplatform.admin' grants full administrative access, which is too broad for read-only needs. Option B is wrong because exporting the model and sharing via a shared drive bypasses version control and security best practices.

Option D is incorrect because Cloud Storage bucket ACLs control access to the underlying bucket, not to the Vertex AI Model Registry; the model is managed through Vertex AI IAM.

214
MCQhard

A financial institution needs to extract structured data from scanned PDFs of loan applications, including text fields and tables. They require a human review step for high-risk applications. Which Google Cloud service and configuration should they use?

A.Document AI with a form parser processor and enable Human-in-the-Loop for high-risk applications
B.Document AI with a custom extractor processor and use Cloud Functions for human review
C.Cloud Vision API to detect text and tables, then send to Cloud Dataflow for processing
D.Vertex AI AutoML Vision to train a custom model for document parsing
AnswerA

Document AI's form parser processor extracts key-value pairs and tables from scanned PDFs via OCR and layout-aware parsing, satisfying the structured-data requirement. Enabling Human-in-the-Loop routes high-risk applications to human reviewers, directly meeting the mandated review step for those cases.

Why this answer

Document AI's Form Parser processor is purpose-built to extract structured key-value pairs and tables from forms like loan applications, and it natively integrates with Human-in-the-Loop (HITL) to route low-confidence or high-risk documents to human reviewers. This combination directly satisfies both the extraction and human review requirements without custom code.

Exam trap

PMLE often tests the misconception that any Document AI processor supports HITL — in reality, HITL is a specific feature that must be enabled and is best paired with Form Parser or custom extractors, not with generic OCR or Vision API.

How to eliminate wrong answers

Option B is wrong because while a custom extractor can be trained, it does not natively provide the HITL workflow — using Cloud Functions for review is a manual, non-integrated workaround that lacks Document AI's built-in confidence-based routing. Option C is wrong because Cloud Vision API only performs OCR/text detection and cannot natively parse structured form fields or tables into key-value pairs; Dataflow is a streaming/batch pipeline, not a document parser. Option D is wrong because AutoML Vision is an image classification/object detection tool, not designed for structured document field extraction from PDFs.

215
MCQhard

An ML engineer is monitoring a deployed model on Vertex AI Endpoints and wants to detect anomalies in the distribution of a categorical feature 'product_category' which has 50 possible values. The engineer configures Vertex AI Model Monitoring with training-serving skew detection. After a week, they receive an alert that the feature is skewed. Upon investigation, they find that a new product category was introduced in the serving data that was not present in the training data. What is the most likely reason for the alert?

A.The monitoring system treats the new category as an unknown value and flags it as skew.
B.The model's predictions are incorrect for the new category, causing the monitoring system to flag the feature.
C.The training data already contained the new category but it was not properly encoded.
D.The monitoring configuration has a bug that misinterprets the new category as a numerical value.
AnswerA

Training-serving skew detection compares the distribution of feature values. If a new category appears in serving data that was not in training, it is considered an unknown value. The monitoring system detects this as a divergence from the training distribution, triggering a skew alert. This is a common occurrence when categorical features have evolving vocabularies.

Why this answer

Vertex AI Model Monitoring's skew detection compares the distribution of feature values between training and serving. When a new category appears in serving that was absent in training, the distribution diverges, and the system flags it as skew. This is because the model has never seen that category and may not handle it well, so monitoring alerts the team to investigate.

Exam trap

The trap here is assuming that skew alerts are triggered by model performance issues, when they are actually triggered by distribution differences in the input features.

216
MCQhard

A team has deployed a model on a Vertex AI Endpoint and enabled Vertex AI Model Monitoring for feature skew. They notice that the skew metric for a categorical feature with high cardinality is consistently high, even though the feature's distribution appears stable to the team. What is the most likely cause of this high skew metric?

A.The training dataset and live data have different sets of categories for that feature, causing divergence.
B.The monitoring job is sampling too few requests, leading to statistical noise.
C.The model's predictions are drifting, which indirectly increases feature skew.
D.The feature is not included in the training dataset schema, so the skew computation fails.
AnswerA

High-cardinality categorical features often have categories present in live data that were not in the training data, or vice versa. This mismatch in category sets leads to large divergence metrics, such as Jensen-Shannon, because the distributions have disjoint support. Even if the overall distribution seems stable, the presence of unseen categories can cause high skew. Therefore, this is the most likely cause.

Why this answer

For high-cardinality categorical features, the presence of categories in live data that were not in the training data (or vice versa) can cause large divergence in distribution comparisons. Vertex AI Model Monitoring uses metrics like Jensen-Shannon divergence, which are sensitive to disjoint category sets. This mismatch can result in consistently high skew even if the overall distribution appears stable.

The correct cause is the difference in category sets between training and live data.

Exam trap

The trap here is assuming that high skew always indicates a shift in the overall distribution, when it can be caused by new or missing categories in high-cardinality features.

217
MCQhard

You have an edge device with limited compute resources. You need to deploy a deep learning model for real-time inference. Which model compression technique should you apply to reduce the model size and latency with minimal accuracy loss?

A.Pruning only
B.Post-training quantization to INT8
C.Knowledge distillation only
D.Use full precision FP32 to maintain accuracy
AnswerB

Post-training quantization to INT8 converts trained weights and activations from FP32 to 8-bit integers, cutting model size roughly fourfold and accelerating inference on constrained edge hardware. It requires no retraining, satisfying the limited-compute constraint, while typically preserving accuracy with minimal degradation for real-time deployment.

Why this answer

Post-training quantization to INT8 reduces model weights and activations from 32-bit floats to 8-bit integers, cutting model size by roughly 4x and speeding up inference on edge hardware with minimal accuracy loss. It requires no retraining and works with most trained models. This makes it the best fit for limited-compute edge deployment.

Exam trap

The trap is assuming pruning or distillation alone delivers the same size/latency reduction as quantization, when quantization is the fastest, most reliable post-training compression for edge inference.

How to eliminate wrong answers

Option A is wrong because pruning alone removes weights but does not reduce numeric precision, so latency and size gains are smaller and often require fine-tuning to recover accuracy. Option C is wrong because knowledge distillation requires training a separate smaller student model, which is more complex and not a pure compression technique applied post-hoc. Option D is wrong because FP32 is the uncompressed baseline and offers no size or latency reduction, defeating the purpose of edge deployment.

218
MCQeasy

A data scientist wants to log prediction inputs and outputs for model monitoring. Which Google Cloud service is best suited for this?

A.Cloud Monitoring
B.Cloud Storage
C.Cloud Logging
D.BigQuery
AnswerC

Cloud Logging captures prediction inputs and outputs as log entries, providing the raw request and response records that Vertex AI Model Monitoring and downstream analysis consume. It is the native service for storing and querying these prediction payloads.

Why this answer

Cloud Logging is the best choice because it is designed to ingest, store, and analyze log data, including custom log entries from applications. The data scientist can use the Cloud Logging API to write structured log entries containing prediction inputs and outputs, then query them using Logs Explorer or export them for further analysis. This aligns with the requirement to log prediction inputs and outputs for model monitoring, as Cloud Logging provides a centralized, scalable, and queryable log management service.

Exam trap

Google Cloud often tests the distinction between logging (Cloud Logging) and monitoring (Cloud Monitoring), where candidates mistakenly choose Cloud Monitoring because they think 'monitoring' includes logging, but Cloud Monitoring is for metrics and alerts, not for storing and querying log data.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring is focused on collecting metrics, uptime checks, and alerting on system performance (e.g., CPU utilization, latency), not on storing and querying arbitrary log data like prediction inputs and outputs. Option B is wrong because Cloud Storage is an object storage service for unstructured data (e.g., images, backups), not a log management service; it lacks native querying capabilities for log entries and is not designed for real-time log ingestion and search. Option D is wrong because BigQuery is a serverless data warehouse for analytical queries on large structured datasets, not a log management service; while it can store logs exported from Cloud Logging, it is not the primary service for ingesting and querying log entries in real time.

219
Multi-Selecthard

A company is migrating from an on-premises ML serving infrastructure to Vertex AI. They have multiple models that need to be served from the same endpoint with different traffic percentages. They also need to monitor prediction quality. Which THREE actions should they take? (Choose 3)

Select 3 answers
A.Deploy multiple model versions on the same endpoint with traffic_split parameter.
B.Deploy each model as a separate endpoint and use Cloud Load Balancing.
C.Enable Vertex AI Model Monitoring to detect prediction drift.
D.Use Cloud Monitoring to create custom metrics based on business outcomes.
E.Export logs to BigQuery for manual analysis only.
AnswersA, C, D

Deploying multiple models to one endpoint with the `traffic_split` parameter directly satisfies the requirement for serving several models from a single endpoint at differing traffic percentages. Vertex AI assigns each deployed model a percentage of incoming prediction requests, enabling gradual rollouts or A/B testing without provisioning separate endpoints.

Why this answer

Option A is correct because Vertex AI endpoints natively support serving multiple DeployedModels behind a single endpoint, and the traffic_split parameter on the endpoint lets you route configurable percentages of prediction traffic to each model version. Option C is correct because Vertex AI Model Monitoring is the managed service that detects training-serving skew and prediction drift on deployed models, directly addressing the requirement to monitor prediction quality. Option D is correct because prediction quality often depends on business outcomes that Model Monitoring does not capture, so Cloud Monitoring custom metrics let you track those outcome-based signals alongside Vertex AI metrics.

Option B is not appropriate because separate endpoints with Cloud Load Balancing adds unnecessary complexity and does not provide Vertex AI's built-in traffic splitting or model monitoring integration. Option E is insufficient because exporting logs to BigQuery for manual analysis only is not an automated monitoring solution and does not satisfy the prediction-quality monitoring requirement.

Exam trap

The trap here is that candidates may think separate endpoints with a load balancer (Option B) are required for traffic distribution, overlooking Google Vertex AI's built-in traffic splitting on a single endpoint, which is simpler and more aligned with the platform's design.

220
Multi-Selecthard

A company is building a document processing pipeline using Document AI to extract data from invoices. They want to ensure high accuracy and handle edge cases where the model may be uncertain. Which THREE steps should they include in their pipeline?

Select 3 answers
A.Regularly retrain the processor using human-verified data
B.Use the pre-built invoice parser without any modifications
C.Use AutoML Vision to classify invoice types
D.Enable Human-in-the-Loop (HITL) to review documents with low confidence scores
E.Use a custom processor trained on their specific invoice format
AnswersA, D, E

Regular retraining with human-verified data directly addresses the accuracy constraint by feeding corrected edge-case extractions back into the processor, letting it learn the specific invoice variations it previously misread. This closes the loop on uncertain predictions, since human review resolves ambiguity that the model alone cannot, progressively raising extraction confidence across the pipeline.

Why this answer

Option A is correct because regularly retraining the processor with human-verified data continuously improves extraction accuracy and adapts the model to new invoice variations and edge cases over time. Option D is correct because enabling Human-in-the-Loop (HITL) routes documents with low confidence scores to human reviewers, ensuring uncertain or edge-case extractions are validated and corrected before entering downstream systems. Option E is correct because a custom processor trained on the company's specific invoice format captures their unique layouts, fields, and terminology, yielding higher accuracy than a generic model.

Option B is not appropriate because using the pre-built invoice parser without modification offers no tuning for the company's specific formats and provides no mechanism for handling uncertainty. Option C is not appropriate because AutoML Vision is an image classification service, not a document entity-extraction tool, and classifying invoice types does not extract the required field data or address low-confidence edge cases.

Exam trap

PMLE often tests the misconception that a pre-built parser is sufficient for all use cases — candidates overlook that custom training and HITL are required for high accuracy on domain-specific documents.

221
MCQmedium

A retailer uses BigQuery ML to build a linear regression model for sales forecasting. The model's evaluation shows high RMSE. Which step should they take first?

A.Use a more complex model like XGBoost
B.Increase the number of features
C.Set a larger training budget
D.Examine the data for outliers and missing values
AnswerD

High RMSE often stems from data quality problems, so inspecting for outliers and missing values addresses the root cause before tuning the model or adding features. This diagnostic step is the logical first action when evaluation error is unexpectedly large.

Why this answer

High RMSE in a linear regression model often indicates issues with data quality, such as outliers or missing values, which can disproportionately skew the model's predictions. BigQuery ML's linear regression is sensitive to such anomalies, so examining and cleaning the data is the most appropriate first step before considering model complexity or feature engineering.

Exam trap

Google Cloud often tests the misconception that high RMSE is always a model complexity issue, leading candidates to jump to advanced algorithms or feature engineering without considering fundamental data quality checks.

How to eliminate wrong answers

Option A is wrong because switching to a more complex model like XGBoost without first addressing data quality issues can amplify overfitting and does not fix the root cause of high RMSE. Option B is wrong because blindly increasing the number of features can introduce noise and multicollinearity, potentially worsening RMSE rather than improving it. Option C is wrong because setting a larger training budget in BigQuery ML does not improve model accuracy; it only allocates more resources for training, which is irrelevant when the issue stems from data problems.

222
MCQmedium

An ML team wants to run a hyperparameter tuning job on Vertex AI using a pre-built pipeline component. Which component should they use?

A.AutoMLTabularTrainingJobRunOp
B.CustomTrainingJobRunOp with hyperparameter arguments.
C.ModelTrainComponent
D.HyperparameterTuningJobRunOp
AnswerD

HyperparameterTuningJobRunOp is the pre-built pipeline component that wraps Vertex AI's HyperparameterTuningJob, launching a tuning job with the specified search space and metrics. It satisfies the stem's requirement to run tuning via a pre-built component rather than custom code.

Why this answer

The HyperparameterTuningJobRunOp is the correct pre-built Vertex AI pipeline component specifically designed to launch a hyperparameter tuning job. It wraps the Vertex AI HyperparameterTuningJob API, allowing you to specify the worker pool spec, metric target, and parameter specifications directly within a Kubeflow Pipelines (KFP) or Vertex AI Pipelines orchestration context.

Exam trap

A common mistake on the Google PMLE exam is confusing the pre-built HyperparameterTuningJobRunOp with CustomTrainingJobRunOp that accepts hyperparameter arguments, but the latter requires manual tuning logic rather than leveraging the built-in hyperparameter tuning service.

How to eliminate wrong answers

Option A is wrong because AutoMLTabularTrainingJobRunOp is used to launch an AutoML training job for tabular data, which does not support custom hyperparameter tuning; it uses AutoML's own search. Option B is wrong because CustomTrainingJobRunOp with hyperparameter arguments is not a pre-built component for tuning; it launches a single custom training job and would require manual orchestration to implement a tuning loop, whereas the question asks for a pre-built component. Option C is wrong because ModelTrainComponent is not a standard pre-built Vertex AI pipeline component; it is a generic name that does not correspond to any official Vertex AI component, and using it would require custom implementation.

223
MCQhard

Two teams independently develop two different versions of a model for the same use case. They both deploy to the same Vertex AI endpoint, causing conflicts. What is the best way to manage multiple model versions and avoid conflicts in a collaborative environment?

A.Have each team work on a separate Google Cloud project
B.Use custom metadata to tag each version and rely on team coordination
C.Deploy each team's model to a separate endpoint
D.Use Vertex AI Model Registry with staging and production channels, and implement CI/CD to control promotions
AnswerD

Vertex AI Model Registry assigns each version a unique resource ID and tracks lineage, so two teams' artefacts never overwrite each other on one endpoint. Staging and production channels, driven by CI/CD promotions, enforce the controlled release the stem's conflict-avoidance constraint requires.

Why this answer

Vertex AI Model Registry with staging and production channels provides a centralized system to manage model versions, track lineage, and control promotions via CI/CD pipelines. This prevents conflicts by enforcing a structured workflow for version updates. Option A is wrong because separate projects increase management overhead and do not address versioning within the same endpoint.

Option B is wrong because custom metadata lacks enforcement of deployment order and can lead to manual errors. Option C is wrong because deploying to separate endpoints does not resolve version conflicts; it merely isolates models, increasing complexity and cost.

224
MCQmedium

A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?

A.Deploy a single model version with minReplicaCount=0 and maxReplicaCount=10, and use a custom prediction routine to handle both versions based on request headers.
B.Create two separate endpoints, one for each model version, and use a global load balancer to distribute traffic equally between them.
C.Deploy the model to a Vertex AI Batch Prediction job and schedule it to run every hour, then use a Cloud Function to serve predictions from the latest batch results.
D.Deploy both model versions to the same endpoint, set minReplicaCount to at least 1 for each, and use traffic splitting to route a percentage of requests to each version.
AnswerD

This configuration ensures high availability with dedicated replicas for each model version, allows precise traffic splitting for gradual rollout, and minimizes cold starts by keeping at least one replica warm. It directly addresses all requirements: availability, traffic control, and latency.

Why this answer

Deploying both versions to the same endpoint with dedicated replicas and traffic splitting satisfies availability, precise rollout control, and low latency. It leverages Vertex AI's native capabilities for version management and gradual traffic shifting without additional infrastructure.

Exam trap

The trap here is assuming that separate endpoints or batch processing can achieve the same level of control and low latency as a single endpoint with traffic splitting and minimum replicas.

225
MCQeasy

A company needs to extract entities (e.g., names, dates) from customer emails using a pre-trained model. Which service should they use?

A.Translation API
B.Natural Language API
C.Dialogflow
D.Vision API
AnswerB

The Natural Language API provides pre-trained entity extraction, recognising people, dates, locations and organisations out of the box. Training a custom model is unnecessary here, so it satisfies the requirement to use a pre-trained model on email text with no labelled data or training effort.

Why this answer

The Natural Language API is designed to extract entities such as names, dates, locations, and other information from text. It provides pre-trained models for entity recognition, sentiment analysis, and syntax analysis, making it the correct choice for extracting entities from customer emails.

Exam trap

PMLE often tests the difference between similar AI services — candidates may confuse Natural Language API with Translation or Vision API, but the key is that entity extraction from text is a natural language processing task.

How to eliminate wrong answers

Option A is wrong because the Translation API is used to translate text from one language to another, not to extract entities. Option C is wrong because Dialogflow is used to build conversational interfaces (chatbots) and does not provide general entity extraction from arbitrary text. Option D is wrong because the Vision API is used to analyze images, such as detecting objects or text in images, not for extracting entities from text emails.

Page 2

Page 3 of 11

Page 4

All pages