Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 1–75

775 questions total · 11pages · All types, answers revealed

Page 1 of 11

Page 2
1
MCQmedium

A company needs to extract key fields from scanned invoices, such as invoice number and total amount, with high accuracy. They want a managed service and plan to use human review for low-confidence results. Which combination of services should they use?

A.Vision API and Natural Language API
B.Document AI and Human-in-the-Loop
C.Translation API and AutoML Vision
D.BigQuery ML and Vertex AI Prediction
AnswerB

Document AI extracts invoice fields via prebuilt or custom models, returning per-field confidence scores. Human-in-the-Loop routes only low-confidence extractions to reviewers, satisfying the stem's accuracy requirement while keeping the service fully managed. This pairing directly matches the stated need for human review of uncertain results.

Why this answer

Document AI is Google Cloud's managed document understanding service with specialized parsers (Invoice, Expense, Form) that extract structured fields like invoice number and total amount with high accuracy. Human-in-the-Loop (HITL) integrates directly with Document AI to route low-confidence predictions to human reviewers, whose corrections feed back to improve the processor. Together they satisfy the managed-service and human-review requirements in one pipeline.

Exam trap

PMLE often tests the confusion between generic Vision/NLP APIs and purpose-built Document AI — candidates assume any OCR service can extract invoice fields, missing that schema-aware extraction plus a native human-review loop is the differentiator.

How to eliminate wrong answers

Option A is wrong because Vision API performs generic OCR/image labeling and Natural Language API does entity/sentiment analysis — neither understands invoice schema or returns structured invoice fields. Option C is wrong because Translation API only converts languages and AutoML Vision is a custom image classifier, not a document field extractor; you would have to build and label everything yourself. Option D is wrong because BigQuery ML and Vertex AI Prediction are general ML training/serving tools — they do not provide pre-built document parsing or a human-review workflow.

2
MCQmedium

You are training a scikit-learn model on Vertex AI using a custom training job. The training dataset is a 2 TB CSV file stored in Cloud Storage, and the job must run on a single CPU-only VM. Loading the entire file into memory fails because the machine has only 32 GB of RAM. You need to train the model without increasing the VM size and without rewriting the training code to use a distributed framework. What should you do?

A.Use Vertex AI Pipelines to split the CSV into many smaller files, then run a separate training job for each file and average the model weights.
B.Convert the CSV to TFRecord format and use tf.data to stream batches, then train the scikit-learn model on the streamed batches.
C.Use the scikit-learn partial_fit method on an estimator that supports incremental learning, reading the CSV in chunks with pandas.read_csv and feeding each chunk to partial_fit.
D.Mount the Cloud Storage bucket as a local filesystem on the training VM and call pandas.read_csv on the mounted path, relying on the OS page cache to keep memory usage low.
AnswerC

Many scikit-learn estimators such as SGDClassifier and MiniBatchKMeans implement partial_fit, which updates the model incrementally on small batches. Reading the 2 TB CSV with pandas.read_csv in chunks keeps only one chunk in memory at a time, so the 32 GB VM is sufficient. This preserves a single model trained over all rows and requires no distributed framework, which matches the constraint exactly.

Why this answer

Incremental learning with partial_fit is the correct approach because it lets a scikit-learn estimator update its parameters one chunk at a time. Reading the CSV in chunks with pandas.read_csv bounds memory to the chunk size rather than the full 2 TB. This keeps the single-model semantics, avoids distributed training, and fits within the 32 GB VM, directly satisfying all the stated constraints.

Exam trap

The trap here is assuming that changing the file format or mounting Cloud Storage reduces the memory required by scikit-learn, when the real fix is an estimator that supports incremental partial_fit updates.

3
Multi-Selectmedium

A company wants to analyze videos to detect objects and track their movement over time. Which TWO Google Cloud services are suitable for this task?

Select 2 answers
A.AutoML Vision
B.Speech-to-Text
C.AutoML Video
D.Video Intelligence API
E.Natural Language API
AnswersC, D

AutoML Video lets you train a custom model on labelled video to detect and track objects specific to your domain, then serve predictions on new footage. This satisfies the movement-tracking requirement when pre-trained labels are insufficient.

Why this answer

AutoML Video supports object tracking, and Video Intelligence API provides both object detection and tracking. AutoML Vision is for images only, Natural Language for text, and Speech-to-Text for audio.

4
Multi-Selecthard

You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?

Select 3 answers
A.Store intermediate data in Cloud Storage with unique run IDs.
B.Pass large datasets between components as serialized in-memory objects.
C.Monitor feature distributions in training data vs. serving data to detect skew.
D.Use the same random seed for every run to ensure reproducibility.
E.Ensure each component produces deterministic outputs given the same inputs.
AnswersA, C, E

Unique run IDs in Cloud Storage paths make each pipeline execution write to distinct locations, so reruns neither overwrite nor reuse stale intermediates. This directly satisfies the idempotency requirement, since identical inputs produce isolated, reproducible outputs rather than colliding with prior runs.

Why this answer

Option A is correct because writing intermediate artifacts to Cloud Storage under a unique run ID (e.g., gs://bucket/run_id/...) isolates each pipeline execution, so reruns don't overwrite or collide with prior outputs — a key requirement for idempotency. Option C is correct because comparing training feature distributions against serving feature distributions (e.g., via statistics like mean, variance, or histogram distance) is the standard way to detect training/serving skew and trigger remediation. Option E is correct because deterministic components that yield identical outputs for identical inputs make reruns safe and reproducible, which is the essence of an idempotent pipeline.

Option B is wrong because passing large datasets as serialized in-memory objects is fragile, memory-bound, and non-idempotent; components should exchange data via durable storage references instead. Option D is wrong because a fixed random seed only aids reproducibility of stochastic steps and does nothing to guarantee idempotency or address training/serving skew.

Exam trap

PMLE often tests the difference between reproducibility (same seed) and idempotency (safe reruns) — candidates pick the seed option because it 'sounds like best practice' but it does not address pipeline idempotency or skew.

5
MCQmedium

You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?

A.Package the model in a custom container without any inference server.
B.Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.
C.Use Vertex AI Model Optimization to quantize the model to FP16 and deploy using the optimized model.
D.Use batch prediction instead of online prediction to reduce latency.
AnswerB

NVIDIA Triton Inference Server provides dynamic batching and concurrent model execution, directly maximising GPU utilisation for PyTorch tensors. This satisfies the stem's twin constraints of minimising prediction latency and maximising throughput, whereas single-request serving containers leave GPU capacity idle between calls.

Why this answer

NVIDIA Triton Inference Server provides advanced features like dynamic batching, concurrent model execution, and GPU scheduling that maximize throughput and minimize latency for GPU-accelerated inference. Vertex AI's prebuilt PyTorch serving container with Triton is specifically designed to handle online prediction workloads efficiently, outperforming a plain custom container without an inference server.

Exam trap

A common pitfall is assuming that model optimization alone (e.g., quantization) is sufficient for low-latency serving, when in fact the inference server's request handling and batching capabilities are critical for minimizing latency and maximizing throughput in online predictions on Vertex AI.

How to eliminate wrong answers

Option A is wrong because a custom container without any inference server lacks request batching, model queuing, and GPU utilization optimizations, leading to higher latency and lower throughput under concurrent requests. Option C is wrong because Vertex AI Model Optimization for FP16 quantization reduces model size and can improve throughput, but it does not address the serving infrastructure needed for low-latency online predictions; the deployment still requires an inference server like Triton to handle request management and GPU scheduling. Option D is wrong because batch prediction is designed for high-throughput, offline processing of large datasets and typically has higher latency per request due to job queuing and resource provisioning, making it unsuitable for minimizing prediction latency in online scenarios.

6
MCQeasy

A data scientist has deployed a model on Vertex AI Endpoints and wants to monitor the model's predictions for any drift over time. Which Vertex AI service should they use?

A.Vertex AI Feature Store
B.Vertex AI Predictions
C.Vertex AI Explainable AI
D.Vertex AI Model Monitoring
AnswerD

Vertex AI Model Monitoring continuously evaluates deployed endpoint predictions against a training baseline, detecting training-serving skew and prediction drift. It satisfies the requirement to monitor predictions over time, unlike feature-level logging or scheduled batch jobs, which capture data but perform no drift computation.

Why this answer

Vertex AI Model Monitoring is the purpose-built service for detecting drift in deployed models on Vertex AI Endpoints. It continuously compares incoming prediction requests against a training baseline and computes statistical drift metrics (e.g., Jensen-Shannon divergence) for features and, optionally, predictions. This is exactly the capability the data scientist needs to detect prediction drift over time.

Exam trap

PMLE often tests the distinction between feature drift (input distribution shift), prediction drift (output distribution shift), and feature skew (training-serving skew) — candidates confuse these three monitoring types and pick the wrong one.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing ML features — it does not monitor deployed endpoints for drift. Option B is wrong because Vertex AI Predictions is the serving/inference component that returns predictions; it does not perform drift analysis. Option C is wrong because Explainable AI provides feature attributions (e.g., SHAP values) for individual predictions to explain why a model produced an output, not to detect distributional drift.

7
MCQmedium

A data engineer wants to use BigQuery ML to train a model that predicts customer churn using a table with customer features and a label column. They want to use a deep neural network. Which model type should they specify?

A.BOOSTED_TREE_CLASSIFIER
B.LOGISTIC_REG
C.DNN_CLASSIFIER
D.DNN_REGRESSOR
AnswerC

DNN_CLASSIFIER fits because churn prediction is binary classification, and the stem specifies a deep neural network. BigQuery ML's DNN_CLASSIFIER trains a feed-forward neural network with a logistic output layer, satisfying both the label-column requirement and the requested architecture, unlike DNN_REGRESSOR, which predicts continuous values.

Why this answer

The DNN_CLASSIFIER model type in BigQuery ML is specifically designed for classification tasks using a deep neural network architecture. Since the problem is predicting customer churn (a binary classification problem) and the data engineer explicitly wants to use a deep neural network, DNN_CLASSIFIER is the appropriate choice.

Exam trap

The trap is that candidates may confuse DNN_REGRESSOR with DNN_CLASSIFIER, but BigQuery ML uses different model types for regression vs. classification. DNN_CLASSIFIER is used for classification tasks like churn prediction, while DNN_REGRESSOR is for continuous values.

How to eliminate wrong answers

Option A is wrong because BOOSTED_TREE_CLASSIFIER uses gradient-boosted decision trees, not a deep neural network, so it does not meet the requirement for a DNN model. Option B is wrong because LOGISTIC_REG is a logistic regression model, which is a linear classifier and not a deep neural network. Option D is wrong because DNN_REGRESSOR is used for regression tasks (predicting continuous values), not for classification tasks like churn prediction.

8
MCQmedium

You are responsible for monitoring a batch prediction pipeline that runs daily. Recently, the pipeline started failing intermittently with out-of-memory errors. The input data volume has not changed. What is the most likely cause?

A.A recent code change that loads the entire dataset into memory before processing
B.Increase in model size due to retraining
C.Decrease in the number of worker machines
D.Increase in input data size
AnswerA

Loading the full dataset into memory scales memory use with input size, so a code change that materialises everything before processing explains intermittent out-of-memory failures despite unchanged data volume; the constraint is that input volume stayed constant, ruling out data growth.

Why this answer

A code change that loads the entire dataset into memory before processing would directly cause out-of-memory (OOM) errors, even if the input data volume remains unchanged. In batch prediction pipelines, data is typically streamed or processed in chunks to manage memory efficiently. A change that bypasses this pattern and loads all data at once can exceed the available heap or container memory, leading to intermittent failures depending on data characteristics or concurrent loads.

Exam trap

The trap here is that candidates may assume OOM errors are always caused by increased data volume or resource scaling issues, but the question explicitly states data volume is unchanged, forcing you to consider code-level changes that alter memory access patterns.

How to eliminate wrong answers

Option B is wrong because an increase in model size due to retraining would affect memory usage during model loading or inference, but it would not cause intermittent OOM errors if the input data volume is unchanged; model size changes are typically gradual and would cause consistent failures, not intermittent ones. Option C is wrong because a decrease in the number of worker machines would reduce total available memory, but the question states the input data volume has not changed, so this would cause consistent OOM errors on every run, not intermittent ones. Option D is wrong because the question explicitly states that input data volume has not changed, so an increase in data size cannot be the cause.

9
MCQeasy

A data science team has trained a TensorFlow model and wants to serve it online with minimal latency. Which Vertex AI deployment option should they use to ensure the model can handle traffic spikes without manual scaling?

A.Use Vertex AI Model Garden.
B.Deploy the model to a Vertex AI Endpoint with automatic scaling.
C.Use Vertex AI Batch Prediction for offline inference.
D.Deploy the model to a Compute Engine VM with a load balancer.
AnswerB

Deploying to a Vertex AI Endpoint with automatic scaling directly satisfies the low-latency and spike-handling constraints. Endpoints provision dedicated prediction nodes, and autoscaling adjusts replica count based on traffic, avoiding cold starts that serverless batch or custom containers would introduce.

Why this answer

Vertex AI Endpoints with automatic scaling (option B) are designed for online serving with minimal latency and can automatically adjust the number of replicas based on traffic load, handling spikes without manual intervention. This is the correct choice for a TensorFlow model requiring real-time inference and elastic scaling.

Exam trap

Google Cloud often tests the misconception that any cloud deployment with a load balancer (like Compute Engine) provides automatic scaling, but the trap here is that Vertex AI Endpoints offer managed autoscaling natively, whereas Compute Engine VMs require additional infrastructure setup and do not automatically scale without configuring managed instance groups.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Garden is a repository of pre-built models and foundation models, not a deployment option for serving custom trained models with automatic scaling. Option C is wrong because Vertex AI Batch Prediction is for offline, asynchronous inference on large datasets, not for real-time online serving with low latency. Option D is wrong because deploying to a Compute Engine VM with a load balancer requires manual scaling configuration (e.g., managed instance groups) and lacks the integrated autoscaling, monitoring, and model versioning capabilities of Vertex AI Endpoints.

10
Multi-Selectmedium

A pipeline uses the Google Cloud Pipeline Components to perform AutoML training and batch prediction. Which two components from the GCPC library should they use? (Choose two.)

Select 2 answers
A.CustomJobRunOp
B.DataflowPythonOp
C.AutoMLTabularTrainingJobRunOp
D.BatchPredictOp
E.EndpointPredictOp
AnswersC, D

AutoMLTabularTrainingJobRunOp submits a tabular AutoML training job directly from the pipeline, satisfying the AutoML training requirement. It wraps the Vertex AI training API as a pipeline component, so the pipeline orchestrates training without custom code. Batch prediction is handled separately by a prediction component, making this one of the two required GCPC components.

Why this answer

Option C, AutoMLTabularTrainingJobRunOp, is correct because it is the Google Cloud Pipeline Components (GCPC) operator that wraps the Vertex AI AutoML tabular training job, allowing the pipeline to launch an AutoMLTabularTrainingJob and produce a model artifact. Option D, BatchPredictOp, is correct because it is the GCPC component that submits a Vertex AI batch prediction job against a trained model and a specified input data source, which is exactly the batch prediction step described in the scenario. Option A, CustomJobRunOp, is not appropriate here because it runs a custom training container rather than an AutoML training job.

Option B, DataflowPythonOp, is a Dataflow-based Python execution component and does not perform AutoML training or batch prediction. Option E, EndpointPredictOp, performs online prediction against a deployed Vertex AI endpoint, not batch prediction, so it does not fit the scenario.

Exam trap

PMLE often tests the confusion between online prediction (EndpointPredictOp) and batch prediction (BatchPredictOp), and between custom training (CustomJobRunOp) and AutoML training (AutoMLTabularTrainingJobRunOp) — candidates pick the wrong op for the stated workload.

11
Multi-Selectmedium

A company is deploying a model on Vertex AI for online predictions with strict latency SLOs. The model requires GPU acceleration. Which TWO configurations should they consider to meet the SLOs while optimizing cost?

Select 2 answers
A.Use n1-highmem-32 machine types without GPU
B.Set min_replica_count to handle base traffic and max_replica_count to handle spikes
C.Use GPU-enabled machine types such as n1-standard-4 with T4
D.Enable autoscaling with min_replica_count=0 and max_replica_count=10
E.Disable autoscaling and set a fixed number of replicas equal to peak load
AnswersB, C

Setting min_replica_count to cover baseline traffic keeps GPU-backed replicas warm, eliminating cold-start latency that would breach the SLO, while max_replica_count caps autoscaling during spikes so cost scales only with genuine demand. This directly satisfies the strict latency requirement without over-provisioning GPUs continuously.

Why this answer

Option B is correct because configuring min_replica_count to cover baseline traffic keeps enough warm replicas to serve steady requests within latency SLOs, while max_replica_count allows Vertex AI autoscaling to absorb traffic spikes without over-provisioning permanently. Option C is correct because the model requires GPU acceleration, and GPU-enabled machine types such as n1-standard-4 with an NVIDIA T4 provide the necessary hardware acceleration for online predictions. Option A is incorrect because n1-highmem-32 without a GPU cannot satisfy the GPU acceleration requirement.

Option D is incorrect because setting min_replica_count=0 can cause cold-start delays and scale-from-zero latency that violate strict latency SLOs. Option E is incorrect because disabling autoscaling and fixing replicas at peak load wastes cost and does not optimize for variable traffic.

Exam trap

PMLE often tests the trade-off between cost and latency, and candidates may incorrectly choose min_replica_count=0 to save cost, ignoring the latency impact.

12
MCQhard

A company uses Vertex AI Pipelines with prebuilt components for data processing, training, and deployment. They need to integrate a custom validation step written in Python. What is the correct way to include this as a component?

A.Package the code in a Docker container and reference it as a custom job
B.Define the step in the YAML pipeline definition using arbitrary Python commands
C.Create a custom component using the Vertex AI Pipelines SDK @component decorator
D.Use a Cloud Function as a pipeline step
E.Write a standalone Python script and call it using a Cloud Shell step
AnswerC

The @component decorator converts a Python function into a reusable Vertex AI pipeline component, packaging its dependencies and generating a component specification. This satisfies the need to add custom Python validation logic alongside prebuilt components without building a container manually.

Why this answer

The Vertex AI Pipelines SDK provides a `@component` decorator that allows you to define a custom Python function as a pipeline component. This decorator automatically handles packaging the Python code into a container image, generating the component specification, and integrating it seamlessly with the pipeline orchestration engine. It is the idiomatic and recommended way to add custom validation logic without manually managing Docker or infrastructure.

Exam trap

The trap here is that candidates often confuse the `@component` decorator with a simple function wrapper and assume they can just write inline Python code in the pipeline YAML (Option B), not realizing that Vertex AI Pipelines requires each step to be a containerized component with explicit input/output definitions.

How to eliminate wrong answers

Option A is wrong because packaging code in a Docker container and referencing it as a custom job would create an independent job outside the pipeline DAG, losing the ability to pass inputs/outputs between pipeline steps and breaking the orchestration flow. Option B is wrong because Vertex AI Pipelines YAML definitions do not support arbitrary Python commands; they require prebuilt or custom component definitions with proper container specifications. Option D is wrong because Cloud Functions are event-driven serverless functions not designed for pipeline step integration; they lack native support for pipeline I/O, artifact tracking, and retry logic within Vertex AI Pipelines.

Option E is wrong because Cloud Shell is an interactive environment for ad-hoc commands, not a pipeline execution step; it cannot be used as a component in a Vertex AI Pipeline and would not support parameter passing or artifact management.

13
MCQhard

A team uses Cloud Composer to orchestrate a complex ML pipeline with many tasks. They notice that the DAG parsing time is very high, causing delays in task scheduling. Which action would most effectively reduce DAG parsing time?

A.Remove all DAG files that are not currently needed from the bucket
B.Increase the parallelism of the Airflow scheduler
C.Optimize DAG files to avoid heavy top-level imports and database queries
D.Combine all DAGs into a single file
AnswerC

The scheduler parses every DAG file repeatedly, so heavy top-level imports and database queries executed at parse time inflate parsing duration and delay scheduling. Moving such work inside tasks or using lazy loading cuts parsing time directly.

Why this answer

Heavy top-level imports and database queries in DAG files are executed every time the scheduler parses the DAG, which happens frequently (default every 30 seconds). By moving imports inside Python callables or using lazy loading, the parsing time is drastically reduced, allowing the scheduler to process DAGs faster and trigger tasks without delay.

Exam trap

Google Cloud often tests the misconception that reducing the number of DAG files or increasing scheduler resources will fix parsing delays, when the real bottleneck is the top-level code execution inside each DAG file.

How to eliminate wrong answers

Option A is wrong because removing unused DAG files reduces clutter but does not address the root cause of high parsing time; the scheduler still parses all present DAG files, and if they contain heavy top-level code, parsing remains slow. Option B is wrong because increasing scheduler parallelism (e.g., `scheduler_parallelism` or `max_threads`) only affects how many tasks the scheduler can process concurrently, not how fast it parses DAG files; parsing is a sequential, per-file operation. Option D is wrong because combining all DAGs into a single file actually increases parsing time, as the scheduler must parse one very large file with all dependencies loaded at once, and it also breaks Airflow's ability to detect changes per DAG.

14
MCQmedium

A healthcare company wants to build a model to predict patient readmission risk using structured data in BigQuery. They have a dataset with 100,000 rows and 30 features, including numerical and categorical variables. They require a model that provides explainable predictions and can be trained quickly. They decide to use BigQuery ML. Which model type should they choose?

A.Matrix factorization
B.Logistic regression
C.K-means clustering
D.Deep neural network (DNN)
AnswerB

Logistic regression is a linear model for binary classification that provides interpretable coefficients, making it suitable for explainable predictions. It trains quickly on structured data and handles both numerical and categorical features after preprocessing. For predicting patient readmission (a binary outcome), logistic regression is a strong choice in BigQuery ML, especially when explainability is required. It also supports regularization to prevent overfitting.

Why this answer

Logistic regression is the best choice because it is interpretable, fast to train, and effective for binary classification with structured data. It provides coefficients that indicate feature importance, which helps explain predictions. BigQuery ML supports logistic regression with options for regularization, and it can handle categorical variables automatically.

This aligns with the requirement for explainable predictions and quick training.

Exam trap

The trap here is assuming that a more complex model like a deep neural network is always better, ignoring the need for explainability.

15
MCQeasy

A team has deployed a model to a Vertex AI Endpoint and wants to monitor the model's performance in production. They need to track the number of prediction requests and the average latency. Which Google Cloud service should they use to collect and visualize these metrics?

A.Cloud Trace
B.Cloud Monitoring
C.Vertex AI Model Monitoring
D.Cloud Logging
AnswerB

Cloud Monitoring collects and visualizes metrics from Google Cloud services, including Vertex AI Endpoints. It automatically ingests metrics such as prediction request count and latency. You can create dashboards and alerts based on these metrics. This is the standard service for monitoring operational metrics of deployed models, providing near real-time visibility into endpoint performance.

Why this answer

Cloud Monitoring is the Google Cloud service designed for collecting, visualizing, and alerting on metrics. Vertex AI Endpoints automatically report metrics such as prediction request count and latency to Cloud Monitoring. Teams can use Cloud Monitoring dashboards to track these operational metrics and set up alerts for anomalies.

This provides a centralized view of endpoint health and performance.

Exam trap

The trap here is confusing operational monitoring with model-specific monitoring, leading to the selection of Vertex AI Model Monitoring or Cloud Logging instead of Cloud Monitoring.

16
MCQhard

An organization uses Cloud Dataflow to preprocess training data. Dataflow jobs are often failing because of insufficient quota for certain resources. The team has requested a quota increase, but the jobs still fail with 'quota exceeded' errors for a different resource. They want to proactively monitor and manage quotas to avoid failures. What is the best approach?

A.Set up Cloud Monitoring alerts for quota usage and automate quota increase requests.
B.Configure Dataflow to use a different pipeline type that avoids the quota.
C.Use Dataflow's autoscaling feature to reduce resource usage.
D.Increase the maximum number of workers in the Dataflow job.
AnswerA

Proactive monitoring and automation allow scaling quotas as needed.

Why this answer

The best approach is to set up Cloud Monitoring alerts for quota usage and automate quota increase requests. This provides proactive visibility into all resource quotas (not just the one initially increased) and enables automated remediation before jobs fail. Cloud Monitoring can track quota metrics for services like Compute Engine, and you can use Cloud Functions or Pub/Sub to trigger quota increase requests via the Service Usage API.

Exam trap

PMLE often tests the difference between reactive fixes (increasing workers) and proactive monitoring, and candidates may choose autoscaling or pipeline changes instead of addressing quota management directly.

How to eliminate wrong answers

Option B is wrong because changing the pipeline type does not address the root cause — quota limits apply regardless of pipeline type, and you may still hit them. Option C is wrong because autoscaling reduces resource usage but does not eliminate the need to monitor and manage quotas; it may even mask the problem temporarily. Option D is wrong because increasing the maximum number of workers would consume more quota, potentially worsening the issue, and does not provide proactive monitoring.

17
MCQeasy

A data analyst wants to build a binary classification model to predict customer churn using SQL queries in BigQuery. Which BigQuery ML model type should they use?

A.MATRIX_FACTORIZATION
B.LINEAR_REG
C.LOGISTIC_REG
D.K_MEANS
AnswerC

LOGISTIC_REG performs binary logistic regression inside BigQuery, outputting probabilities for two-class labels such as churn or no churn. It satisfies the stem's requirement for a binary classification model built directly through SQL, with no data export or external tooling needed.

Why this answer

LOGISTIC_REG is the BigQuery ML model type for binary classification, predicting a binary outcome such as churn (yes/no). It uses logistic regression to estimate the probability of the binary outcome. This is the correct choice for predicting customer churn.

Exam trap

PMLE often tests the confusion between regression and classification model types, leading candidates to pick LINEAR_REG for binary outcomes or K_MEANS for supervised tasks.

How to eliminate wrong answers

Option A is wrong because MATRIX_FACTORIZATION is used for recommendation systems and collaborative filtering, not binary classification. Option B is wrong because LINEAR_REG is for linear regression, predicting continuous numeric values, not binary outcomes. Option D is wrong because K_MEANS is for clustering, an unsupervised learning technique, not classification.

18
MCQeasy

Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?

A.It eliminates the need for a load balancer.
B.It reduces costs by scaling down to zero replicas when no requests are received.
C.It reduces model training time.
D.It automatically upgrades the model version.
AnswerB

Scale-to-zero removes all replicas when no requests arrive, so you pay nothing during idle periods while still serving traffic when demand returns. This satisfies the cost-reduction benefit, unlike always-on endpoints that bill for idle capacity.

Why this answer

Vertex AI Endpoints with autoscaling and scale-to-zero allow the number of serving replicas to dynamically adjust based on incoming traffic. When no requests are received, the endpoint can scale down to zero replicas, meaning you are not charged for idle compute resources. This directly reduces operational costs compared to maintaining a minimum number of always-on instances.

Exam trap

A common misconception is that autoscaling eliminates the need for a load balancer, but in Vertex AI Endpoints, the load balancer is a separate component that remains essential for request distribution even when scaling to zero.

How to eliminate wrong answers

Option A is wrong because Vertex AI Endpoints still require a load balancer (the built-in Google Cloud Load Balancer) to distribute incoming requests across replicas; autoscaling does not eliminate this need. Option C is wrong because model training time is a function of training infrastructure and algorithm, not of serving endpoint configuration like autoscaling. Option D is wrong because Vertex AI Endpoints do not automatically upgrade model versions; you must explicitly deploy a new model version or use a traffic split to route requests to a different version.

19
MCQmedium

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to automatically deploy the model to an endpoint only if the evaluation metric (e.g., accuracy) exceeds a threshold. The pipeline is defined using the Kubeflow Pipelines SDK. Which approach should the engineer use to implement this conditional deployment?

A.Set the deployment component's 'condition' parameter to the metric value, and rely on the component to skip itself if the value is below threshold.
B.Configure the pipeline to always deploy, and use a Cloud Function triggered by a Pub/Sub message from the evaluation step to roll back if the metric is below threshold.
C.Use a Vertex AI Model Evaluation component and set its 'deploy' flag to true, which automatically deploys if the metric meets the threshold.
D.Use a Vertex AI Pipelines condition with a comparison of the metric output to the threshold, and place the deployment component inside the condition's 'then' branch.
AnswerD

Vertex AI Pipelines supports dsl.Condition to control execution flow based on pipeline parameters or component outputs. By comparing the evaluation metric to a threshold, the deployment component runs only when the condition is true. This is the standard approach for conditional steps in Kubeflow Pipelines.

Why this answer

Conditional deployment in Vertex AI Pipelines is achieved by using dsl.Condition to evaluate a metric against a threshold. The deployment component is placed inside the condition's true branch, ensuring it only executes when the metric meets the requirement. This provides a clean, orchestrated way to gate deployment without external services or manual intervention.

Exam trap

The trap here is assuming that evaluation components or deployment components have built-in conditional flags, when actually the condition must be explicitly defined in the pipeline graph.

20
MCQmedium

An ML engineer is designing a Vertex AI Pipeline that includes a custom training component. The component must read a dataset from a Cloud Storage bucket and write the trained model to another Cloud Storage location. The engineer wants the component to be reusable across pipelines and to ensure that the pipeline tracks the exact dataset and model artifacts. Which approach should the engineer take?

A.Pass the Cloud Storage URIs as component inputs and outputs, and declare them as artifacts of type Dataset and Model.
B.Hardcode the Cloud Storage URIs inside the component's container code and use environment variables for configuration.
C.Use a single string parameter for the dataset URI and a single string parameter for the model URI, without declaring artifact types.
D.Mount a Cloud Storage bucket as a volume in the component container and write the model directly to the mounted path.
AnswerA

Declaring inputs and outputs as artifacts of type Dataset and Model allows Vertex AI Pipelines to track lineage and metadata automatically. The component remains reusable because the URIs are parameterized. This approach also enables the pipeline to visualize the artifacts and their relationships in the Vertex AI Pipelines UI, and supports artifact-based triggering and caching.

Why this answer

Using artifact inputs and outputs with types Dataset and Model ensures that Vertex AI Pipelines tracks lineage and metadata. It keeps the component reusable because the actual URIs are passed at runtime. Declaring artifacts also enables the pipeline to display them in the UI and to use them for caching and conditional execution.

Exam trap

The trap here is assuming that passing URIs as plain string parameters is sufficient for artifact tracking, when only declared artifacts provide lineage and metadata.

21
MCQmedium

Your Vertex AI custom training job is failing with an out-of-memory error on a single GPU. You need to reduce memory usage without changing the model architecture. Which approach should you try first?

A.Decrease the batch size
B.Implement model parallelism across GPUs
C.Use gradient accumulation
D.Enable mixed precision training (FP16)
AnswerA

Reducing the batch size lowers the number of samples held in GPU memory per step, directly cutting activation and gradient memory consumption. It requires no architecture change, satisfying the stem's constraint, and is the least invasive first remedy before considering gradient checkpointing or mixed precision.

Why this answer

Decreasing the batch size directly reduces the memory required for activations and gradients, making it the simplest and most immediate way to resolve out-of-memory errors without altering the model architecture. It is a standard first step because it requires no code changes beyond a hyperparameter and often resolves OOM with minimal impact on convergence if adjusted with learning rate.

Exam trap

PMLE often tests the order of operations for troubleshooting OOM, trapping candidates who jump to advanced techniques like mixed precision or model parallelism instead of the simplest, most direct fix: reducing batch size.

How to eliminate wrong answers

Option B is wrong because model parallelism across GPUs is complex, requires significant code changes, and is not the first thing to try for a single-GPU OOM. Option C is wrong because gradient accumulation increases memory usage by storing gradients for multiple steps, which would worsen OOM. Option D is wrong because mixed precision training (FP16) can reduce memory, but it requires code changes and may introduce numerical instability; it is not the first, simplest step.

22
MCQmedium

A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?

A.Update the endpoint's deployed model to set a lower autoscaling metric threshold, such as 70% CPU utilization.
B.Enable request-response logging on the endpoint and use Cloud Monitoring alerts to manually add replicas when CPU exceeds 70%.
C.Increase the endpoint's minReplicaCount to match the peak traffic level so that replicas are always available.
D.Redeploy the model with a larger machine type so that each replica handles more traffic and CPU never reaches the threshold.
AnswerA

Vertex AI Endpoints expose an autoscaling configuration on the DeployedModel, including the target metric and threshold. Lowering the CPU utilization target to 70% makes the autoscaler add replicas sooner, which reduces latency during ramp-up. This is a configuration change on the deployed model and does not require retraining or replacing the model artifact.

Why this answer

The autoscaler on a Vertex AI Endpoint uses a target metric and threshold defined in the deployed model's autoscaling configuration. Adjusting the CPU utilization target to a lower value makes the system add replicas before latency degrades. Other approaches either do not change the scaling trigger or introduce manual processes that are slower and less reliable.

Exam trap

The trap here is assuming that increasing the minimum replica count changes when the autoscaler scales, when in fact it only raises the floor of running replicas.

23
MCQmedium

You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?

A.Write a Cloud Composer workflow that runs the preprocessing and then triggers the batch prediction job.
B.Use Dataflow to read from both BigQuery tables, perform the join and preprocessing, write the results to GCS, then run Vertex AI batch prediction with GCS source.
C.Use Vertex AI batch prediction with a custom container that includes logic to read and join tables on the fly.
D.Use BigQuery to create a materialized view that joins the tables and directly use that as the batch prediction source.
AnswerB

Dataflow performs the complex two-table BigQuery join and preprocessing, writing results to Cloud Storage, which Vertex AI batch prediction then reads as its source. This satisfies the constraint that preprocessing must complete before inference, since batch prediction cannot join BigQuery tables itself.

Why this answer

Dataflow (Apache Beam) is designed for complex, stateful data processing like joining two BigQuery tables and performing custom preprocessing. It can read from BigQuery, execute the join logic, and write the preprocessed results to Cloud Storage (GCS). Vertex AI batch prediction then reads the preprocessed data from GCS, which is the recommended pattern for non-trivial transformations before inference, as it decouples preprocessing from prediction and avoids resource contention.

Exam trap

Google often tests the misconception that batch prediction can handle live data transformations within the prediction container, but the correct design is to preprocess data in a separate, scalable data processing service like Dataflow before feeding it to batch prediction.

How to eliminate wrong answers

Option A is wrong because Cloud Composer (Apache Airflow) is an orchestration tool, not a data processing engine; using it to run the preprocessing itself would be inefficient and error-prone, as it lacks native support for large-scale data joins and transformations. Option C is wrong because Vertex AI batch prediction with a custom container that reads and joins tables on the fly violates the principle of separation of concerns, leading to longer inference latency, higher memory usage, and potential timeouts during prediction, as batch prediction expects preprocessed input, not live database joins. Option D is wrong because BigQuery materialized views are precomputed, read-only snapshots that cannot be used directly as a batch prediction source; batch prediction requires input data in GCS (JSON/CSV) or BigQuery tables, but a materialized view is not a table and cannot be referenced as a source URI.

24
MCQmedium

A company wants to analyze customer reviews for sentiment (positive, negative, neutral) using a pre-trained model with no training. They have text data stored in BigQuery. Which Google Cloud service should they use?

A.Translation API
B.Speech-to-Text API
C.Natural Language API
D.AutoML NLP
AnswerC

The Natural Language API provides pre-trained sentiment analysis, returning positive, negative and neutral scores without any custom training. It integrates directly with BigQuery data, satisfying the requirement for a pre-trained model applied to stored review text.

Why this answer

The Natural Language API is a pre-trained service that provides sentiment analysis out of the box, requiring no training. It can directly process text data from BigQuery (via integration or export) to classify sentiment as positive, negative, or neutral. This matches the requirement of using a pre-trained model with no training.

Exam trap

PMLE often tests the distinction between pre-trained APIs and custom training services; candidates may mistakenly choose AutoML NLP thinking it's needed for sentiment, but AutoML requires training data.

How to eliminate wrong answers

Option A is wrong because the Translation API translates text between languages, not analyze sentiment. Option B is wrong because Speech-to-Text transcribes audio, not text sentiment. Option D is wrong because AutoML NLP requires custom training with labeled data, which contradicts the 'no training' requirement.

25
MCQeasy

An MLOps team has deployed a model on Vertex AI Endpoints and wants to monitor for skew between training and serving data distributions. Which Vertex AI service should they use?

A.Vertex AI Explainability
B.Vertex AI Model Monitoring
C.Vertex AI Model Registry
D.Vertex AI Continuous Training
AnswerB

Vertex AI Model Monitoring detects training-serving skew and drift by comparing incoming prediction request distributions against a baseline schema and statistics. It satisfies the requirement to monitor distributional skew on deployed Endpoints without custom tooling.

Why this answer

Vertex AI Model Monitoring is the service designed to monitor for skew and drift between training and serving data distributions. It can detect anomalies in feature distributions and alert when skew or drift exceeds thresholds.

Exam trap

PMLE often tests the distinction between Model Monitoring and other Vertex AI services, and candidates may confuse monitoring with explainability or model registry.

How to eliminate wrong answers

Option A is wrong because Vertex AI Explainability provides feature attributions for predictions, not monitoring of data distributions. Option C is wrong because Vertex AI Model Registry is for managing model versions and metadata, not monitoring. Option D is wrong because Vertex AI Continuous Training is for automating retraining pipelines, not for monitoring skew.

26
MCQeasy

Which Vertex AI service is best suited for finding similar items in a large dataset based on embedding vectors, such as product recommendations or image similarity search?

A.Vertex AI Prediction Endpoint
B.Vertex AI Model Monitoring
C.Vertex AI Feature Store
D.Vertex AI Matching Engine
AnswerD

Matching Engine performs approximate nearest-neighbour search over embedding vectors, returning semantically similar items at scale. It satisfies the stem's similarity-search requirement for recommendations and image search, unlike tabular or forecasting services that do not index vector embeddings.

Why this answer

Vertex AI Matching Engine is specifically designed for high-performance vector similarity search (also known as approximate nearest neighbor search) using embedding vectors. It scales to billions of vectors and is ideal for use cases like product recommendations and image similarity search, where you need to find the most similar items based on dense vector representations.

Exam trap

Candidates often confuse Vertex AI Prediction (model serving) with Vertex AI Matching Engine (vector similarity search). The key distinction is that Prediction serves model inference on input data, while Matching Engine retrieves similar items based on embedding vectors.

How to eliminate wrong answers

Option A is wrong because Vertex AI Prediction Endpoint serves model predictions via HTTP requests but does not provide built-in vector similarity search or indexing capabilities. Option B is wrong because Vertex AI Model Monitoring tracks prediction quality and data drift over time, not similarity search. Option C is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature data, but it does not perform nearest neighbor search on embedding vectors.

27
Multi-Selecthard

You are preparing to train a large image classification model on Vertex AI using a custom training job. You want to optimize the training job for cost and performance. The dataset is stored in Cloud Storage as TFRecords and is about 2 TB. You plan to use a machine with 4 NVIDIA V100 GPUs. Which two actions should you take to improve training efficiency? (Choose two.)

Select 2 answers
A.Use a smaller model architecture to reduce training time.
B.Store the TFRecords in a regional Cloud Storage bucket instead of a multi-regional bucket.
C.Use a larger batch size and scale the learning rate accordingly.
D.Enable mixed precision training using NVIDIA Apex or TensorFlow's mixed precision API.
E.Increase the number of vCPUs on the machine to improve data loading throughput.
AnswersC, D

Increasing batch size can improve GPU utilization and reduce training time by processing more samples per iteration. Scaling the learning rate helps maintain convergence. This is a standard technique for efficient training on multiple GPUs. It directly leverages the available GPU memory and compute, leading to better throughput and cost efficiency.

Why this answer

Mixed precision training and increasing batch size with learning rate scaling are both effective ways to improve training efficiency on V100 GPUs. Mixed precision leverages tensor cores for faster computation and reduced memory, while larger batches improve GPU utilization. These actions directly target performance and cost without altering the model architecture or data storage.

Exam trap

The trap here is focusing on infrastructure changes like bucket location or more vCPUs, when the most impactful optimizations are at the training algorithm level, such as precision and batch size.

28
MCQeasy

A team wants to track the lineage of ML pipeline runs, including which datasets, parameters, and models were used in each execution. Which Vertex AI service should they use?

A.Vertex AI Metadata
B.Vertex AI Feature Store
C.Vertex AI Model Registry
D.Vertex AI Experiments
AnswerA

Vertex AI Metadata stores artefacts and executions in a lineage graph, recording which datasets, parameters and models each pipeline run consumed and produced. This directly satisfies the requirement to track lineage across ML pipeline executions.

Why this answer

Vertex AI Metadata (part of Vertex ML Metadata) is the service designed to record and query ML metadata, including artifacts (datasets, models), executions (pipeline runs), and events, forming a lineage graph. It captures which datasets, parameters, and models were used in each pipeline execution, enabling reproducibility and auditability. This directly matches the requirement to track lineage of ML pipeline runs.

Exam trap

PMLE often tests the confusion between Vertex AI Experiments (run tracking and comparison) and Vertex ML Metadata (lineage and artifact relationships), causing candidates to pick Experiments when lineage is the requirement.

How to eliminate wrong answers

Option B (Vertex AI Feature Store) is wrong because it manages and serves feature values for training and online serving, not pipeline run lineage or artifact tracking. Option C (Vertex AI Model Registry) is wrong because it stores and versions trained models and their metadata, but it does not track the full pipeline execution lineage including datasets and parameters. Option D (Vertex AI Experiments) is wrong because it tracks experiment runs, metrics, and parameters for comparison, but it is not the underlying lineage/metadata service that records artifact relationships across pipeline executions.

29
MCQeasy

A team needs to quickly create a visual interface for data exploration and model building without writing code. They want to run AutoML jobs and visualize results. Which Google Cloud tool should they use?

A.Vertex AI Workbench
B.Cloud Datalab
C.Cloud Composer
D.Google Colab
AnswerA

Provides a managed notebook environment with visual data exploration and one-click AutoML integration.

Why this answer

Vertex AI Workbench provides a managed JupyterLab environment with a low-code interface for data exploration, AutoML model training, and result visualization without writing code. It integrates directly with Vertex AI's AutoML and custom training services, allowing users to run AutoML jobs and view evaluation metrics, feature importance, and predictions through its UI.

Exam trap

Google Cloud often tests the distinction between code-based notebook tools (Colab, Datalab) and managed low-code platforms (Vertex AI Workbench), expecting candidates to recognize that AutoML job execution and visual result exploration require the latter's integrated UI and API access.

How to eliminate wrong answers

Option B (Cloud Datalab) is wrong because it is a deprecated tool that required code-based notebooks and does not support AutoML job execution or low-code visual interfaces. Option C (Cloud Composer) is wrong because it is a workflow orchestration service based on Apache Airflow, designed for scheduling and monitoring pipelines, not for interactive data exploration or AutoML. Option D (Google Colab) is wrong because it is a free, code-centric notebook environment that lacks native integration with Vertex AI AutoML and does not provide a low-code visual interface for model building.

30
Multi-Selecteasy

A data scientist is creating a Vertex AI pipeline using the Kubeflow Pipelines SDK v2. Which TWO statements about pipeline parameters are correct? (Choose two.)

Select 2 answers
A.Pipeline parameters are defined as inputs to the pipeline function decorated with @dsl.pipeline.
B.Pipeline parameters must be serialized to JSON before use.
C.Pipeline parameters can only be of type str.
D.Pipeline parameters can be overridden at pipeline run time.
E.Pipeline parameters can be used to pass large datasets between components.
AnswersA, D

Declaring parameters as typed arguments on the @dsl.pipeline-decorated function is how KFP v2 compiles them into pipeline-level inputs, letting callers pass values per run. This satisfies the requirement that parameters be defined at the pipeline level rather than hard-coded inside individual components.

Why this answer

Option A is correct because in the Kubeflow Pipelines SDK v2, pipeline parameters are declared as typed arguments of the function decorated with @dsl.pipeline, which defines the pipeline's input interface. Option D is correct because these parameters are runtime inputs: when submitting a run (e.g., via the Vertex AI Pipelines API or the SDK's create_run_from_pipeline_func), you can supply new values that override the defaults specified in the pipeline definition. Option B is wrong because KFP v2 handles parameter serialization automatically; you do not manually JSON-serialize parameters before use.

Option C is wrong because KFP v2 parameters support multiple types such as int, float, bool, str, list, and dict, not only str. Option E is wrong because parameters are intended for small scalar/structured configuration values; large datasets should be passed between components as artifacts (e.g., Dataset, Model), not as parameters.

Exam trap

The trap is that candidates often assume pipeline parameters must be JSON-serialized or limited to strings due to older Kubeflow v1 conventions, but Vertex AI's Kubeflow Pipelines SDK v2 natively supports multiple Python types and automatic serialization.

31
Multi-Selectmedium

An ML team uses Delta Lake on Dataproc for data versioning. Which THREE benefits does Delta Lake provide?

Select 3 answers
A.Automatic data encryption at rest
B.Time travel for accessing previous versions
C.Schema enforcement and evolution
D.ACID transactions on data lakes
E.Built-in real-time streaming
AnswersB, C, D

Delta Lake's transaction log records every write as a numbered commit, so querying a prior version by timestamp or version number retrieves the exact historical snapshot. This directly satisfies the team's data versioning requirement on Dataproc, enabling reproducible ML training and rollback without duplicating datasets.

Why this answer

Delta Lake provides time travel (B), allowing queries against earlier snapshots of a table via version numbers or timestamps, which directly supports the team's data versioning requirement. It also delivers schema enforcement and evolution (C), rejecting writes that violate the table schema while permitting controlled schema changes such as adding columns. Additionally, Delta Lake brings ACID transactions (D) to data lakes, ensuring atomic, consistent, isolated, and durable writes on top of object storage like GCS.

Options A and E are not Delta Lake benefits: encryption at rest is handled by the underlying storage service (e.g., Google Cloud Storage), not Delta Lake itself, and Delta Lake is a storage/transaction layer rather than a built-in real-time streaming engine.

32
MCQeasy

You are a machine learning engineer working on a team that uses Vertex AI Feature Store. A colleague has created a new feature and wants to make it available to other teams for training and serving. You need to ensure that the feature can be discovered and reused across projects. What should you do?

A.Use Vertex AI Feature Store's built-in feature sharing capability by setting the feature's visibility to 'Public' within the organization.
B.Create the feature in Vertex AI Feature Store and use Dataplex to catalog it, making it searchable and accessible to other teams.
C.Create the feature in a Vertex AI Feature Store instance and share the instance's project ID with other teams.
D.Register the feature in Vertex AI Feature Store and assign it a label that other teams can search for in the Vertex AI console.
AnswerB

Dataplex provides a centralized data catalog that can include Vertex AI Feature Store features. By cataloging the feature, other teams can discover it through search, understand its metadata, and request access. This promotes reuse and collaboration across projects while maintaining proper governance.

Why this answer

Using Dataplex to catalog Vertex AI Feature Store features enables centralized discovery and metadata management. Other teams can search the catalog, find the feature, and request access, which fosters reuse and collaboration. This approach integrates with existing governance and access controls, unlike ad-hoc sharing methods.

Exam trap

The trap here is believing that Vertex AI Feature Store has a native public sharing or cross-project visibility toggle, when in fact you need an external catalog like Dataplex.

33
MCQhard

A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?

A.Brute-force index
B.Hash-based index
C.Tree-based index
D.Approximate nearest neighbor (ANN) index with ScaNN
AnswerD

ScaNN is Google's ANN algorithm optimised for high-dimensional vectors, using anisotropic vector quantisation to deliver sub-10ms latency and high throughput. It satisfies both stated constraints on 10 million embeddings, whereas exact or tree-based indexes cannot meet that latency at this scale.

Why this answer

For large datasets requiring low latency, an approximate nearest neighbor (ANN) index is appropriate. The Scann algorithm (ScaNN) is used by Vertex AI Vector Search for ANN.

34
MCQeasy

A data scientist finishes training a model in a Vertex AI Workbench notebook and wants to save it so that the deployment team can later deploy it to an endpoint without re-running the notebook. The deployment team needs to see the model's version history and assign a production alias. Which action should the data scientist take?

A.Upload the model to a Vertex AI Feature Store entity type so the deployment team can retrieve it.
B.Create a Vertex AI Pipeline that retrains the model and send the pipeline template to the deployment team.
C.Register the model in Vertex AI Model Registry and share the model resource name with the deployment team.
D.Export the model artifacts to a Cloud Storage bucket and share the bucket path with the deployment team.
AnswerC

Vertex AI Model Registry stores model versions with metadata and supports aliases, so the deployment team can review history and assign a production alias before deploying to an endpoint. It is the intended handoff mechanism between experimentation and serving, satisfying both version tracking and alias requirements.

Why this answer

The handoff from experimentation to deployment is standardized through Vertex AI Model Registry, which versions models and supports aliases. Sharing raw artifacts or retraining pipelines does not provide the deployment team with a governed, aliasable model resource, and Feature Store is unrelated to model storage.

Exam trap

The trap here is assuming that a Cloud Storage path is an acceptable substitute for a registered model resource.

35
MCQmedium

A data science team is using AI Platform for training. They want to track hyperparameters and metrics across multiple experiments. What should they use?

A.Cloud Logging with custom metrics
B.Vertex AI Experiments
C.Store metrics in Cloud Storage and compare manually
D.Cloud Monitoring dashboards
AnswerB

Vertex AI Experiments records parameters, metrics and artefacts across training runs, letting the team compare and track experiments systematically. This directly satisfies the requirement to monitor hyperparameters and metrics across multiple experiments, integrating with Vertex AI Training and TensorBoard.

Why this answer

Vertex AI Experiments is the correct choice because it is the native service within Vertex AI designed specifically for tracking, comparing, and analyzing hyperparameters and metrics across multiple training runs. It provides a centralized UI and SDK to log parameters, metrics, and artifacts, enabling systematic experiment management without manual effort or external tools.

Exam trap

Google Cloud often tests the distinction between logging/monitoring services (Cloud Logging, Cloud Monitoring) and ML-specific experiment tracking (Vertex AI Experiments), leading candidates to pick a generic monitoring tool instead of the purpose-built ML service.

How to eliminate wrong answers

Option A is wrong because Cloud Logging is intended for collecting and querying log data (e.g., application logs, error messages), not for structured tracking of hyperparameters and metrics across experiments; it lacks built-in experiment comparison features. Option C is wrong because storing metrics in Cloud Storage and comparing manually is inefficient, error-prone, and does not provide automated tracking, visualization, or versioning of experiments, which is the core requirement. Option D is wrong because Cloud Monitoring dashboards are designed for monitoring infrastructure and application performance metrics (e.g., CPU usage, latency), not for tracking ML experiment hyperparameters and metrics across multiple runs.

36
MCQeasy

A startup wants to build a product recommendation engine without writing custom training code. They have user-item interaction data stored in BigQuery. Which Google Cloud service should they use?

A.Cloud Dataflow with ML APIs
B.BigQuery ML matrix factorization
C.Vertex AI AutoML Tables
D.Vertex AI Matching Engine
AnswerB

BigQuery ML matrix factorization trains collaborative-filtering models directly on user-item interaction tables using CREATE MODEL, requiring only SQL. No custom training code or data export is needed, matching the startup's constraint of building recommendations in place.

Why this answer

BigQuery ML matrix factorization is the correct choice because it allows building a recommendation engine directly in BigQuery using SQL, without writing custom training code. It supports implicit and explicit user-item interaction data and provides built-in evaluation metrics, making it ideal for low-code ML solutions on existing BigQuery data.

Exam trap

Google Cloud often tests the distinction between services that require custom code (Dataflow) versus those that offer SQL-based low-code ML (BigQuery ML), and the trap here is assuming any ML service like AutoML or Matching Engine is suitable for recommendation without recognizing the specific need for matrix factorization on interaction data.

How to eliminate wrong answers

Option A is wrong because Cloud Dataflow is a data processing pipeline service, not a low-code ML training service; using ML APIs would require custom code to orchestrate and train models. Option C is wrong because Vertex AI AutoML Tables is designed for tabular data with structured features, not specifically for user-item interaction matrices, and requires exporting data from BigQuery. Option D is wrong because Vertex AI Matching Engine is for vector similarity search and nearest neighbor retrieval, not for training matrix factorization models from interaction data.

37
MCQmedium

An ML engineer is using Cloud Build to trigger a Vertex AI Pipeline on every commit to a repository. The pipeline takes 2 hours. The engineer wants to only run the pipeline when changes are made to specific directories. How can this be achieved?

A.Use Cloud Composer to poll the repository periodically
B.Configure Cloud Build trigger with included file globs
C.Use a Cloud Function to evaluate changes and invoke the pipeline
D.Modify the pipeline to ignore unrelated changes
E.Add a conditional step in the pipeline to abort if no relevant changes
AnswerB

Cloud Build triggers support included file globs, which restrict invocation to commits touching specified directories. This satisfies the constraint of running the two-hour Vertex AI Pipeline only when changes occur in particular paths, avoiding unnecessary executions.

Why this answer

Cloud Build triggers support 'included file globs' and 'ignored file globs' to filter which file changes should invoke the trigger. By specifying glob patterns for the directories of interest, the trigger will only fire when commits modify files matching those patterns, avoiding unnecessary pipeline runs for unrelated changes.

Exam trap

The trap here is that candidates may think a pipeline-level conditional check (Option E) is sufficient, but they overlook that Cloud Build triggers can filter at the trigger level, avoiding any pipeline startup cost for irrelevant changes.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is an orchestration service for workflows, not a polling mechanism for repository changes; it would add unnecessary complexity and latency. Option C is wrong because using a Cloud Function to evaluate changes and invoke the pipeline is an overengineered solution; Cloud Build triggers natively support file glob filtering without needing an intermediary. Option D is wrong because modifying the pipeline to ignore unrelated changes would still consume resources to start the pipeline and then abort, wasting time and cost.

Option E is wrong because adding a conditional step in the pipeline to abort if no relevant changes still requires the pipeline to start and run until the conditional check, incurring unnecessary execution time and cost.

38
Multi-Selecteasy

An ML engineer wants to monitor the performance of a Vertex AI Endpoint. Which TWO metrics are available in Cloud Monitoring for Vertex AI Endpoints? (Choose 2)

Select 2 answers
A.Model accuracy
B.Error count
C.Feature skew score
D.SHAP values
E.Prediction latency (p50, p95, p99)
AnswersB, E

Error count is exported automatically to Cloud Monitoring for every Vertex AI Endpoint, aggregated per deployed model and response code. It satisfies the stem's monitoring requirement by surfacing failed prediction requests, letting the engineer alert on serving errors without configuring custom logging or additional instrumentation.

Why this answer

Option B (Error count) is correct because Vertex AI Endpoints automatically publish request-level metrics to Cloud Monitoring, including aiplatform.googleapis.com/endpoint/error_count, which tracks failed prediction requests and is essential for detecting serving problems. Option E (Prediction latency (p50, p95, p99)) is correct because Vertex AI Endpoints expose latency metrics such as aiplatform.googleapis.com/endpoint/prediction_latencies, which Cloud Monitoring reports as percentiles (p50, p95, p99) to characterize response-time distribution. Option A (Model accuracy) is not a built-in Cloud Monitoring metric for Endpoints; accuracy must be computed separately, for example via Vertex AI Model Monitoring or custom evaluation jobs.

Option C (Feature skew score) belongs to Vertex AI Model Monitoring's skew/drift detection outputs, not to the standard Endpoint metrics in Cloud Monitoring. Option D (SHAP values) are explainability artifacts produced by Vertex Explainable AI, not time-series metrics available for an Endpoint in Cloud Monitoring.

Exam trap

The trap here is confusing model-quality metrics (accuracy, skew, SHAP) with operational serving metrics (error count, latency); candidates often assume that because Vertex AI offers Model Monitoring, those quality metrics are automatically available in Cloud Monitoring for every endpoint.

39
Multi-Selecthard

A team is using Vertex AI Pipelines to orchestrate a training workflow. They want to ensure that the pipeline can be reproduced exactly six months later for auditing purposes. They need to capture all necessary information to rerun the pipeline and obtain identical results. (Choose two.)

Select 2 answers
A.Rely on the default caching behavior of Vertex AI Pipelines to reuse previous step outputs.
B.Use the same pipeline name and run it again with the same parameters.
C.Use a fixed pipeline template with pinned component versions and container image digests.
D.Store the pipeline parameters and input artifacts in a versioned location, such as a Cloud Storage bucket with versioning enabled.
E.Enable Vertex AI Experiments to log metrics and parameters for each run.
AnswersC, D

Pinning component versions and container image digests ensures that the exact same code and dependencies are used when the pipeline is rerun. This is critical for reproducibility because container images can be updated, and using digests guarantees immutability. Pinned versions also prevent unexpected changes in component behavior.

Why this answer

To reproduce a pipeline run exactly, you must pin the pipeline template and container images to immutable versions, and store the exact parameters and input artifacts in a versioned location. These two practices ensure that both the code and data are preserved, allowing an identical rerun. Other options like caching or experiment tracking do not provide the necessary immutability.

Exam trap

The trap here is assuming that caching or experiment tracking alone provides reproducibility, but they only log or reuse outputs without guaranteeing the exact code and data are preserved.

40
Multi-Selectmedium

You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)

Select 2 answers
A.Set minReplicaCount to 0 to allow scale-to-zero.
B.Use a custom container that loads the model during startup.
C.Increase maxReplicaCount to a high number.
D.Use batch prediction instead of online prediction.
E.Set minReplicaCount to 1 to always have at least one replica running.
AnswersB, E

Loading the model during container startup means weights are already in memory when the first request arrives, rather than fetched lazily per request. This directly reduces cold-start latency for the GPU-backed deep learning model described in the stem, complementing replica-warming measures.

Why this answer

Option B is correct because using a custom container that loads the model during startup lets you control and optimize the initialization process (for example, preloading weights and warming up the GPU) so the model is ready as soon as the container starts, which directly reduces cold-start latency. Option E is correct because setting minReplicaCount to 1 keeps at least one replica always provisioned and running on Vertex AI, so incoming requests hit an already-loaded model instead of triggering a new replica spin-up, eliminating the cold start entirely. Option A is wrong because setting minReplicaCount to 0 enables scale-to-zero, meaning replicas are torn down when idle and every new request incurs a full cold start, which is the opposite of the goal.

Option C is wrong because increasing maxReplicaCount only raises the ceiling for horizontal scaling under load and does nothing to reduce the latency of starting an individual replica. Option D is wrong because batch prediction is an asynchronous, job-based mode that does not serve real-time online requests and therefore is irrelevant to minimizing endpoint cold-start latency.

Exam trap

Google often tests the misconception that scale-to-zero (minReplicaCount=0) reduces latency, when in fact it increases cold-start latency; the correct approach is to keep at least one replica always warm (minReplicaCount=1) and pre-load the model during container startup.

41
MCQeasy

You have trained a scikit-learn model and want to deploy it to Vertex AI for online predictions. You need to minimize the effort to create a custom container and ensure the model is served with the default pre-built container. What should you do?

A.Package the model into a custom container with a Flask app and deploy it to Vertex AI.
B.Export the model as a PMML file and deploy using the pre-built XGBoost container.
C.Convert the model to TensorFlow SavedModel format and deploy using the pre-built TensorFlow container.
D.Save the model using joblib and upload it to Vertex AI Model Registry, then deploy using the pre-built scikit-learn container.
AnswerD

Vertex AI provides pre-built containers for scikit-learn that expect the model to be saved in a specific format, typically joblib or pickle. By saving the model with joblib and uploading it to Model Registry, you can deploy it directly using the pre-built container without building a custom image. This minimizes effort and leverages the managed serving stack.

Why this answer

The least-effort path is to use the pre-built scikit-learn container on Vertex AI. Saving the model with joblib and uploading it to Model Registry allows you to deploy without writing any serving code or building a container. The pre-built container handles loading the model and serving predictions.

Other options involve unnecessary conversion or custom packaging, which increase effort and risk.

Exam trap

The trap here is overcomplicating the deployment by assuming a custom container is needed, when Vertex AI provides a pre-built container for scikit-learn models.

42
Multi-Selectmedium

Your organization wants to automate the retraining of a model when new data is available and also on a weekly schedule. Which TWO services would you use together to achieve this? (Choose two.)

Select 2 answers
A.Cloud Functions
B.Cloud Composer
C.Cloud Tasks
D.Dataflow
E.Cloud Scheduler
AnswersA, E

Cloud Functions provides the event-driven compute: a trigger fires when new data lands, invoking code that submits the retraining pipeline. This satisfies the stem's requirement to automate retraining on data arrival, complementing a scheduler for the weekly cadence.

Why this answer

Cloud Scheduler (E) is the correct service for the time-based trigger, since it can invoke a target on a cron schedule such as weekly, which satisfies the 'weekly schedule' requirement. Cloud Functions (A) is correct because it provides the serverless compute that actually executes the retraining logic, and it can be triggered both by Cloud Scheduler for the weekly run and by an event (for example, a Cloud Storage object finalize event when new data lands) for the data-availability requirement. Together, Cloud Scheduler fires on the cron schedule and Cloud Functions runs the retraining code, covering both triggers.

Cloud Composer (B) is a managed Airflow workflow orchestrator and could schedule DAGs, but it is heavier than needed and is not the marked pairing for this scenario. Cloud Tasks (C) is a queue for asynchronous task dispatch, not a cron scheduler or compute runtime, and Dataflow (D) is a managed Apache Beam service for data processing pipelines, not for triggering or hosting model-retraining logic.

Exam trap

Google often tests the distinction between orchestration (Cloud Composer) and simple scheduling/event-driven triggers (Cloud Scheduler + Cloud Functions), leading candidates to over-engineer the solution by choosing Cloud Composer when a lightweight combination suffices.

43
MCQhard

An ML engineer manages a Vertex AI Endpoint serving a recommendation model. The team wants to detect when the distribution of a specific numerical feature, average session duration, shifts significantly from its training distribution. They have configured Vertex AI Model Monitoring with a training dataset baseline and a monitoring frequency of one hour. After a week, no drift alerts have fired even though the feature's daily mean has visibly moved. What is the most likely cause?

A.The monitoring frequency of one hour is too infrequent to detect daily mean shifts.
B.Drift alerts require at least 30 days of live data before they can trigger.
C.The training dataset baseline was overwritten by the latest live data automatically.
D.The feature was not included in the monitoring configuration's feature list, so it is not being analyzed.
AnswerD

Vertex AI Model Monitoring only analyzes features explicitly listed in the monitoring configuration. If average session duration was omitted from the monitored feature list, the service will not compute drift for it regardless of the baseline or frequency. This is the most likely reason no alerts fired despite an obvious shift, because unlisted features are silently ignored.

Why this answer

Vertex AI Model Monitoring only evaluates features that are explicitly included in the monitoring configuration. If average session duration was left out of the monitored feature list, drift for that feature is never computed, so no alert can fire even when the mean shifts. The other options either describe non-existent behavior or misattribute the cause to frequency or baseline updates.

Exam trap

The trap here is focusing on frequency or baseline mechanics while overlooking that unlisted features are simply not monitored.

44
MCQhard

Your company uses a custom container for model serving on Vertex AI. After a recent update, the model returns predictions but they are clearly wrong (e.g., negative probabilities for a classification model). The logs show no errors. What is the most likely cause?

A.The preprocessing code in the container was updated but the model was not retrained on the new preprocessing
B.The model file is corrupted
C.The model file was accidentally replaced with a different model
D.The container is using an incompatible version of the serving framework
AnswerA

A preprocessing change alters the feature transformation applied before inference, so the model receives inputs on a different scale or encoding than it was trained on. Predictions remain numerically valid but semantically wrong, and no runtime error is raised because the container executes successfully.

Why this answer

The most likely cause of a model returning predictions without errors, but with clearly wrong outputs like negative probabilities, is a mismatch between the preprocessing logic used during training and inference. If the preprocessing code in the container was updated (e.g., scaling, normalization, or feature engineering steps changed) but the model was not retrained on data processed with that new logic, the model receives inputs that are out of distribution, leading to nonsensical outputs. Vertex AI containers run inference with the deployed code, so any change in preprocessing directly affects the input tensor values without raising runtime errors.

Exam trap

Google Cloud often tests the concept that silent prediction errors (no logs, no crashes) are almost always due to data or preprocessing mismatches, not infrastructure or model file issues, which would generate explicit errors.

How to eliminate wrong answers

Option B is wrong because a corrupted model file would typically cause loading failures, runtime errors, or crashes, not silent generation of plausible but wrong predictions like negative probabilities. Option C is wrong because replacing the model file with a different model would likely produce predictions that are consistently wrong in a different pattern (e.g., all zeros, constant values) or cause shape mismatches, not specifically negative probabilities from a classification model. Option D is wrong because an incompatible serving framework version would usually manifest as import errors, missing symbols, or version mismatch warnings in logs, not silent incorrect predictions with no errors.

45
Multi-Selecteasy

A company wants to transcribe audio from customer service calls and then analyze the sentiment of the transcribed text. Which TWO Google Cloud services should they use?

Select 2 answers
A.Natural Language API
B.Document AI
C.Speech-to-Text
D.Translation API
E.Vision API
AnswersA, C

The Natural Language API performs sentiment analysis on text, satisfying the requirement to analyse the transcribed call content. Speech-to-Text handles the audio transcription; Natural Language API then classifies each transcript's sentiment. Together they cover both stages of the pipeline, making this one of the two required services.

Why this answer

Speech-to-Text (option C) is correct because it converts audio from the customer service calls into written text, which is the required transcription step. Natural Language API (option A) is correct because it performs sentiment analysis on text, allowing the company to analyze the sentiment of the transcribed call content. Document AI (option B) is not appropriate here because it processes documents and forms rather than audio or general sentiment analysis.

Translation API (option D) only translates text between languages and does not transcribe audio or analyze sentiment. Vision API (option E) analyzes images, so it cannot handle audio transcription or text sentiment analysis.

Exam trap

PMLE often tests the combination of services for multi-step tasks; candidates might choose Document AI for transcription or Translation API for sentiment, but these are incorrect service mappings.

46
MCQeasy

A company wants to automatically retrain their model every night at 2 AM using Vertex AI Pipelines. Which approach should they use to trigger the pipeline on a schedule?

A.Use Cloud Scheduler to call the Vertex AI pipeline creation API
B.Deploy the pipeline as a Cloud Run job with a cron trigger
C.Use Vertex AI Experiments to schedule runs
D.Configure a cron job inside the pipeline definition
AnswerA

Cloud Scheduler provides cron-based triggering, satisfying the nightly 2 AM requirement. It invokes the Vertex AI Pipelines API endpoint directly, which compiles and runs the pipeline on the specified schedule without manual intervention. This decouples scheduling from pipeline logic, letting Vertex AI handle orchestration and execution.

Why this answer

Cloud Scheduler is the correct approach because it can directly invoke the Vertex AI Pipeline creation API via an HTTP trigger at a specified cron schedule (e.g., 2 AM daily). This integrates natively with Vertex AI's pipeline orchestration, allowing the scheduler to submit a pipeline run without additional infrastructure. The other options either lack native Vertex AI pipeline support or introduce unnecessary complexity.

Exam trap

A common mistake is confusing scheduling a pipeline run (using Cloud Scheduler + Vertex AI API) with scheduling tasks inside a pipeline (using cron within the pipeline definition). Neither Vertex AI Experiments nor Cloud Run jobs are designed for scheduled pipeline orchestration.

How to eliminate wrong answers

Option B is wrong because Cloud Run jobs are designed for stateless container execution and do not natively support Vertex AI Pipelines; they would require custom code to call the API, adding overhead and breaking the managed pipeline lifecycle. Option C is wrong because Vertex AI Experiments is used for tracking and comparing model training runs, not for scheduling or triggering pipeline executions. Option D is wrong because a cron job inside the pipeline definition would only schedule tasks within a single pipeline run, not trigger the pipeline itself on a recurring schedule.

47
MCQmedium

A data scientist trained a custom TensorFlow model using Vertex AI Training and wants to deploy it for online predictions with low latency (<100ms). Which deployment option on Google Cloud is best?

A.Deploy on Cloud Run with a custom container
B.Deploy on Cloud Functions
C.Deploy on AI Platform Prediction (legacy)
D.Deploy on Vertex AI Endpoints
AnswerD

Vertex AI Endpoints serve models for online prediction with autoscaling and low-latency inference, meeting the sub-100ms requirement. Batch prediction cannot serve real-time requests, and deploying the TensorFlow model directly to Endpoints uses the managed serving stack rather than custom infrastructure.

Why this answer

Vertex AI Endpoints is the correct choice because it is purpose-built for deploying TensorFlow models with optimized serving infrastructure, including automatic scaling, GPU/TPU support, and built-in monitoring for latency-sensitive online predictions. It provides a managed endpoint that can achieve sub-100ms latency by leveraging model optimization techniques like TensorFlow Serving and hardware accelerators, which are not available in the other options.

Exam trap

Google Cloud often tests the misconception that any serverless option (like Cloud Run or Cloud Functions) is sufficient for low-latency ML inference, ignoring the need for GPU acceleration and optimized serving infrastructure that only Vertex AI Endpoints provides.

How to eliminate wrong answers

Option A is wrong because Cloud Run, while supporting custom containers, lacks native GPU/TPU acceleration and has a cold-start latency that can exceed 100ms, making it unsuitable for low-latency online predictions. Option B is wrong because Cloud Functions has a maximum timeout of 9 minutes and no GPU support, and its cold-start latency often exceeds 100ms, making it impractical for real-time inference. Option C is wrong because AI Platform Prediction (legacy) is being deprecated and does not offer the same level of integration with Vertex AI's model registry, monitoring, and autoscaling features, and it may not achieve the same low-latency guarantees as Vertex AI Endpoints.

48
Multi-Selectmedium

A data science team is building a real-time feature engineering pipeline for ML model training and serving. They need to compute features from streaming data, store them for low-latency serving, and ensure consistency between training and serving. Which TWO Google Cloud services should they use?

Select 2 answers
A.Vertex AI Feature Store
B.BigQuery
C.Cloud Functions
D.Cloud Dataflow
E.Cloud SQL
AnswersA, D

Vertex AI Feature Store provides a centralised repository for feature values, enabling low-latency online serving alongside consistent offline retrieval for training. It directly satisfies the stem's requirement for training-serving consistency and low-latency access, ingesting features computed from streaming data without duplicating logic between pipelines.

Why this answer

Vertex AI Feature Store (A) is correct because it provides a centralized repository for storing, serving, and sharing feature data with low-latency online serving and batch serving for training, ensuring consistency between training and serving through point-in-time lookups and feature value time-stamping. Cloud Dataflow (D) is correct because it is a fully managed stream and batch processing service based on Apache Beam, enabling real-time feature engineering from streaming data with exactly-once processing semantics and automatic scaling.

Exam trap

A common trap in Google PMLE exams is assuming BigQuery can serve as a low-latency online feature store for real-time inference, but it is designed for analytical queries with seconds-to-minutes latency, not sub-millisecond serving required for real-time ML inference.

49
MCQeasy

A company uses Vertex AI Model Registry to manage multiple model versions. They want to designate a model version as 'champion' for production deployment and another as 'challenger' for A/B testing. Which feature of the registry should they use?

A.Model version labels
B.Model lineage
C.Model aliases
D.Model evaluation metrics
AnswerC

Model aliases are mutable, named pointers (for example 'champion', 'challenger') that reference a specific version within a registered model, so traffic can be switched between versions without redeploying or changing version IDs. This directly supports designating production and A/B testing versions.

Why this answer

Model aliases in Vertex AI Model Registry let you assign a mutable, named reference (e.g., 'champion', 'challenger') to a specific model version, so deployment endpoints can point to the alias rather than a fixed version ID. This enables seamless promotion or rollback by reassigning the alias, and supports A/B testing by directing traffic to different aliased versions. It is the intended feature for designating champion/challenger roles.

Exam trap

PMLE often tests the confusion between labels (static tags for filtering) and aliases (mutable pointers for deployment), leading candidates to choose labels for champion/challenger designation.

How to eliminate wrong answers

Option A (Model version labels) is wrong because labels are key-value tags for organization and filtering, not mutable pointers that endpoints can resolve to a specific version for deployment. Option B (Model lineage) is wrong because lineage tracks the provenance and relationships of artifacts, not deployment role assignment. Option D (Model evaluation metrics) is wrong because metrics describe model performance (e.g., AUC, accuracy) and do not control which version is served in production or testing.

50
Multi-Selectmedium

A marketing team wants to build a model to predict which customers are likely to churn. They have a BigQuery table with customer demographics, usage metrics, and a binary churn label. They want to use BigQuery ML and need to evaluate the model's performance. Which two statements are true regarding model evaluation in BigQuery ML? (Choose two.)

Select 2 answers
A.ML.FEATURE_INFO returns the importance of each feature in the model.
B.ML.PREDICT automatically calculates the model's accuracy on the input data.
C.ML.TRAINING_INFO provides the evaluation metrics for the trained model.
D.ML.CONFUSION_MATRIX can be used to visualize the confusion matrix of a classification model.
E.ML.EVALUATE returns metrics such as precision, recall, accuracy, and AUC for classification models.
AnswersD, E

ML.CONFUSION_MATRIX generates a confusion matrix showing true positives, true negatives, false positives, and false negatives. This is useful for understanding the types of errors the model makes, which is critical when the cost of false positives and false negatives differs, as in churn prediction.

Why this answer

ML.EVALUATE computes standard classification metrics, and ML.CONFUSION_MATRIX provides a detailed breakdown of prediction outcomes. Both are essential for assessing a churn model. The other functions serve different purposes: ML.PREDICT for scoring, ML.TRAINING_INFO for training details, and ML.FEATURE_INFO for feature statistics.

Exam trap

The trap here is confusing prediction with evaluation; ML.PREDICT does not evaluate, and ML.TRAINING_INFO does not provide classification metrics.

51
MCQeasy

A small marketing team has a CSV file of 2,000 labeled customer support tickets (each with a category such as 'billing' or 'technical'). They have no ML engineers and want a fully managed, low-code way to train a text classification model that they can later call from their internal web app. Which Google Cloud service should they use?

A.BigQuery ML with a CREATE MODEL statement using the LOGISTIC_REG model type
B.Vertex AI Pipelines with a custom Kubeflow component for text preprocessing
C.Cloud Natural Language API classifyText method
D.Vertex AI AutoML text classification
AnswerD

AutoML text classification is a managed, low-code service that trains on a labeled CSV in a Vertex AI dataset and exposes a prediction endpoint for app integration. It handles model selection and tuning, matching the team's lack of ML engineers and need for a callable API.

Why this answer

AutoML text classification is purpose-built for teams with labeled text and no ML expertise: it ingests a CSV, trains, tunes, and serves a model behind a managed endpoint. The other services either require custom code, expect structured numeric features, or rely on a fixed taxonomy that cannot represent the team's custom ticket categories.

Exam trap

The trap here is assuming a general NLP API with a classify method can be trained on custom labels, when it actually applies only a predefined taxonomy.

52
MCQeasy

A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?

A.Use Vertex AI Dataset service to create a dataset and export it to BigQuery.
B.Use BigQuery snapshots to capture a versioned dataset and reference the snapshot in the training pipeline.
C.Train the model directly on the BigQuery table and let AutoML handle versioning.
D.Export the data to a timestamped CSV file and store it in Cloud Storage before each training run.
AnswerB

BigQuery snapshots preserve table data at a point in time and are immutable, so each training run references a fixed snapshot rather than the daily-changing source table. This guarantees a consistent dataset version and reproducibility without manual CSV exports.

Why this answer

BigQuery snapshots provide a consistent, versioned view of the dataset at a specific point in time, ensuring reproducibility without duplicating data. By referencing the snapshot in the Vertex AI training pipeline, the team can train models on the exact same data snapshot, even as the source table is updated daily. This approach avoids the overhead of exporting to CSV and Cloud Storage while maintaining data integrity and lineage.

Exam trap

Google Cloud often tests the misconception that exporting to CSV or using Vertex AI Dataset is sufficient for versioning, when in fact BigQuery snapshots provide the native, scalable, and auditable mechanism for point-in-time data consistency without data duplication.

How to eliminate wrong answers

Option A is wrong because the Vertex AI Dataset service is designed for managing training data within Vertex AI, but exporting to BigQuery does not inherently create a versioned snapshot; it simply moves data back to BigQuery without preserving a consistent point-in-time copy. Option C is wrong because training directly on a live BigQuery table does not guarantee a consistent snapshot; AutoML does not handle versioning, and the table may change between training runs, breaking reproducibility. Option D is wrong because exporting to a timestamped CSV file in Cloud Storage is a manual workaround that introduces storage overhead, potential data drift from export timing, and lacks the built-in versioning and query capabilities of BigQuery snapshots.

53
MCQmedium

A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?

A.Increase maxReplicaCount to 10 and configure the autoscaling metric to CPU utilization with a target of 60%.
B.Deploy the model to a new endpoint with minReplicaCount=1 and maxReplicaCount=10, and use a traffic split to gradually shift traffic from the old endpoint.
C.Set minReplicaCount to 3 and maxReplicaCount to 10, and enable autoscaling based on CPU utilization with a target of 80%.
D.Keep minReplicaCount=1 but set maxReplicaCount to 10, and configure autoscaling based on a custom metric that tracks the number of incoming requests per second.
AnswerD

This approach allows the endpoint to scale out rapidly during traffic spikes by using a request-based metric that directly reflects load, while keeping the minimum replica count at 1 to save costs during idle periods. A custom metric such as requests per second is more responsive for sudden spikes than CPU utilization, which may lag. Increasing max replicas provides headroom.

Why this answer

The best solution is to increase the maximum replica count to handle spikes and use a custom metric that directly measures request load, such as requests per second, for autoscaling. This allows rapid scaling during flash sales while keeping the minimum replicas low to control costs. CPU-based autoscaling may not react quickly enough for sudden spikes, and raising the minimum replicas increases baseline cost unnecessarily.

Exam trap

The trap here is assuming that CPU utilization is always the best autoscaling metric, when in fact request-based metrics can be more responsive for sudden traffic spikes.

54
MCQmedium

A data-processing pipeline using Dataflow needs to incorporate a custom ML prediction step. The team wants to maintain fast processing and minimize latency. What is the optimal approach?

A.Write the data to Cloud Storage, trigger a Cloud Function to call the model, and write results back
B.Use a custom ParDo transform in Dataflow that calls Vertex AI Prediction API directly
C.Send data to a Pub/Sub topic and have a separate subscriber that runs predictions
D.Stream data through Cloud Functions that serve predictions and write to BigQuery
AnswerB

A custom ParDo transform runs the prediction call inside the Dataflow worker pipeline, streaming each element to the Vertex AI Prediction API without an intermediate storage hop. This preserves fast processing and minimises latency, satisfying the low-latency constraint better than batch export-and-reimport patterns.

Why this answer

Using a custom ParDo transform in Dataflow allows the pipeline to call the Vertex AI Prediction API synchronously within each worker, avoiding the overhead of external triggers, intermediate storage, or asynchronous messaging. This keeps the data in-memory and minimizes latency by processing predictions inline with the Dataflow streaming or batch pipeline.

Exam trap

Google Cloud often tests the misconception that adding external services like Cloud Functions or Pub/Sub improves modularity without considering the latency penalty, leading candidates to choose options that introduce unnecessary hops instead of keeping prediction inline within the Dataflow pipeline.

How to eliminate wrong answers

Option A is wrong because writing data to Cloud Storage and triggering a Cloud Function introduces significant I/O latency and additional orchestration overhead, breaking the low-latency requirement. Option C is wrong because sending data to Pub/Sub and having a separate subscriber decouples the prediction step, adding network round-trips and potential backpressure issues that increase end-to-end latency. Option D is wrong because streaming data through Cloud Functions for predictions and then writing to BigQuery creates a multi-hop architecture with cold-start risks and no native Dataflow optimization for parallelism or state management.

55
MCQmedium

A data scientist has deployed a model with Vertex AI Endpoints and enabled request/response logging to BigQuery. They want to compute a confusion matrix over time to monitor model quality. What should they do?

A.Use Vertex AI Model Monitoring to automatically generate confusion matrices
B.Use Cloud Monitoring to create a confusion matrix dashboard
C.Upload ground truth labels to BigQuery and join with prediction logs, then compute confusion matrix in a scheduled query
D.Enable Vertex AI Explainability to get confusion matrix
AnswerC

Confusion matrices need both predictions and ground truth labels. Vertex AI request/response logging writes predictions to BigQuery, so joining uploaded ground truth labels there and computing the matrix in a scheduled query satisfies the stem's monitoring requirement.

Why this answer

To compute a confusion matrix over time for a deployed model on Vertex AI Endpoints with request/response logging to BigQuery, you need ground truth labels. Vertex AI does not automatically generate confusion matrices from prediction logs alone. You must upload the actual labels (ground truth) to BigQuery and join them with the prediction logs, then compute the confusion matrix using a scheduled query.

This allows you to monitor model quality over time. Vertex AI Model Monitoring (A) can detect skew and drift but does not automatically generate confusion matrices. Cloud Monitoring (B) is for infrastructure and application metrics, not model quality.

Explainability (D) provides feature attributions, not confusion matrices.

Exam trap

PMLE often tests the misconception that Vertex AI Model Monitoring or Explainability automatically provides confusion matrices, when in fact confusion matrices require ground truth labels and custom computation.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Monitoring focuses on detecting training-serving skew and prediction drift, not on computing confusion matrices, which require ground truth labels. Option B is wrong because Cloud Monitoring is designed for operational metrics (CPU, latency, etc.) and does not have built-in capabilities to compute confusion matrices from prediction logs. Option D is wrong because Vertex AI Explainability provides explanations for individual predictions (e.g., feature attributions) and does not generate confusion matrices.

56
MCQhard

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline includes a step that trains a model and a subsequent step that evaluates the model. The engineer wants to ensure that the evaluation step runs only if the training step succeeds and that the pipeline fails if the model's accuracy is below a threshold. Which approach should the engineer use?

A.Set the evaluation step's retry policy to zero and use the pipeline's built-in accuracy threshold parameter.
B.Use a condition in the pipeline that checks the training step's status and a custom component that raises an exception if accuracy is below the threshold.
C.Use a Vertex AI Model resource to store the model and enable Model Monitoring to trigger a pipeline failure if accuracy drops.
D.Configure the evaluation step to always run and rely on Vertex AI Pipelines to automatically fail the pipeline if the accuracy metric is below the threshold.
AnswerB

Vertex AI Pipelines supports conditions to control step execution based on the status of previous steps. A custom component can evaluate the model and raise an exception if the accuracy does not meet the threshold. This exception will cause the pipeline to fail, ensuring that only models meeting the criteria proceed. This approach provides the required control flow and failure behavior.

Why this answer

To conditionally run a step and fail the pipeline based on a metric, the engineer should use a condition to check the training step's status and a custom component that raises an exception if accuracy is too low. Vertex AI Pipelines conditions allow steps to run only if previous steps succeed, and raising an exception in a component causes the pipeline to fail. This combination provides the required control flow and failure semantics.

Exam trap

The trap here is assuming that Vertex AI Pipelines automatically fails on low accuracy metrics or that Model Monitoring can be used within a pipeline for this purpose.

57
MCQeasy

A machine learning model deployed on Vertex AI is returning erroneous predictions. The team needs to investigate the root cause by examining the prediction request and response details. Which Google Cloud tool is best suited for this?

A.Cloud Monitoring
B.Cloud Debugger
C.Cloud Logging
D.Cloud Trace
AnswerC

Cloud Logging captures the raw prediction request and response payloads emitted by Vertex AI endpoints, letting the team inspect exact inputs and outputs to trace erroneous predictions. It satisfies the need to examine request and response details, which aggregate metrics alone cannot expose.

Why this answer

Cloud Logging is the correct tool because it captures detailed logs of prediction requests and responses, including input features, model outputs, and any errors. By examining these logs, the team can trace the exact data flow and identify discrepancies causing erroneous predictions, such as data preprocessing issues or model version mismatches.

Exam trap

The trap here is that candidates confuse Cloud Monitoring (which shows aggregate health metrics) with Cloud Logging (which provides granular request/response data), leading them to choose a tool that cannot reveal the specific prediction details needed for root cause analysis.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring focuses on metrics and alerting (e.g., latency, error rates) but does not capture the content of individual prediction requests or responses. Option B is wrong because Cloud Debugger is designed for inspecting live application code state (e.g., variable values) in production, not for logging request/response payloads of ML predictions. Option D is wrong because Cloud Trace provides latency analysis and distributed tracing of requests across services, but it does not log the actual prediction data or response details needed to debug prediction errors.

58
Multi-Selecthard

A company is using Vertex AI Pipelines for ML workflows. They want to implement best practices for idempotent components and data passing. Which THREE practices should they adopt?

Select 3 answers
A.Pass large datasets between components using GCS URIs instead of in-memory values.
B.Avoid hard-coding file paths; use pipeline parameters to pass URIs.
C.Read data into memory in the first component and pass the in-memory object to subsequent components.
D.Use global variables in the pipeline code to store intermediate results.
E.Design components to be idempotent so that the same input always produces the same output.
AnswersA, B, E

Passing GCS URIs keeps components stateless and idempotent, since each run reads from an immutable location rather than relying on in-memory state that cannot survive retries or cross-component boundaries. This satisfies the stem's data-passing best practice for large datasets.

Why this answer

Option A is correct because Vertex AI Pipelines components exchange data through artifacts, and passing large datasets as GCS URIs (e.g., gs://bucket/path) avoids serializing huge in-memory objects into the pipeline's metadata/execution store, keeping runs efficient and reproducible. Option B is correct because hard-coded paths break portability and reproducibility across environments; using pipeline parameters (or input artifacts) to supply URIs lets the same pipeline definition run against dev, test, or prod buckets without code changes. Option E is correct because idempotent components—where identical inputs always yield identical outputs and re-execution has no additional side effects—are a core Vertex AI Pipelines best practice, enabling safe retries and caching.

Option C is not appropriate because passing in-memory objects between components requires serialization, can exceed metadata limits, and undermines the artifact-based data-passing model. Option D is not appropriate because global variables introduce hidden state, are not tracked as pipeline parameters or artifacts, and break reproducibility and caching across pipeline runs.

Exam trap

A common misconception in Vertex AI Pipelines is that in-memory data passing is acceptable in containerized components, but the correct pattern is to use GCS URIs and artifact references to ensure idempotency and scalability.

59
MCQeasy

An ML engineer is designing a CI/CD pipeline for ML models using Cloud Build and Cloud Deploy. They want to automatically test model performance on a validation set before promoting to production. Which step should be included in the CI/CD pipeline?

A.Run unit tests on the training code
B.Use Cloud Composer to schedule evaluation
C.Deploy to production immediately after training
D.Train the model in the CI/CD pipeline
E.Run a Vertex AI Pipeline for model evaluation and register the model only if metrics exceed thresholds
AnswerE

Embedding a Vertex AI Pipeline evaluation step gates promotion on measured validation metrics, so Cloud Deploy only releases a model whose performance exceeds defined thresholds. This satisfies the stem's requirement to test performance automatically before production, rather than relying on manual review or post-deployment monitoring.

Why this answer

It directly integrates model evaluation into the CI/CD pipeline using Vertex AI Pipelines, which allows automated validation of model performance against predefined thresholds before promotion. This ensures that only models meeting quality criteria are deployed, aligning with MLOps best practices for gated promotions.

Exam trap

Google Cloud often tests the distinction between code testing (unit tests) and model validation (performance metrics), leading candidates to choose A because they conflate software testing with ML evaluation.

How to eliminate wrong answers

Option A is wrong because unit tests on training code verify code correctness but do not assess model performance on a validation set, which is the requirement. Option B is wrong because Cloud Composer is an orchestration tool for workflows, not a CI/CD step for automatic model evaluation before promotion; it would introduce scheduling latency rather than inline gating. Option C is wrong because deploying immediately after training bypasses validation, risking production degradation from underperforming models.

Option D is wrong because training the model in the CI/CD pipeline is possible but does not include the evaluation step needed to gate promotion; it focuses on the training process itself, not validation.

60
MCQmedium

A team wants to share feature definitions across multiple projects in their organization using Vertex AI Feature Store. What is the recommended approach?

A.Export features to BigQuery datasets in each project
B.Use Vertex AI Feature Store's feature view for cross-project access
C.Create separate feature stores in each project and synchronize them with Dataflow
D.Use a centralized feature store in a shared project and grant access to other projects via IAM
AnswerD

A centralized feature store hosted in one shared project lets multiple projects consume identical feature definitions, with IAM policies granting cross-project access. This satisfies the requirement to share definitions organisation-wide while avoiding duplicated, divergent feature stores per project.

Why this answer

Vertex AI Feature Store is project-scoped, so to share features across projects, the recommended pattern is to create a centralized feature store in a shared (host) project and grant IAM roles to users/service accounts in other projects. This avoids duplication and ensures a single source of truth for feature definitions and online serving.

Exam trap

The trap is thinking feature views or BigQuery exports provide cross-project sharing; the correct pattern is a centralized feature store with IAM grants, which candidates often overlook in favor of data duplication.

How to eliminate wrong answers

Option A is wrong because exporting features to BigQuery datasets in each project duplicates data and does not provide online serving or feature freshness management — it is a batch workaround, not a feature store sharing mechanism. Option B is wrong because a feature view is a logical grouping within a feature store, not a cross-project access mechanism; feature views do not grant cross-project permissions by themselves. Option C is wrong because creating separate feature stores and synchronizing with Dataflow introduces complexity, latency, and consistency issues — it is an anti-pattern for sharing definitions.

61
MCQmedium

A team is scaling a prototype ML model to production on Vertex AI. The model was developed using scikit-learn and requires custom preprocessing. They want to minimize operational overhead and ensure consistency between training and serving. Which approach should they use?

A.Train on a local machine and upload the model artifacts to Cloud Storage, then create an endpoint with a pre-built container.
B.Use a pre-built Vertex AI container for scikit-learn and provide a custom training Python package with preprocessing code included.
C.Deploy the model as a custom prediction routine on Vertex AI Endpoints with a custom container.
D.Export the model as a .pkl file and use Vertex AI's 'Import Model' with a default container for inference.
AnswerB

Packaging preprocessing inside a custom training Python package lets Vertex AI's pre-built scikit-learn container run both training and prediction, so the same code executes at serving time. This satisfies the consistency constraint directly, while the managed container removes the operational overhead of building and maintaining your own image.

Why this answer

Using a pre-built Vertex AI container for scikit-learn with a custom training Python package ensures that the same preprocessing code runs during both training and serving, minimizing operational overhead. This approach leverages Vertex AI's managed infrastructure to handle scaling, monitoring, and consistency without requiring custom container maintenance.

Exam trap

The trap here is that candidates often assume a pre-built container cannot handle custom preprocessing, leading them to choose a custom container (Option C) or a simpler import (Option D), but Vertex AI allows embedding preprocessing in the training package or model artifact to maintain consistency with minimal overhead.

How to eliminate wrong answers

Option A is wrong because training on a local machine and uploading model artifacts to Cloud Storage, then creating an endpoint with a pre-built container, does not guarantee consistency between training and serving preprocessing logic, as the preprocessing code is not bundled with the model. Option C is wrong because deploying the model as a custom prediction routine with a custom container introduces unnecessary operational overhead for a scikit-learn model that can be served with a pre-built container, and it requires building and maintaining a custom Docker image. Option D is wrong because exporting the model as a .pkl file and using Vertex AI's 'Import Model' with a default container for inference does not include custom preprocessing code, leading to potential inconsistencies between training and serving.

62
MCQeasy

A non-technical user wants to build a binary classification model using Vertex AI. Which UI should they use?

A.Vertex AI AutoML
B.Vertex AI Workbench
C.Vertex AI Pipelines
D.Vertex AI Prediction
AnswerA

AutoML lets non-technical users train models through a guided point-and-click interface, automatically handling feature engineering, algorithm selection and hyperparameter tuning. This satisfies the stem's constraint that the user lacks ML expertise, unlike custom training in notebooks or pipelines, which demands coding and modelling knowledge.

Why this answer

Vertex AI AutoML is the correct choice because it provides a no-code graphical user interface specifically designed for non-technical users to build, train, and deploy machine learning models, including binary classification models, without writing any code. It automates the entire ML pipeline—feature engineering, model selection, hyperparameter tuning—allowing users to simply upload labeled data and get a production-ready model.

Exam trap

Google Cloud often tests the distinction between 'building/training' tools (AutoML) and 'deploying/serving' tools (Prediction), leading candidates to mistakenly choose Vertex AI Prediction because they confuse the deployment phase with the model creation phase.

How to eliminate wrong answers

Option B is wrong because Vertex AI Workbench is a Jupyter notebook-based development environment intended for data scientists and ML engineers who write custom code, not for non-technical users seeking a low-code solution. Option C is wrong because Vertex AI Pipelines is a tool for orchestrating and automating ML workflows using code-defined pipelines (e.g., Kubeflow Pipelines SDK), requiring programming skills to define steps and dependencies. Option D is wrong because Vertex AI Prediction is a serving endpoint for deploying and running inference on already-trained models, not a UI for building or training models from scratch.

63
MCQmedium

A company wants to track the cost of their Vertex AI prediction endpoint. They use a custom machine type with 1 n1-standard-4 (4 vCPU, 15 GB memory) and 1 NVIDIA T4 GPU. The endpoint is configured for automatic scaling with min=1, max=5 replicas. Which cost monitoring approach should they use?

A.Use Cloud Billing budget alerts and export cost data to BigQuery for analysis.
B.Calculate cost manually based on replica count and GPU hours from endpoint logs.
C.Use Vertex AI Experiments to track cost.
D.Monitor only the CPU utilisation metrics to infer cost.
AnswerA

Cloud Billing budget alerts notify on spend thresholds, and exporting cost data to BigQuery enables granular analysis by label, including the custom machine type and GPU replicas. This satisfies the stem's requirement to track prediction endpoint costs with autoscaling replicas.

Why this answer

For tracking Vertex AI prediction endpoint costs, the most comprehensive approach is to use Cloud Billing budget alerts and export cost data to BigQuery for detailed analysis. This allows the company to monitor actual spend, set alerts, and analyze costs by labels, projects, or services, including Vertex AI endpoints with custom machine types and GPUs.

Exam trap

The trap is choosing manual calculation or unrelated tools like Vertex AI Experiments, when the correct approach is to leverage native GCP cost management tools (Cloud Billing + BigQuery export) for accurate and scalable cost monitoring.

How to eliminate wrong answers

Option B is wrong because manually calculating cost from replica count and GPU hours is error-prone and does not account for actual billing nuances like sustained use discounts, committed use discounts, or network egress; it also lacks integration with billing alerts. Option C is wrong because Vertex AI Experiments is designed for tracking ML experiment metrics and parameters, not for cost monitoring or billing analysis. Option D is wrong because monitoring only CPU utilization does not provide cost data; cost depends on many factors including GPU usage, memory, and replica count, and CPU metrics alone cannot infer cost accurately.

64
Multi-Selecthard

Which TWO strategies help ensure data consistency when multiple teams are contributing features to a shared Vertex AI Feature Store?

Select 2 answers
A.Each team should create their own feature store to avoid conflicts.
B.Use only batch ingestion to keep features synchronized.
C.Define and enforce feature schemas using the Feature Store API.
D.Allow each team to independently define feature engineering logic.
E.Set up monitoring and alerting on feature value distributions to detect drift.
AnswersC, E

Schemas ensure consistent data types and values.

Why this answer

Defining and enforcing feature schemas using the Vertex AI Feature Store API ensures that all teams adhere to a consistent data structure (e.g., fixed feature names, data types, and value ranges). This prevents schema drift and ingestion conflicts, which are common when multiple teams independently push features to the same feature store. Without schema enforcement, one team might inadvertently change a feature's data type or add unexpected values, breaking downstream models.

Exam trap

Google Cloud often tests the misconception that 'separate stores' or 'batch-only ingestion' are valid consistency strategies, when in fact the correct approach is centralized schema governance with monitoring to detect drift.

65
Multi-Selecthard

A financial institution uses a machine learning model to approve loans. They must monitor for fairness and bias. Which THREE Google Cloud tools or features can help them achieve this? (Choose 3.)

Select 3 answers
A.What-If Tool
B.Vertex AI Model Monitoring
C.Cloud Data Loss Prevention
D.Cloud Healthcare API
E.Explainable AI
AnswersA, B, E

The What-If Tool lets you probe a trained model with counterfactual examples, swapping protected attributes such as gender or ethnicity to reveal whether predictions shift unfairly. This directly satisfies the stem's fairness and bias monitoring requirement by exposing disparate impact across loan approval decisions.

Why this answer

The What-If Tool (A) is a visualization toolkit integrated with Vertex AI that lets analysts probe a trained model with counterfactual and hypothetical inputs to inspect how predictions change across sensitive attributes like gender or ethnicity, directly supporting fairness and bias analysis. Vertex AI Model Monitoring (B) continuously tracks deployed models for training-serving skew and prediction drift, and can be configured with fairness-related metrics so the institution detects when a model's behavior degrades or becomes biased in production. Explainable AI (E) provides feature attributions (e.g., Sampled Shapley and integrated gradients) through Vertex AI, showing which features drive each loan decision so reviewers can confirm that protected attributes are not improperly influencing approvals.

Cloud Data Loss Prevention (C) is for discovering and redacting sensitive data such as PII, not for evaluating model fairness, and Cloud Healthcare API (D) is a healthcare-specific data interoperability service unrelated to loan-model bias monitoring.

Exam trap

Google Cloud often tests the distinction between data security tools (like DLP) and ML fairness tools, so candidates mistakenly select Cloud DLP thinking it addresses bias because it handles sensitive attributes, but DLP does not analyze model predictions or fairness metrics.

66
MCQhard

A large enterprise has multiple ML models deployed in production across different regions. They want to implement a centralized monitoring dashboard that tracks key performance indicators such as prediction accuracy, latency, and error rates for all models, with the ability to drill down into individual model versions. Which approach best meets these requirements?

A.Use Vertex AI Experiments to log metrics and compare across runs
B.Use Cloud Logging to search logs from each model and create a dashboard
C.Use BigQuery to store prediction logs and then visualize in Looker
D.Use Cloud Monitoring with custom metrics reported by each model deployment, and create a unified dashboard with filterable resources
AnswerD

Custom metrics let each deployment report accuracy, latency and error rates into Cloud Monitoring, where a single dashboard with filterable resource labels aggregates all models and supports drilling into individual versions. This satisfies the centralised, cross-region visibility requirement.

Why this answer

Cloud Monitoring with custom metrics allows each model deployment to report key performance indicators (e.g., prediction accuracy, latency, error rates) as metric time series. These custom metrics can be aggregated into a single unified dashboard, and the dashboard can be configured with filterable resources (e.g., region, model version) to enable drill-down into individual model versions. This approach provides centralized, real-time monitoring without relying on log-based or batch analytics.

Exam trap

Google Cloud often tests the distinction between logging (Cloud Logging) and monitoring (Cloud Monitoring), where candidates mistakenly think log-based dashboards are sufficient for real-time KPI tracking, ignoring the need for structured, low-latency custom metrics.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is designed for tracking and comparing training runs (e.g., hyperparameter tuning), not for real-time monitoring of deployed models in production across regions. Option B is wrong because Cloud Logging is a log management service that requires parsing unstructured log entries to extract metrics, which is inefficient for real-time KPIs and lacks native metric aggregation and dashboard drill-down capabilities. Option C is wrong because BigQuery is a data warehouse for storing and querying large datasets, and while Looker can visualize it, this approach introduces latency from batch loading and is not designed for real-time monitoring of live model deployments.

67
MCQeasy

Refer to the exhibit. A data scientist runs this Vertex AI training job code. What will be the outcome?

A.The job runs as a regular custom training with 10 replicas.
B.A HyperparameterTuningJob is created and runs trials.
C.A CustomJob is created with hyperparameters from the spec.
D.The job fails because parallel_trial_count cannot be less than max_trial_count.
AnswerB

Passing a hyperparameter tuning configuration alongside the training script causes Vertex AI to launch a HyperparameterTuningJob rather than a single CustomJob. The service then runs multiple trials, each training with different hyperparameter values, and selects the best-performing trial.

Why this answer

The code uses `HyperparameterTuningJob` with `parallel_trial_count=1` and `max_trial_count=10`. This creates a hyperparameter tuning job that runs up to 10 trials, each trial being a separate training run with different hyperparameter values. The `parallel_trial_count=1` means trials run sequentially, not in parallel, but this is valid and does not cause failure.

Exam trap

Google Cloud often tests the misconception that `parallel_trial_count` must be equal to or greater than `max_trial_count`, when in reality it can be any value from 1 to `max_trial_count`, and sequential trials are perfectly valid.

How to eliminate wrong answers

Option A is wrong because the code explicitly creates a `HyperparameterTuningJob`, not a regular custom training job; a regular custom training job would use `CustomJob` or `CustomContainerTrainingJob` without hyperparameter tuning parameters. Option C is wrong because a `CustomJob` does not accept hyperparameter tuning parameters like `parallel_trial_count` or `max_trial_count`; those are specific to `HyperparameterTuningJob`. Option D is wrong because `parallel_trial_count` can be less than `max_trial_count`; the constraint is that `parallel_trial_count` must be less than or equal to `max_trial_count`, and 1 ≤ 10 is valid.

68
MCQmedium

A fraud detection model is deployed to a Vertex AI Endpoint and configured with Vertex AI Model Monitoring for feature drift. The team wants the drift monitor to compare live production traffic against the exact statistics captured from the training dataset, so that alerts reflect deviation from the model's original data distribution rather than from recent traffic. Which configuration should they use?

A.Enable prediction drift instead of feature drift and let Vertex AI infer the baseline automatically.
B.Configure the monitor to use a rolling 24-hour window of live requests as the baseline distribution.
C.Attach the training dataset as a Vertex AI managed dataset and set the monitoring frequency to daily.
D.Set the monitoring training dataset to the original training data and enable drift detection.
AnswerD

Vertex AI Model Monitoring computes drift by comparing live feature distributions against a baseline derived from a training dataset or a saved baseline. Pointing the monitor at the original training data fixes the reference distribution to the model's training period, so drift alerts reflect deviation from that stable baseline rather than from rolling production windows.

Why this answer

Feature drift detection in Vertex AI Model Monitoring requires a baseline distribution to compare against. Using the original training data as the baseline anchors the comparison to the model's training distribution, which is exactly what the team wants. Other options either change the comparison window, switch to prediction drift, or alter scheduling without defining the reference, so they do not meet the stated requirement.

Exam trap

The trap here is assuming that any monitoring configuration with drift enabled will automatically compare against training data, when the baseline must be explicitly specified.

69
MCQmedium

A financial services company uses BigQuery ML to build a logistic regression model for fraud detection. The model is trained on the last 6 months of transaction data (about 50 million rows). After deployment, the fraud detection team notices a high false positive rate, causing customer dissatisfaction and extra manual review costs. The model is currently retrained monthly. The team wants to reduce false positives without sacrificing recall. They have access to real-time transaction streaming and can compute new features quickly. What is the most effective approach?

A.Replace logistic regression with gradient boosted trees (XGBoost) in BigQuery ML
B.Use Vertex AI AutoML Tables to train a more complex model
C.Increase retraining frequency to daily
D.Add engineered features like rolling transaction count and velocity per user
AnswerD

Rolling transaction count and velocity per user capture behavioural patterns over recent windows, giving the logistic regression model stronger discriminative signals than raw transaction fields. These features directly reduce false positives while preserving recall, and real-time streaming lets them be computed at scoring time.

Why this answer

Adding engineered features like rolling transaction count and velocity per user directly addresses the high false positive rate by providing the logistic regression model with more discriminative temporal signals. Since the team has access to real-time streaming and can compute features quickly, these features capture behavioral patterns that reduce false positives without sacrificing recall, and logistic regression can effectively leverage them with proper feature engineering.

Exam trap

The trap here is that candidates often assume a more complex model (XGBoost or AutoML) is always better for reducing false positives, but the question specifically tests the principle that feature engineering—especially temporal aggregations—is the most effective lever when the model is already appropriate and data is streaming.

How to eliminate wrong answers

Option A is wrong because replacing logistic regression with gradient boosted trees (XGBoost) may improve model capacity but does not directly target the root cause of high false positives—lack of informative features—and could increase complexity without guaranteed recall preservation. Option B is wrong because using Vertex AI AutoML Tables to train a more complex model similarly addresses model complexity rather than feature insufficiency, and may introduce overfitting or latency issues without solving the false positive problem. Option C is wrong because increasing retraining frequency to daily does not change the underlying feature set or model architecture; it only refreshes weights on the same features, which will not reduce false positives if the model lacks discriminative signals.

70
MCQeasy

A financial company is building a fraud detection model. The dataset has 1% fraud cases and 99% legitimate transactions. Which technique should they use to handle the class imbalance?

A.Use class weighting or synthetic oversampling (SMOTE) during training
B.Randomly undersample the majority class to balance the dataset
C.Collect more data until the fraud rate increases
D.Train without any modifications; the model will naturally handle it
AnswerA

With only 1% fraud, a model optimising raw accuracy predicts legitimate for everything. Class weighting penalises minority-class errors more heavily, while SMOTE synthesises new minority samples, both shifting the decision boundary so fraud patterns are actually learned during training.

Why this answer

For highly imbalanced datasets like fraud detection (1% fraud), class weighting or synthetic oversampling (SMOTE) during training helps the model learn the minority class patterns without discarding valuable majority-class data. Class weighting adjusts the loss function to penalize minority-class misclassifications more heavily, while SMOTE generates synthetic minority samples to balance the class distribution, improving recall and F1 for the fraud class.

Exam trap

PMLE often tests whether candidates recognize that accuracy is not a valid metric for imbalanced data and that techniques like SMOTE or class weighting are necessary—undersampling is a common distractor but can discard critical information.

How to eliminate wrong answers

Option B is wrong because randomly undersampling the majority class discards potentially useful legitimate transaction data, which can lead to loss of information and degraded overall performance, especially if the majority class has diverse patterns. Option C is wrong because collecting more data until the fraud rate increases is impractical and does not address the inherent imbalance; fraud rates are typically low and may not change. Option D is wrong because training without modifications on a 1% fraud dataset will likely cause the model to predict 'legitimate' for almost all cases, achieving 99% accuracy but failing to detect fraud (high false negatives).

71
MCQhard

A team is training a TensorFlow model on Vertex AI using a custom container. The training script writes checkpoints to a local directory inside the container. The job runs for 14 hours, and when it completes, the team cannot find the checkpoints in Cloud Storage. They need the checkpoints to be persisted so they can resume training and deploy the best model. What should they do?

A.Enable Vertex AI TensorBoard integration and configure the training script to log checkpoints as TensorBoard artifacts, then retrieve them from the TensorBoard instance.
B.Set the training job's base output directory to a local path and rely on Vertex AI to automatically upload everything under that path to the job's Cloud Storage output directory at the end of training.
C.Increase the boot disk size of the training VM and re-run the job, then copy the checkpoints from the boot disk to Cloud Storage after training finishes.
D.Configure the training job to write checkpoints to a Cloud Storage URI by passing a gs:// path to the checkpoint directory in the training script, and ensure the Vertex AI service account has storage.objectAdmin on the bucket.
AnswerD

Writing checkpoints directly to a gs:// URI makes TensorFlow use its GCS filesystem implementation, so checkpoint files are persisted in Cloud Storage as they are written. The Vertex AI custom training service account must have permission to write to the bucket, which storage.objectAdmin grants. This is the standard pattern for durable checkpoints in Vertex AI training and directly solves the missing-checkpoint problem.

Why this answer

Checkpoints must be written to a durable location, and in Vertex AI custom training the durable location is Cloud Storage. Passing a gs:// URI to the checkpoint directory makes TensorFlow persist each checkpoint through its GCS filesystem, and granting storage.objectAdmin to the training service account authorizes those writes. This allows both resuming training and retrieving the best model for deployment.

Exam trap

The trap here is assuming that Vertex AI automatically uploads local training artifacts to Cloud Storage, when only files written to a gs:// path are persisted.

72
MCQhard

An ML team has set up automated retraining triggered by Cloud Monitoring alerts. When a feature drift alert fires, a Cloud Function publishes to Pub/Sub, which triggers a Vertex AI Pipeline. However, the retraining pipeline is failing because the training data is not updated. What is the most likely cause?

A.The Cloud Function does not have permission to start the pipeline
B.The Pub/Sub topic is incorrectly configured
C.The training data in the pipeline input is stale or not refreshed
D.The model endpoint is overloaded
AnswerC

The pipeline executes correctly but trains on an unchanged dataset, so drift persists. The alert and Pub/Sub trigger fire as designed; the failure lies in the input source, which is not being refreshed from the updated feature store or data location before the pipeline runs.

Why this answer

The pipeline is triggering correctly (the alert fires, the Cloud Function runs, Pub/Sub delivers, and the Vertex AI Pipeline starts), so the failure is downstream of orchestration — it is the data itself. When a drift alert fires, the pipeline's input dataset or feature table must be refreshed from the source system before training; if the pipeline references a static/stale snapshot or a BigQuery view that is not re-materialized, training runs on old data and fails validation or produces a useless model. The most likely cause is therefore that the training data in the pipeline input is stale or not refreshed.

Exam trap

The trap here is that candidates fixate on the orchestration layer (permissions, Pub/Sub, endpoints) because those are the visible components, when the question explicitly says the pipeline is failing due to training data not being updated — a data-pipeline problem, not an infrastructure problem.

How to eliminate wrong answers

Option A is wrong because a missing IAM permission (e.g., roles/aiplatform.user on the service account) would prevent the pipeline from starting at all, producing an authorization error rather than a data-not-updated failure. Option B is wrong because a misconfigured Pub/Sub topic would break message delivery entirely, so the pipeline would never be triggered — the scenario states the pipeline is running and failing. Option D is wrong because an overloaded model endpoint affects online prediction serving, not the batch training pipeline, and would surface as latency/5xx errors on inference, not stale training data.

73
MCQmedium

A data analyst wants to use Vision API to detect custom objects in manufacturing images, but the pre-trained API does not recognize their specific components. They have 1000 labeled images. Which path offers the fastest time-to-value with minimal coding?

A.Store images in BigQuery and use ML.PREDICT with a custom model
B.Use AutoML Vision for object detection
C.Use a Cloud Function to call the Vision API and post-process results
D.Train a custom object detection model using TensorFlow on Vertex AI
AnswerB

AutoML Vision object detection trains on your labelled images and handles training infrastructure automatically, requiring no model code. With 1000 labelled images it satisfies the custom-component requirement while delivering the fastest time-to-value compared with building and tuning a custom model.

Why this answer

AutoML Vision for object detection is the fastest path because it requires no custom coding—users simply upload labeled images, and the platform automatically trains a model tailored to their custom components. This directly addresses the need to detect objects the pre-trained Vision API cannot recognize, while minimizing time-to-value compared to manual TensorFlow training or custom infrastructure setup.

Exam trap

Google Cloud often tests the misconception that any cloud function or API call can be adapted to custom objects via post-processing, but the pre-trained Vision API's fixed label set cannot be extended without retraining, making AutoML the only low-code solution that actually learns new object classes.

How to eliminate wrong answers

Option A is wrong because BigQuery ML.PREDICT is designed for structured data and tabular models, not for image-based object detection; storing images in BigQuery and using ML.PREDICT would require converting images to embeddings or using a pre-trained model, which does not solve the custom object recognition problem efficiently. Option C is wrong because calling the pre-trained Vision API via Cloud Function and post-processing results still relies on the same pre-trained model that cannot recognize the custom components, so it fails to address the core requirement. Option D is wrong because training a custom model using TensorFlow on Vertex AI requires significant coding, manual architecture design, and hyperparameter tuning, which is far slower and more complex than using AutoML Vision's no-code automated training pipeline.

74
MCQhard

A machine learning pipeline in Vertex AI produces a dataset artifact, a trained model, and evaluation metrics. The team wants to query the lineage to find all downstream artifacts that depend on a particular dataset. Which Vertex AI service should they use?

A.Vertex AI Feature Store
B.Vertex AI Experiments
C.Vertex AI Model Registry
D.Vertex AI Metadata
AnswerD

Vertex AI Metadata stores artefacts, executions and contexts in a lineage graph, so querying it returns every downstream artefact derived from a given dataset. This satisfies the stem's need to trace dependencies from a specific dataset artefact.

Why this answer

Vertex AI Metadata is the service that records and stores ML metadata — artifacts, executions, and contexts — and their relationships, forming the lineage graph. It lets you query upstream and downstream dependencies of any artifact, such as finding all models and metrics derived from a dataset. This is exactly the lineage-query capability the team needs.

Exam trap

PMLE often tests the overlap between Vertex AI Experiments (run tracking) and Vertex AI Metadata (lineage graph), so candidates pick Experiments when the question explicitly asks for upstream/downstream dependency queries.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store manages and serves feature values, not artifact lineage across pipeline runs. Option B is wrong because Vertex AI Experiments tracks runs, parameters, and metrics for comparison, but does not expose a queryable lineage graph of artifact dependencies. Option C is wrong because Vertex AI Model Registry catalogs and versions models for deployment, not the full dataset-to-model lineage graph.

75
MCQeasy

An ML team wants to automatically track training runs, including hyperparameters and metrics, with minimal code changes. Which Vertex AI service should they use?

A.Vertex AI Prediction
B.Vertex AI Workbench
C.Vertex AI Metadata
D.Vertex AI Experiments with autologging
AnswerD

Autologging hooks into supported frameworks (for example scikit-learn, TensorFlow, XGBoost) and records parameters, metrics and artifacts to an Experiment run automatically, satisfying the minimal-code-change constraint. Manual logging via the SDK would require explicit calls in every training script.

Why this answer

Vertex AI Experiments with autologging automatically captures parameters, metrics, and artifacts from training runs with minimal code changes—often just a single call to initialize and start a run. This directly satisfies the requirement to track training runs with minimal code. It integrates with the Vertex AI SDK to log framework-specific details (e.g., TensorFlow, PyTorch) automatically.

Exam trap

PMLE often tests the confusion between Vertex AI Experiments (high-level run tracking with autologging) and Vertex ML Metadata (low-level lineage store), causing candidates to choose Metadata when minimal-code tracking is required.

How to eliminate wrong answers

Option A (Vertex AI Prediction) is wrong because it is for serving predictions from deployed models, not tracking training runs. Option B (Vertex AI Workbench) is wrong because it is a managed notebook environment for development, not a run-tracking service. Option C (Vertex AI Metadata) is wrong because it is the underlying lineage store, but it does not provide the high-level autologging convenience for experiments; you would have to log metadata manually.

Page 1 of 11

Page 2

All pages