Courseiva

Google Professional Machine Learning Engineer (PMLE) — Questions 751–775

775 questions total · 11pages · All types, answers revealed

Page 10

Page 11 of 11

751
MCQmedium

An ML team trains a model using a dataset stored in a BigQuery table. They want to ensure that the exact data snapshot used for training is recorded and can be reproduced later for auditing. Which approach should they take?

A.Use BigQuery table snapshots and log the snapshot ID as a parameter in the Vertex AI Pipeline run.
B.Schedule a daily export of the BigQuery table to Cloud Storage and use the latest export for training.
C.Export the BigQuery table to a Cloud Storage bucket and record the URI in Vertex AI Experiments.
D.Enable BigQuery audit logging and rely on the logs to reconstruct the data state.
AnswerA

BigQuery table snapshots provide a point-in-time, immutable copy of the table. By capturing the snapshot ID as a pipeline parameter, the training run is linked to the exact data version. This ensures reproducibility and auditability, as the snapshot can be restored or queried later. It directly addresses the need to record the data snapshot.

Why this answer

Using BigQuery table snapshots captures an immutable, point-in-time copy of the training data. Recording the snapshot ID in the pipeline run creates a direct link between the model and the exact data version, enabling reproducibility and audit compliance. Other methods either do not capture the precise data state or lack automatic linkage to the training run.

Exam trap

The trap here is assuming that exporting data or enabling audit logs provides a reproducible snapshot, when only a table snapshot guarantees an immutable, point-in-time copy.

752
MCQmedium

A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?

A.Cloud Composer
B.Dataproc
C.Cloud Functions
D.Dataflow
AnswerD

Dataflow runs Apache Beam pipelines that read from BigQuery, compute aggregates at scale, and write to Vertex AI Feature Store for both online serving and offline training. Its managed batch processing satisfies the daily large-scale computation requirement without provisioning servers.

Why this answer

Dataflow is Google Cloud's managed Apache Beam service, purpose-built for batch and streaming data pipelines that read from BigQuery, transform data, and write to sinks like Vertex AI Feature Store. It handles autoscaling, sharding, and windowing natively, making it the canonical choice for daily batch feature engineering at scale. Its native BigQuery and Feature Store I/O connectors mean minimal glue code.

Exam trap

PMLE often tests the boundary between orchestration (Cloud Composer/Airflow) and computation (Dataflow/Beam) — candidates who pick Composer because it 'runs pipelines' miss that it does not process the data itself.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is an orchestration service (managed Airflow) — it schedules and coordinates jobs but does not itself perform large-scale data transformation. Option B is wrong because Dataproc is a managed Spark/Hadoop cluster; while it can process BigQuery data, it requires cluster management and is better suited to Spark/Hadoop workloads than Beam-style pipelines. Option C is wrong because Cloud Functions is a serverless, event-driven compute service with short execution timeouts (up to 60 minutes for HTTP, 10 minutes for events) and is unsuitable for large-scale batch aggregation.

753
Multi-Selecteasy

A data analyst wants to use BigQuery ML to train a linear regression model (LINEAR_REG) to predict house prices. They have a table with features like square footage, number of bedrooms, and location. Which TWO statements about the training process are correct?

Select 2 answers
A.The analyst must call ML.TRAIN after CREATE MODEL to start training
B.The trained model is stored in Cloud Storage
C.The model must be exported to Vertex AI for prediction
D.The model is automatically evaluated on a held-out test set if data splitting is enabled
E.Training is performed using the CREATE MODEL statement
AnswersD, E

Enabling data splitting in CREATE MODEL reserves a portion of the input rows as a held-out test set, and BigQuery ML automatically computes evaluation metrics on it after training, satisfying the requirement to assess model quality without a separate manual step.

Why this answer

When data splitting is enabled in BigQuery ML, the `CREATE MODEL` statement automatically reserves a portion of the input data as a held-out test set. After training completes, BigQuery ML evaluates the model on this test set and reports metrics like mean absolute error and R², without requiring any manual split or separate evaluation step.

Exam trap

A common misconception is that BigQuery ML requires an explicit training command (like `ML.TRAIN`) or that models are stored in Cloud Storage by default, when in fact training is fully encapsulated in `CREATE MODEL` and models reside in BigQuery's internal storage.

754
MCQhard

A team is using Vertex AI Explainability with a deployed model. They need to generate explanations for image classification predictions. Which explanation method should they configure in the ExplanationSpec?

A.XRAI
B.SHAP with KernelExplainer
C.Sampled Shapley
D.Integrated Gradients
AnswerA

XRAI integrates gradients over regions rather than individual pixels, producing saliency maps grouped into coherent areas. For image classification, this regional attribution satisfies the need to highlight which image regions drove the predicted class, which pixel-level methods like Integrated Gradients express less intuitively.

Why this answer

XRAI (eXplainable AI) is Google Cloud's explanation method specifically designed for image classification models, producing saliency-based heatmaps that highlight regions of the image contributing to the prediction. It is the recommended method in Vertex AI ExplanationSpec for image data because it outperforms simpler gradient methods on visual tasks.

Exam trap

PMLE often tests which explanation method maps to which data modality — candidates confuse Sampled Shapley (tabular) with XRAI (image) or assume Integrated Gradients is always the default.

How to eliminate wrong answers

Option B is wrong because SHAP with KernelExplainer is a model-agnostic method suited to tabular data and is not natively supported by Vertex AI's ExplanationSpec for image classification. Option C is wrong because Sampled Shapley is Vertex AI's method for tabular and certain structured data, not for image classification. Option D is wrong because Integrated Gradients is a gradient-based attribution method supported for tabular and some image use cases, but XRAI is the method Google specifically recommends and optimizes for image classification.

755
MCQeasy

A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?

A.Delete the endpoint's service account permissions on Cloud Logging so the platform cannot write any log entries, including metrics.
B.Disable request-response logging on the deployed model while leaving endpoint monitoring and Cloud Monitoring metrics enabled.
C.Set the endpoint's traffic split to route all requests to a canary deployment that has logging disabled, while the primary deployment keeps logging enabled.
D.Enable request-response logging but set the sampling rate to a very small value so almost no payloads are captured.
AnswerB

Request-response logging is a separate, opt-in setting on the deployed model. Leaving it disabled means payloads are never captured, while the endpoint still emits standard metrics like request count, error rate, and latency percentiles to Cloud Monitoring, satisfying both the compliance constraint and the observability requirement.

Why this answer

Payload logging in Vertex AI is controlled by an explicit setting on each deployed model and is off unless enabled. Disabling it prevents request and response content from reaching Cloud Logging, while the endpoint continues to publish aggregate metrics such as request counts, error rates, and latency to Cloud Monitoring, which is exactly the separation this compliance scenario requires.

Exam trap

The trap here is conflating operational metrics with payload logging and assuming both must be turned off together to protect sensitive data.

756
MCQhard

A data engineer is troubleshooting a Vertex AI Endpoint that serves a large BERT model. After deployment, many prediction requests fail with 'Out of Memory' errors. The machine type is n1-standard-8 (30 GB memory) with no accelerator. Which action will most likely resolve the issue?

A.Change the machine type to n1-highmem-16 (104 GB memory).
B.Use batch prediction instead of online prediction.
C.Add a GPU accelerator (e.g., NVIDIA T4) to offload computation.
D.Quantize the model from FP32 to INT8.
AnswerA

The n1-standard-8's 30 GB cannot hold the large BERT model plus inference tensors, causing out-of-memory failures. Moving to n1-highmem-16 raises memory to 104 GB, satisfying the model's footprint without accelerators. More vCPUs alone would not increase memory proportionally.

Why this answer

The Out of Memory errors occur because the n1-standard-8 has only 30 GB RAM, which is insufficient to load the large BERT model plus runtime overhead. Switching to n1-highmem-16 provides 104 GB RAM, giving ample headroom for the model and inference batch, directly resolving the OOM condition.

Exam trap

PMLE often tests the misconception that adding a GPU fixes OOM errors, when OOM is a host memory issue and the correct fix is a high-memory machine type.

How to eliminate wrong answers

Option B is wrong because batch prediction does not solve the underlying memory constraint for online serving and changes the use case. Option C is wrong because adding a GPU offloads computation but does not increase host memory, and the OOM is a memory issue, not a compute issue. Option D is wrong because quantization reduces model size and memory footprint but may degrade accuracy and is not the most direct fix compared to simply provisioning adequate memory.

757
MCQeasy

An ML engineer needs to monitor the error rate of prediction jobs on a Vertex AI Endpoint. Where can they view the number of failed prediction requests over time?

A.Cloud Monitoring
B.Cloud Console endpoint details page
C.Vertex AI Experiments
D.Cloud Logging
AnswerA

Cloud Monitoring captures Vertex AI Endpoint request metrics, including the count of failed prediction requests, and plots them over time. This directly satisfies the requirement to view error rate trends, since prediction failures are surfaced as monitored metrics rather than logs.

Why this answer

Cloud Monitoring (formerly Stackdriver) is the Google Cloud service that collects, stores, and visualizes metrics such as prediction request counts and error rates for Vertex AI Endpoints. It provides dashboards, alerting, and time-series charts for failed prediction requests over time. This is the correct place to monitor error rate metrics.

Exam trap

PMLE often tests the distinction between Cloud Monitoring (metrics) and Cloud Logging (logs), so candidates who see 'failed prediction requests' and pick Cloud Logging forget that aggregated error rate over time is a metric, which belongs in Cloud Monitoring.

How to eliminate wrong answers

Option B is wrong because the Cloud Console endpoint details page shows configuration and some basic metrics, but it is not the primary monitoring tool for time-series error rate analysis and alerting. Option C is wrong because Vertex AI Experiments is for tracking training runs, hyperparameters, and model metrics — not for monitoring online prediction error rates. Option D is wrong because Cloud Logging stores log entries (including prediction errors), but it is not designed for aggregated metric visualization of error rates over time; that is Cloud Monitoring's role.

758
MCQmedium

Your PyTorch training script uses DistributedDataParallel (DDP) across 4 vertices each with 4 GPUs (16 GPUs total). You submit a Vertex AI custom training job. How should you configure the worker pool spec?

A.Create one worker pool with 4 replicas, each with machine type having 4 GPUs
B.Create a chief worker pool with 1 replica (4 GPUs) and a parameter server pool with 4 replicas (no GPUs)
C.Create 4 separate jobs, each with 1 replica and 4 GPUs
D.Create one worker pool with 16 replicas, each with 1 GPU
AnswerA

This matches the requirement: 4 workers, each with 4 GPUs.

Why this answer

For DDP across multiple machines, use MultiWorkerMirroredStrategy equivalent in PyTorch: set replicas to 4, each with machine type having 4 GPUs. The TF_CONFIG env var is not needed; Vertex AI sets necessary environment variables for distributed training.

759
MCQeasy

An ML team is moving from a prototype Jupyter notebook to a production training pipeline. They want to ensure reproducibility. Which approach should they take?

A.Use interactive parameter tuning.
B.Use a container with fixed dependencies and record hyperparameters.
C.Export the notebook's output model directly.
D.Save the notebook as a .py file.
AnswerB

Pinning dependencies inside a container image plus logging hyperparameters captures every input that affects training, so any run can be reproduced exactly. Notebooks alone leave library versions and random seeds unpinned, breaking reproducibility across environments.

Why this answer

Using a container with fixed dependencies and recording hyperparameters ensures that the training environment and configuration are captured, enabling exact reproduction. Option A is wrong because interactive parameter tuning is not reproducible—it introduces manual adjustments. Option C is wrong because exporting the notebook's output model directly lacks environment tracking and hyperparameter records.

Option D is wrong because saving the notebook as a .py file does not capture the full environment or dependencies.

760
MCQmedium

Your team is using Vertex AI Pipelines to train a model weekly. You want to monitor the pipeline for failures and receive a notification when a pipeline run fails. You have configured the pipeline to send logs to Cloud Logging. What should you do to receive an alert on pipeline failure?

A.Set up a Cloud Function that triggers on pipeline completion and sends an email if the status is failed.
B.Use Cloud Monitoring's built-in Vertex AI Pipeline failure metric to create an alert.
C.Configure the pipeline to publish a custom metric to Cloud Monitoring on failure and create an alert on that metric.
D.Create a log-based metric that counts error log entries and an alerting policy on that metric.
AnswerD

Vertex AI Pipelines logs pipeline execution events to Cloud Logging. By creating a log-based metric that filters for error-level logs from the pipeline, you can then create an alerting policy that triggers when the metric exceeds a threshold (e.g., >0 errors). This is the standard way to alert on pipeline failures using logs.

Why this answer

To alert on pipeline failures, you can create a log-based metric that filters for error logs from Vertex AI Pipelines and then create a Cloud Monitoring alerting policy on that metric. This leverages existing logs and integrates with Cloud Monitoring's alerting system without custom code.

Exam trap

The trap here is assuming there is a built-in pipeline failure metric in Cloud Monitoring, while in reality you must derive it from logs.

761
MCQmedium

A data scientist trains an XGBoost model on Vertex AI with a custom container. The model performs well on a held-out test set but fails to generalize in production. They suspect data leakage between training and validation. What is the best practice to prevent this?

A.Store and serve features using Vertex AI Feature Store with point-in-time correctness
B.Implement feature engineering in Vertex AI Pipelines to ensure temporal ordering
C.Store all features in BigQuery and join on timestamp during training and serving
D.Use Vertex AI AutoML instead of custom training
AnswerA

Vertex AI Feature Store with point-in-time correctness retrieves feature values as they existed at each training timestamp, preventing future information leaking into training rows. This directly addresses the leakage causing poor generalisation, which a held-out test set alone cannot expose.

Why this answer

Vertex AI Feature Store with point-in-time correctness ensures that for each training example, only feature values that were known at the time of the prediction (i.e., before the label occurred) are used. This prevents future data from leaking into the training set, which is the most common cause of poor generalization when temporal ordering matters. The Feature Store automatically retrieves the latest feature value as of a specified timestamp, eliminating the need for manual joins and windowing logic.

Exam trap

Google Cloud often tests the misconception that simply using a pipeline or a data warehouse with timestamps is sufficient to prevent leakage, but the key is the automated enforcement of point-in-time correctness, which only a dedicated feature store with time-travel capabilities provides.

How to eliminate wrong answers

Option B is wrong because implementing feature engineering in Vertex AI Pipelines ensures reproducible workflows but does not inherently enforce temporal ordering or prevent data leakage; pipelines can still join future features if the data is not time-aware. Option C is wrong because storing all features in BigQuery and joining on timestamp during training and serving is a manual approach that is error-prone and does not guarantee point-in-time correctness; it requires careful windowing logic and can still leak future data if the join is not correctly scoped. Option D is wrong because using Vertex AI AutoML does not automatically solve data leakage; AutoML models are equally susceptible to leakage if the training data contains future information, and the user still needs to ensure temporal integrity of the input features.

762
MCQeasy

A hospital wants to build a system that automatically transcribes doctors' dictated notes into text and then identifies key medical terms such as diagnoses and medications. They have no ML expertise and want to use Google Cloud's pre-trained APIs. Which combination of services should they use?

A.Speech-to-Text API for transcription and AutoML Natural Language for custom entity extraction.
B.Speech-to-Text API for transcription and Healthcare Natural Language API for medical term extraction.
C.Text-to-Speech API for transcription and Natural Language API for entity extraction.
D.Dialogflow for transcription and Healthcare Natural Language API for medical term extraction.
AnswerB

Speech-to-Text API converts audio to text accurately, and the Healthcare Natural Language API is specifically designed to extract medical entities like diagnoses and medications from text. This combination uses fully managed pre-trained models, requiring no ML expertise, and directly addresses both transcription and medical term identification in a single pipeline.

Why this answer

The hospital needs two capabilities: converting audio to text and extracting medical terms. Speech-to-Text API handles audio transcription with high accuracy, and Healthcare Natural Language API is purpose-built for medical text analysis, identifying entities like medications and diagnoses. Both are pre-trained, requiring no ML expertise, and integrate easily.

Other options use incorrect services for transcription or require custom model training, which is unnecessary given the availability of specialized pre-trained APIs.

Exam trap

The trap here is confusing Text-to-Speech with Speech-to-Text, or assuming that general Natural Language API can handle medical terminology as well as the specialized Healthcare Natural Language API.

763
MCQhard

A team wants to implement automated model documentation that captures training data, feature importance, evaluation metrics, and intended use. Which Vertex AI feature supports this?

A.Vertex AI Model Registry with model cards
B.Vertex AI Metadata
C.Vertex AI Explainable AI
D.Vertex AI Pipelines
AnswerA

Vertex AI Model Registry stores model cards alongside each version, capturing training data, feature importance, evaluation metrics and intended use. This satisfies the automated documentation requirement by centralising governance artefacts with the model lineage rather than in separate manual documents.

Why this answer

Vertex AI Model Registry supports model cards, which are structured documents that capture training data, feature importance, evaluation metrics, intended use, and limitations. Model cards are designed specifically for automated, standardized model documentation and governance. This makes Model Registry with model cards the correct feature for the stated requirement.

Exam trap

The trap is confusing metadata/lineage tracking (Vertex AI Metadata) or explainability (Explainable AI) with formal model documentation, which is specifically the model card feature in Model Registry.

How to eliminate wrong answers

Option B is wrong because Vertex AI Metadata stores lineage and artifact metadata for ML workflows (experiments, runs, artifacts), not human-readable model documentation like intended use and feature importance. Option C is wrong because Explainable AI provides feature attributions for predictions, not documentation of training data, metrics, and intended use. Option D is wrong because Vertex AI Pipelines orchestrates ML workflows; it does not itself produce model documentation artifacts.

764
MCQmedium

A model deployed on Vertex AI Prediction repeatedly exits with code 137. What is the most likely cause?

A.The model has a disk I/O bottleneck.
B.The model is using too much CPU.
C.The container image is incompatible with the machine type.
D.The model is using more memory than allocated (4GB).
AnswerD

Exit code 137 signals SIGKILL, typically delivered by the Linux OOM killer when a container exceeds its memory cgroup limit. The 4GB allocation is therefore being breached by the model's runtime footprint, making memory exhaustion — not CPU, timeout, or network failure — the constraint the stem describes.

Why this answer

Exit code 137 indicates that the container was killed by the Linux kernel's Out-Of-Memory (OOM) killer. In Vertex AI Prediction, each model deployment has a fixed memory allocation (default 4GB for custom containers). When the model's inference process exceeds this limit, the OOM killer terminates the container, resulting in exit code 137.

This is the most direct and common cause for this specific exit code in Vertex AI.

Exam trap

Google Cloud often tests the distinction between exit codes: candidates may confuse exit code 137 (OOM kill) with exit code 1 (generic error) or exit code 139 (segmentation fault), leading them to incorrectly attribute the issue to CPU or disk problems.

How to eliminate wrong answers

Option A is wrong because disk I/O bottlenecks typically cause slow performance or timeouts, not exit code 137 (SIGKILL from OOM). Option B is wrong because high CPU usage may cause throttling or latency, but does not trigger the OOM killer; exit code 137 is specifically memory-related. Option C is wrong because an incompatible container image would result in a different error, such as a crash loop with exit code 1 or 139 (segfault), not the OOM-specific exit code 137.

765
MCQmedium

A data scientist trained a model on a single GPU but needs to train on multiple GPUs for a larger dataset. They observe that training time does not decrease linearly with additional GPUs. Which common issue is most likely?

A.Overfitting.
B.Model architecture too simple.
C.Learning rate too high.
D.Data pipeline bottleneck.
AnswerD

Scaling GPUs only speeds up the compute graph; if the input pipeline cannot feed data fast enough, accelerators idle between steps. The bottleneck lies in data loading and preprocessing, so adding GPUs yields sublinear gains.

Why this answer

The most likely issue is a data pipeline bottleneck. When training on multiple GPUs, if the data input pipeline cannot feed data fast enough, the GPUs will idle waiting for data, leading to sublinear scaling. This is a common problem in distributed training where I/O or preprocessing becomes the limiting factor.

Exam trap

PMLE often tests the assumption that adding more GPUs always linearly reduces training time; candidates must consider Amdahl's law and identify bottlenecks like data loading or communication overhead.

How to eliminate wrong answers

Option A is wrong because overfitting affects model generalization, not training speed scaling. Option B is wrong because a simple model would train faster, not slower, and would not cause sublinear scaling. Option C is wrong because a high learning rate might cause divergence or slow convergence, but it does not directly explain why adding GPUs doesn't reduce training time.

766
MCQhard

An ML engineer is troubleshooting a Vertex AI Model Monitoring setup on a deployed model. The monitoring configuration uses a training dataset baseline and monitors several numerical features. After several days, the engineer notices that drift scores are being computed, but no alerts have fired even though one feature's distribution has shifted dramatically. The monitoring configuration specifies a drift threshold of 0.3, and the observed drift score for that feature is 0.45. What is the most likely explanation?

A.The training dataset baseline was not properly saved, so drift scores are computed against a default baseline.
B.The Cloud Monitoring alerting policy was not created or its condition does not match the drift metric.
C.The drift threshold of 0.3 is applied only to categorical features, not numerical ones.
D.Vertex AI Model Monitoring requires a minimum of 10,000 prediction requests before any alert can fire.
AnswerB

Vertex AI Model Monitoring computes drift scores and writes them as metrics, but alerts only fire if a Cloud Monitoring alerting policy exists and its condition evaluates the correct metric. If the policy is missing or misconfigured, no notification is sent even when the drift score exceeds the threshold. This is the most likely reason for the discrepancy.

Why this answer

Drift scores above threshold are necessary but not sufficient for alerts. Vertex AI Model Monitoring publishes metrics to Cloud Monitoring, and an alerting policy must be configured to watch those metrics and trigger notifications. If the policy is absent or its condition does not match the drift metric, no alert fires despite a high score.

The other options describe non-existent constraints or misattribute the cause.

Exam trap

The trap here is assuming that exceeding the monitoring configuration's drift threshold automatically triggers an alert, when Cloud Monitoring alerting policies are a separate requirement.

767
MCQeasy

A retail company wants to build a demand forecasting model for thousands of product SKUs. They have historical sales data in BigQuery and limited ML expertise. They want to minimize coding and automatically handle seasonality and promotions. Which approach should they use?

A.Use BigQuery ML with the ARIMA_PLUS model type and provide the time series data.
B.Use Dataflow to preprocess the data and then train a custom TensorFlow model on Vertex AI.
C.Use BigQuery ML with the LINEAR_REG model type and include date features.
D.Use Vertex AI AutoML Forecasting with the sales data exported to Cloud Storage.
AnswerA

BigQuery ML's ARIMA_PLUS is designed for univariate time series forecasting and automatically handles seasonality, holiday effects, and outliers. It requires only SQL, making it ideal for teams with limited ML expertise. It also supports large-scale training across many time series, such as product SKUs, and can incorporate covariates like promotions. This directly meets the company's need for minimal coding and automatic seasonality handling.

Why this answer

BigQuery ML's ARIMA_PLUS is purpose-built for time series forecasting and automatically addresses seasonality, holidays, and outliers without manual feature engineering. It operates entirely within BigQuery using SQL, aligning with the team's limited ML expertise and desire to minimize coding. It scales to many time series, making it suitable for thousands of SKUs.

Exam trap

The trap here is assuming that any BigQuery ML model can handle time series seasonality, when only specialized model types like ARIMA_PLUS provide that capability out of the box.

768
MCQmedium

A team deploys a model using Vertex AI Endpoint with automatic scaling. They observe that during traffic spikes, new instances take a long time to become ready, causing high latency for some requests. What should they configure to reduce this startup time?

A.Increase the max replicas
B.Use a custom container with a smaller footprint
C.Enable predictive autoscaling
D.Set a higher target CPU utilization
AnswerB

Instance startup time is dominated by pulling and initialising the container image. A smaller custom container reduces image size, so new replicas become ready faster during spikes, directly addressing the slow scale-out latency described in the stem.

Why this answer

A custom container with a smaller footprint reduces image pull time and container initialization overhead, which are the dominant contributors to Vertex AI Endpoint replica startup latency during scale-out. Smaller images pull faster from Artifact Registry and start faster, so new replicas become ready sooner and absorb traffic spikes with less queuing delay.

Exam trap

The trap here is conflating scaling policy knobs (max replicas, predictive autoscaling, target utilization) with startup-time reduction — only changes to the container/model itself shorten per-replica readiness time.

How to eliminate wrong answers

Option A is wrong because increasing max replicas only raises the ceiling on how many instances can run — it does nothing to reduce the time each new instance takes to become ready. Option C is wrong because predictive autoscaling forecasts traffic and pre-warms replicas ahead of demand, but it does not reduce per-replica startup time; it just shifts when scaling begins. Option D is wrong because raising target CPU utilization makes scaling less aggressive (replicas are added later), which would worsen latency during spikes rather than improve startup time.

769
MCQmedium

An engineer needs to perform sentiment analysis on customer reviews. They have a large volume of text and need a solution that requires minimal customisation. Which option is most efficient?

A.Use Vertex AI Prediction with a pre-trained model
B.Use BigQuery ML with LOGISTIC_REG
C.Train a custom model using AutoML NLP
D.Use the Natural Language API
AnswerD

The Natural Language API provides pre-trained sentiment analysis, satisfying the minimal-customisation constraint without labelled data or model training. It handles large text volumes through a managed endpoint, unlike custom model approaches that demand dataset preparation and tuning. This makes it the most efficient fit for the stated scenario.

Why this answer

The Natural Language API provides pre-built sentiment analysis with minimal setup. AutoML NLP would require custom training, BigQuery ML is for tabular data, and Vertex AI Prediction needs a deployed model.

770
MCQeasy

A retail company wants to predict customer churn using their transaction history and customer demographics. They have limited ML expertise and want to use a managed service on Google Cloud. Which service should they use?

A.AI Platform Notebooks
B.Vertex AI AutoML (Tables)
C.Cloud TPU
D.BigQuery ML
AnswerB

Vertex AI AutoML (Tables) trains tabular classification models on transaction and demographic data without requiring coding, satisfying the limited ML expertise constraint. Its managed pipeline handles feature engineering, model selection and tuning automatically, so the retail team can deploy churn predictions without building custom infrastructure.

Why this answer

Vertex AI AutoML (Tables) is the correct choice because it is a managed service specifically designed for tabular data, requiring no ML expertise. It automates model training, hyperparameter tuning, and deployment for classification tasks like churn prediction, directly handling transaction history and demographic features.

Exam trap

The trap here is that candidates may confuse BigQuery ML as a fully managed no-code solution, but it still requires SQL proficiency and manual model selection, whereas Vertex AI AutoML is the true zero-code managed service for tabular data.

How to eliminate wrong answers

Option A is wrong because AI Platform Notebooks provides a Jupyter-based development environment for custom ML coding, not a managed no-code solution, and requires ML expertise to build and train models. Option C is wrong because Cloud TPU is a hardware accelerator for training large deep learning models (e.g., NLP or vision), not a managed service for tabular churn prediction, and is overkill for this use case. Option D is wrong because BigQuery ML enables SQL-based model creation directly in BigQuery, but it requires some ML knowledge to write queries and tune models, and is less automated than AutoML for users with limited ML expertise.

771
Multi-Selectmedium

A financial services company has deployed a credit risk ML model on Vertex AI. They want to monitor the model for fairness across demographic groups to ensure no biased outcomes. Which TWO actions should they take as best practices? (Choose TWO.)

Select 2 answers
A.Eliminate all features that are correlated with protected attributes from the model input to ensure fairness.
B.Use Vertex Explainable AI to understand feature attributions and compare their distributions across demographic groups.
C.Periodically compare the model's performance metrics (e.g., AUC) on the overall population versus the holdout test set.
D.Store all model predictions in BigQuery but do not capture ground truth labels to avoid privacy issues.
E.Set up alerts on the Vertex AI Model Monitoring fairness metrics, such as equal opportunity difference, and configure a slack channel for notifications.
AnswersB, E

Feature attribution analysis helps identify if the model relies disproportionately on sensitive attributes.

Why this answer

Vertex Explainable AI provides feature attribution scores that can be compared across demographic groups to detect if the model relies on sensitive attributes or proxies. This enables fairness auditing by revealing whether the model's decision logic differs systematically for protected groups, which is a best practice for monitoring bias.

Exam trap

Google Cloud often tests the misconception that removing protected attributes or correlated features is sufficient for fairness, when in reality proxy features and complex interactions can still cause bias, making monitoring with explainability and fairness metrics essential.

772
MCQmedium

A company uses BigQuery to store feature data for ML training. A data engineer notices that a Vertex AI Training job is failing with 'Access Denied' errors when reading from a BigQuery table. The training job uses a custom service account that has been granted the 'bigquery.dataViewer' role on the dataset. What is the most likely cause of the failure?

A.The service account is not in the same project as the BigQuery dataset.
B.The BigQuery table is partitioned and requires row-level access.
C.The service account lacks the 'bigquery.jobs.create' permission in the project.
D.The training job does not have the required network access to BigQuery.
AnswerC

Reading a BigQuery table requires a query job, which needs bigquery.jobs.create in the project. The bigquery.dataViewer role on the dataset grants table-level read access only, so the custom service account cannot start the job and Vertex AI returns Access Denied.

Why this answer

The 'bigquery.dataViewer' role grants permissions to read BigQuery data (e.g., bigquery.tables.getData), but it does not include the 'bigquery.jobs.create' permission. When a Vertex AI training job reads from BigQuery, it must first create a BigQuery job (a query job) to retrieve the data. Without 'bigquery.jobs.create' at the project level, the service account cannot initiate the read operation, resulting in an 'Access Denied' error even though it has data-level access.

Exam trap

The trap here is that candidates often assume 'bigquery.dataViewer' is sufficient for all read operations, overlooking the requirement for 'bigquery.jobs.create' to initiate the query job that actually reads the data.

How to eliminate wrong answers

Option A is wrong because the service account does not need to be in the same project as the BigQuery dataset; cross-project access is supported as long as IAM permissions are granted at the dataset or table level. Option B is wrong because partitioned tables do not require row-level access by default; row-level access is controlled via BigQuery row-level security policies, which are not automatically required for partitioned tables. Option D is wrong because Vertex AI training jobs run within Google Cloud's internal network and have built-in access to BigQuery via the Cloud API; network access is not a common cause of 'Access Denied' errors for BigQuery reads.

773
Multi-Selecthard

Which TWO statements are true about canary deployments for Vertex AI endpoints?

Select 2 answers
A.Canary deployments are only supported for custom containers, not prebuilt frameworks.
B.You can roll back a canary by resetting traffic to 0% for the new version.
C.You can use traffic splitting to gradually shift 1-100% of traffic to a new version.
D.Canary deployments require the use of Vertex AI Model Registry.
E.Once a canary receives 50% traffic, you cannot increase it further.
AnswersB, C

Traffic splitting is reversible: setting the new version's split to 0% routes all requests back to the stable version, effectively rolling back the canary. This satisfies the rollback requirement without redeploying or deleting the endpoint.

Why this answer

Option B is correct because in Vertex AI traffic splitting, a canary version is simply a DeployedModel receiving a percentage of prediction traffic; setting that version's traffic split to 0% (and returning 100% to the stable version) effectively rolls back the canary without undeploying it. Option C is correct because Vertex AI endpoints support traffic splitting where you assign each deployed model a percentage from 0 to 100, letting you gradually shift traffic from the existing version to the new version in controlled increments. Option A is wrong because canary/traffic splitting works with any deployed model on an endpoint, including AutoML and prebuilt-framework models, not only custom containers.

Option D is wrong because traffic splitting operates on DeployedModels on an endpoint and does not require the model to be registered in Vertex AI Model Registry. Option E is wrong because traffic percentages are configurable at any time up to 100%, so a canary at 50% can be increased further.

Exam trap

PMLE often tests the misconception that canary deployments require specific model types or the Model Registry; candidates may incorrectly believe that prebuilt frameworks cannot use traffic splitting or that the Model Registry is mandatory.

774
MCQmedium

You are using TensorFlow Transform (tf.Transform) to preprocess data for a model that will be deployed on Vertex AI. What is the primary benefit of using tf.Transform over Dataflow alone?

A.Support for GPUs during preprocessing
B.Built-in feature store integration
C.Faster data processing
D.Training/serving skew prevention through a consistent transformation graph
AnswerD

tf.Transform records the full preprocessing graph and applies it identically at training and prediction, eliminating training/serving skew. Dataflow alone executes transforms but does not preserve that reusable graph, so the same feature engineering cannot be replayed consistently at serving time on Vertex AI.

Why this answer

tf.Transform computes statistics (e.g., min, max) on the full dataset, then generates a TensorFlow graph that applies the same transformation consistently at training and serving time. Dataflow alone does not ensure this consistency.

775
MCQeasy

A machine learning engineer wants to define a lightweight pipeline component that runs custom Python code without building a container image. Which KFP SDK feature should they use?

A.Importer component
B.Python function component with @dsl.component
C.Container component
D.Vertex AI Training job
AnswerB

The @dsl.component decorator converts a plain Python function into a lightweight pipeline component, letting KFP build the container automatically. This satisfies the stem's requirement to run custom Python code without manually building a container image.

Why this answer

The `@dsl.component` decorator in KFP SDK allows you to define a lightweight Python function component that runs custom code without requiring a container image. It automatically generates a container specification from the function's dependencies, making it ideal for simple, non-containerized pipeline steps.

Exam trap

The trap here is that candidates may confuse 'lightweight' with 'no container at all,' but KFP always runs components in containers; the `@dsl.component` feature automates container creation, not eliminates it.

How to eliminate wrong answers

Option A is wrong because the Importer component is used to import existing artifacts (like datasets or models) into a pipeline, not to run custom Python code. Option C is wrong because a Container component requires you to specify a pre-built container image, which contradicts the requirement of not building a container image. Option D is wrong because Vertex AI Training job is a managed service for running training jobs on Vertex AI, not a lightweight KFP SDK feature for running custom code without containers.

Page 10

Page 11 of 11

All pages