Courseiva

CCNA Pmle Scaling Models Questions

75 of 103 questions · Page 1/2 · Pmle Scaling Models topic · Answers revealed

1
MCQmedium

You are training a scikit-learn model on Vertex AI using a custom training job. The training dataset is a 2 TB CSV file stored in Cloud Storage, and the job must run on a single CPU-only VM. Loading the entire file into memory fails because the machine has only 32 GB of RAM. You need to train the model without increasing the VM size and without rewriting the training code to use a distributed framework. What should you do?

A.Use Vertex AI Pipelines to split the CSV into many smaller files, then run a separate training job for each file and average the model weights.
B.Convert the CSV to TFRecord format and use tf.data to stream batches, then train the scikit-learn model on the streamed batches.
C.Use the scikit-learn partial_fit method on an estimator that supports incremental learning, reading the CSV in chunks with pandas.read_csv and feeding each chunk to partial_fit.
D.Mount the Cloud Storage bucket as a local filesystem on the training VM and call pandas.read_csv on the mounted path, relying on the OS page cache to keep memory usage low.
AnswerC

Many scikit-learn estimators such as SGDClassifier and MiniBatchKMeans implement partial_fit, which updates the model incrementally on small batches. Reading the 2 TB CSV with pandas.read_csv in chunks keeps only one chunk in memory at a time, so the 32 GB VM is sufficient. This preserves a single model trained over all rows and requires no distributed framework, which matches the constraint exactly.

Why this answer

Incremental learning with partial_fit is the correct approach because it lets a scikit-learn estimator update its parameters one chunk at a time. Reading the CSV in chunks with pandas.read_csv bounds memory to the chunk size rather than the full 2 TB. This keeps the single-model semantics, avoids distributed training, and fits within the 32 GB VM, directly satisfying all the stated constraints.

Exam trap

The trap here is assuming that changing the file format or mounting Cloud Storage reduces the memory required by scikit-learn, when the real fix is an estimator that supports incremental partial_fit updates.

2
MCQmedium

Your Vertex AI custom training job is failing with an out-of-memory error on a single GPU. You need to reduce memory usage without changing the model architecture. Which approach should you try first?

A.Decrease the batch size
B.Implement model parallelism across GPUs
C.Use gradient accumulation
D.Enable mixed precision training (FP16)
AnswerA

Reducing the batch size lowers the number of samples held in GPU memory per step, directly cutting activation and gradient memory consumption. It requires no architecture change, satisfying the stem's constraint, and is the least invasive first remedy before considering gradient checkpointing or mixed precision.

Why this answer

Decreasing the batch size directly reduces the memory required for activations and gradients, making it the simplest and most immediate way to resolve out-of-memory errors without altering the model architecture. It is a standard first step because it requires no code changes beyond a hyperparameter and often resolves OOM with minimal impact on convergence if adjusted with learning rate.

Exam trap

PMLE often tests the order of operations for troubleshooting OOM, trapping candidates who jump to advanced techniques like mixed precision or model parallelism instead of the simplest, most direct fix: reducing batch size.

How to eliminate wrong answers

Option B is wrong because model parallelism across GPUs is complex, requires significant code changes, and is not the first thing to try for a single-GPU OOM. Option C is wrong because gradient accumulation increases memory usage by storing gradients for multiple steps, which would worsen OOM. Option D is wrong because mixed precision training (FP16) can reduce memory, but it requires code changes and may introduce numerical instability; it is not the first, simplest step.

3
Multi-Selecthard

You are preparing to train a large image classification model on Vertex AI using a custom training job. You want to optimize the training job for cost and performance. The dataset is stored in Cloud Storage as TFRecords and is about 2 TB. You plan to use a machine with 4 NVIDIA V100 GPUs. Which two actions should you take to improve training efficiency? (Choose two.)

Select 2 answers
A.Use a smaller model architecture to reduce training time.
B.Store the TFRecords in a regional Cloud Storage bucket instead of a multi-regional bucket.
C.Use a larger batch size and scale the learning rate accordingly.
D.Enable mixed precision training using NVIDIA Apex or TensorFlow's mixed precision API.
E.Increase the number of vCPUs on the machine to improve data loading throughput.
AnswersC, D

Increasing batch size can improve GPU utilization and reduce training time by processing more samples per iteration. Scaling the learning rate helps maintain convergence. This is a standard technique for efficient training on multiple GPUs. It directly leverages the available GPU memory and compute, leading to better throughput and cost efficiency.

Why this answer

Mixed precision training and increasing batch size with learning rate scaling are both effective ways to improve training efficiency on V100 GPUs. Mixed precision leverages tensor cores for faster computation and reduced memory, while larger batches improve GPU utilization. These actions directly target performance and cost without altering the model architecture or data storage.

Exam trap

The trap here is focusing on infrastructure changes like bucket location or more vCPUs, when the most impactful optimizations are at the training algorithm level, such as precision and batch size.

4
Multi-Selectmedium

A data science team is building a real-time feature engineering pipeline for ML model training and serving. They need to compute features from streaming data, store them for low-latency serving, and ensure consistency between training and serving. Which TWO Google Cloud services should they use?

Select 2 answers
A.Vertex AI Feature Store
B.BigQuery
C.Cloud Functions
D.Cloud Dataflow
E.Cloud SQL
AnswersA, D

Vertex AI Feature Store provides a centralised repository for feature values, enabling low-latency online serving alongside consistent offline retrieval for training. It directly satisfies the stem's requirement for training-serving consistency and low-latency access, ingesting features computed from streaming data without duplicating logic between pipelines.

Why this answer

Vertex AI Feature Store (A) is correct because it provides a centralized repository for storing, serving, and sharing feature data with low-latency online serving and batch serving for training, ensuring consistency between training and serving through point-in-time lookups and feature value time-stamping. Cloud Dataflow (D) is correct because it is a fully managed stream and batch processing service based on Apache Beam, enabling real-time feature engineering from streaming data with exactly-once processing semantics and automatic scaling.

Exam trap

A common trap in Google PMLE exams is assuming BigQuery can serve as a low-latency online feature store for real-time inference, but it is designed for analytical queries with seconds-to-minutes latency, not sub-millisecond serving required for real-time ML inference.

5
MCQhard

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline includes a step that trains a model and a subsequent step that evaluates the model. The engineer wants to ensure that the evaluation step runs only if the training step succeeds and that the pipeline fails if the model's accuracy is below a threshold. Which approach should the engineer use?

A.Set the evaluation step's retry policy to zero and use the pipeline's built-in accuracy threshold parameter.
B.Use a condition in the pipeline that checks the training step's status and a custom component that raises an exception if accuracy is below the threshold.
C.Use a Vertex AI Model resource to store the model and enable Model Monitoring to trigger a pipeline failure if accuracy drops.
D.Configure the evaluation step to always run and rely on Vertex AI Pipelines to automatically fail the pipeline if the accuracy metric is below the threshold.
AnswerB

Vertex AI Pipelines supports conditions to control step execution based on the status of previous steps. A custom component can evaluate the model and raise an exception if the accuracy does not meet the threshold. This exception will cause the pipeline to fail, ensuring that only models meeting the criteria proceed. This approach provides the required control flow and failure behavior.

Why this answer

To conditionally run a step and fail the pipeline based on a metric, the engineer should use a condition to check the training step's status and a custom component that raises an exception if accuracy is too low. Vertex AI Pipelines conditions allow steps to run only if previous steps succeed, and raising an exception in a component causes the pipeline to fail. This combination provides the required control flow and failure semantics.

Exam trap

The trap here is assuming that Vertex AI Pipelines automatically fails on low accuracy metrics or that Model Monitoring can be used within a pipeline for this purpose.

6
MCQmedium

A team is scaling a prototype ML model to production on Vertex AI. The model was developed using scikit-learn and requires custom preprocessing. They want to minimize operational overhead and ensure consistency between training and serving. Which approach should they use?

A.Train on a local machine and upload the model artifacts to Cloud Storage, then create an endpoint with a pre-built container.
B.Use a pre-built Vertex AI container for scikit-learn and provide a custom training Python package with preprocessing code included.
C.Deploy the model as a custom prediction routine on Vertex AI Endpoints with a custom container.
D.Export the model as a .pkl file and use Vertex AI's 'Import Model' with a default container for inference.
AnswerB

Packaging preprocessing inside a custom training Python package lets Vertex AI's pre-built scikit-learn container run both training and prediction, so the same code executes at serving time. This satisfies the consistency constraint directly, while the managed container removes the operational overhead of building and maintaining your own image.

Why this answer

Using a pre-built Vertex AI container for scikit-learn with a custom training Python package ensures that the same preprocessing code runs during both training and serving, minimizing operational overhead. This approach leverages Vertex AI's managed infrastructure to handle scaling, monitoring, and consistency without requiring custom container maintenance.

Exam trap

The trap here is that candidates often assume a pre-built container cannot handle custom preprocessing, leading them to choose a custom container (Option C) or a simpler import (Option D), but Vertex AI allows embedding preprocessing in the training package or model artifact to maintain consistency with minimal overhead.

How to eliminate wrong answers

Option A is wrong because training on a local machine and uploading model artifacts to Cloud Storage, then creating an endpoint with a pre-built container, does not guarantee consistency between training and serving preprocessing logic, as the preprocessing code is not bundled with the model. Option C is wrong because deploying the model as a custom prediction routine with a custom container introduces unnecessary operational overhead for a scikit-learn model that can be served with a pre-built container, and it requires building and maintaining a custom Docker image. Option D is wrong because exporting the model as a .pkl file and using Vertex AI's 'Import Model' with a default container for inference does not include custom preprocessing code, leading to potential inconsistencies between training and serving.

7
MCQhard

A team is training a TensorFlow model on Vertex AI using a custom container. The training script writes checkpoints to a local directory inside the container. The job runs for 14 hours, and when it completes, the team cannot find the checkpoints in Cloud Storage. They need the checkpoints to be persisted so they can resume training and deploy the best model. What should they do?

A.Enable Vertex AI TensorBoard integration and configure the training script to log checkpoints as TensorBoard artifacts, then retrieve them from the TensorBoard instance.
B.Set the training job's base output directory to a local path and rely on Vertex AI to automatically upload everything under that path to the job's Cloud Storage output directory at the end of training.
C.Increase the boot disk size of the training VM and re-run the job, then copy the checkpoints from the boot disk to Cloud Storage after training finishes.
D.Configure the training job to write checkpoints to a Cloud Storage URI by passing a gs:// path to the checkpoint directory in the training script, and ensure the Vertex AI service account has storage.objectAdmin on the bucket.
AnswerD

Writing checkpoints directly to a gs:// URI makes TensorFlow use its GCS filesystem implementation, so checkpoint files are persisted in Cloud Storage as they are written. The Vertex AI custom training service account must have permission to write to the bucket, which storage.objectAdmin grants. This is the standard pattern for durable checkpoints in Vertex AI training and directly solves the missing-checkpoint problem.

Why this answer

Checkpoints must be written to a durable location, and in Vertex AI custom training the durable location is Cloud Storage. Passing a gs:// URI to the checkpoint directory makes TensorFlow persist each checkpoint through its GCS filesystem, and granting storage.objectAdmin to the training service account authorizes those writes. This allows both resuming training and retrieving the best model for deployment.

Exam trap

The trap here is assuming that Vertex AI automatically uploads local training artifacts to Cloud Storage, when only files written to a gs:// path are persisted.

8
Multi-Selecthard

A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?

Select 3 answers
A.Using Vertex AI Vizier to optimize the model parallelism strategy
B.Enabling Vertex AI AutoML to automatically distribute the model
C.Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe
D.A custom container with the distributed framework (e.g., PyTorch DDP) installed
E.Setting the --worker-machine-count flag when submitting the job
AnswersC, D, E

Manual implementation of pipeline parallelism is required as Vertex AI does not provide built-in model parallelism.

Why this answer

Pipeline parallelism requires explicit implementation in the training script, such as using `torch.distributed.pipeline.sync.Pipe` in PyTorch, to split the model layers across multiple GPUs. This is necessary when the model does not fit on a single GPU, and Vertex AI does not automatically handle model parallelism—it must be coded by the user.

Exam trap

This question tests the misconception that Vertex AI automatically handles model parallelism (e.g., via AutoML or Vizier), when in reality the user must manually implement it in the training script using frameworks like PyTorch or TensorFlow.

9
MCQeasy

You are deploying a scikit-learn model to Vertex AI for online prediction. The model was trained on a dataset with numerical features and expects input data in a specific JSON format. You have created a custom container that serves the model using a Flask app. After deploying the model to a Vertex AI Endpoint, you send a prediction request with a JSON payload, but the response is an error indicating that the input format is invalid. What is the most likely reason for this error?

A.The custom container's Flask app does not parse the incoming JSON correctly because it expects a different key or structure than what was sent.
B.The Vertex AI Endpoint requires the input data to be base64-encoded.
C.The Vertex AI Endpoint automatically converts the input to a TensorFlow tensor, which is incompatible with scikit-learn models.
D.The model artifact was not uploaded to Cloud Storage, so the container cannot load the model and returns an input format error.
AnswerA

When you deploy a custom container, the prediction request payload is passed to your container as-is. If your Flask app expects a specific JSON schema (e.g., a key named 'instances' or 'data'), but you send a different structure, it will fail to parse the input. This is a common mistake when the container's preprocessing logic does not match the request format.

Why this answer

The correct answer is that the custom container's Flask app does not parse the incoming JSON correctly. Vertex AI passes the raw request to the container, so the container must implement logic to extract and format the input as expected by the model. If the JSON structure sent by the client does not match what the container expects, an input format error occurs.

Ensuring alignment between the client request and container preprocessing is essential.

Exam trap

The trap here is assuming that Vertex AI automatically transforms the input for scikit-learn, when in fact the custom container is fully responsible for parsing and preprocessing the request.

10
MCQhard

An ML engineer is training a very large PyTorch model on Vertex AI using a TPU v3 pod. The training is slower than expected, and the TPU utilization is low. What is the most likely cause?

A.The data pipeline is a bottleneck; the TPU is waiting for data.
B.The learning rate schedule is too aggressive.
C.The model is using a single TensorFlow operation not supported by TPU.
D.The batch size is too large for the TPU memory.
AnswerA

TPUs stall when input tensors arrive slower than the accelerator can consume them, so low utilisation with slow training points to host-side starvation. Feeding data via tf.data with parallel extraction and prefetching keeps the TPU v3 pod saturated rather than idle waiting on I/O.

Why this answer

The most likely cause of low TPU utilization is a data pipeline bottleneck, where the TPU spends a significant amount of time idle waiting for the next batch of data to be loaded and preprocessed. TPU v3 pods are designed for high-throughput matrix operations and can process data far faster than a typical CPU-based data loader can supply it, especially if the data pipeline uses inefficient I/O, lacks prefetching, or has insufficient workers. This mismatch starves the TPU, leading to low utilization and slower training.

Exam trap

Google often tests the misconception that low utilization is caused by model architecture or hyperparameter issues, when in reality the most common bottleneck in distributed TPU training is the data pipeline, not the compute or memory limits.

How to eliminate wrong answers

Option B is wrong because an aggressive learning rate schedule may cause training instability or divergence, but it does not directly cause low TPU utilization; utilization is a measure of hardware activity, not training convergence. Option C is wrong because the question explicitly states the model is a PyTorch model, and while PyTorch has limited TPU support compared to TensorFlow, the issue is not a single unsupported operation—such an operation would typically raise an error or fall back to CPU, not cause low utilization across the entire pod. Option D is wrong because a batch size that is too large for TPU memory would cause an out-of-memory (OOM) error, not low utilization; the TPU would fail to allocate the batch, not run slowly.

11
Multi-Selectmedium

A company wants to train a custom machine learning model on Vertex AI using a pre-built container for scikit-learn. They want to use spot VMs to reduce costs. However, the training job fails intermittently due to preemption. Which TWO actions should they take to ensure the training job completes successfully?

Select 2 answers
A.Use a larger machine type to reduce training time
B.Increase the number of parallel trials in hyperparameter tuning
C.Set the worker_pool_specs to use spot VMs by setting spot=True
D.Set the max_retry_count in the worker pool spec to a value greater than 0
E.Implement checkpointing in the training code to save model state periodically to Cloud Storage
AnswersD, E

Vertex AI will retry the job if preempted up to max_retry_count times.

Why this answer

To handle spot VM preemptions, the training job must be restartable. Using checkpoints allows the job to resume from the last saved state. Vertex AI automatically retries on preemption if the job is restartable (managed by the service).

Setting max_retry_count in the worker pool spec allows Vertex AI to automatically restart the job after preemption. Also, reducing machine type or increasing parallel trials are not direct solutions.

12
MCQmedium

A company wants to use Vertex AI Vizier to tune hyperparameters for a PyTorch model. They have a limited budget of 50 training jobs. The objective metric is validation accuracy, and they want to find the best configuration efficiently. Which algorithm should they choose?

A.Bayesian optimization using Vertex AI Vizier.
B.Random search with 50 random configurations.
C.Use a custom algorithm implemented in the training code.
D.Grid search with 50 evenly spaced points.
AnswerA

Bayesian optimisation builds a probabilistic surrogate model of validation accuracy and selects each configuration to maximise expected improvement, converging in far fewer trials than grid or random search. This suits the 50-job budget while targeting the stated objective metric.

Why this answer

Vertex AI Vizier is a hyperparameter tuning service that uses Bayesian optimization to efficiently search the hyperparameter space. With a limited budget of 50 trials, Bayesian optimization is more sample-efficient than random or grid search, making it the best choice.

Exam trap

The trap is underestimating the efficiency of Bayesian optimization. Candidates might think random search is sufficient with 50 trials, but Bayesian optimization is designed for sample efficiency, which is critical when the budget is limited.

How to eliminate wrong answers

Option B is wrong because random search does not use information from previous trials and is less efficient with a small budget. Option C is wrong because implementing a custom algorithm in the training code would require significant effort and may not be as optimized as Vizier. Option D is wrong because grid search with 50 points is a brute-force approach that does not adapt to the results of previous trials and is inefficient for high-dimensional spaces.

13
Multi-Selectmedium

An ML engineer is using Vertex AI for distributed training of a PyTorch model across multiple nodes. The training job must use TPUs for high throughput. The engineer sets up the job configuration. Which THREE components are required for the training to work correctly? (Select 3)

Select 3 answers
A.A startup script to configure the TPU pod (e.g., `xla_lib.sh`)
B.A MultiWorkerMirroredStrategy configuration
C.A Docker image that includes PyTorch and the TPU library (torch-xla)
D.A TF_CONFIG environment variable set for each worker
E.A CustomJob with a TPU accelerator type (e.g., v3-32)
AnswersA, C, E

Startup scripts are often needed to initialize TPU devices.

Why this answer

A is correct because TPU pods require a startup script (e.g., `xla_lib.sh`) to initialize the XLA runtime, configure the TPU mesh, and set environment variables like `XRT_TPU_CONFIG`. Without this script, the TPU devices will not be discoverable by the PyTorch/XLA process, causing the training to fail with device-not-found errors.

Exam trap

Google Cloud often tests the distinction between TensorFlow and PyTorch distributed training configurations, and the trap here is assuming that `TF_CONFIG` or `MultiWorkerMirroredStrategy` are universal for all frameworks, when in fact PyTorch uses its own environment variables and the `torch-xla` library for TPU training.

14
Multi-Selectmedium

You are preparing to deploy a trained scikit-learn model to Vertex AI for online prediction. You need to create a custom container that serves the model. Which two of the following steps are required to ensure the container works with Vertex AI? (Choose two.)

Select 2 answers
A.Ensure the container runs as a non-root user for security.
B.Include the model artifacts in the container image or ensure they are accessible at runtime.
C.Implement a web server that listens on the port specified by the AIP_HTTP_PORT environment variable.
D.Expose a /health endpoint that returns a 200 status code.
E.Use the Vertex AI SDK to build and push the container image.
AnswersB, C

The container must have access to the trained model artifacts to perform predictions. You can either bake the model into the image or mount it from a Cloud Storage location at runtime. Without the model, the container cannot serve predictions. Vertex AI expects the model to be loaded when the container starts, so providing the artifacts is essential for a functional deployment.

Why this answer

For a custom container to work with Vertex AI, it must listen on the port specified by the AIP_HTTP_PORT environment variable and have access to the model artifacts. These are the two essential steps. Other aspects like health endpoints, SDK usage, or non-root user are optional or best practices but not mandatory for the container to serve predictions on Vertex AI.

Exam trap

The trap here is assuming that a /health endpoint or non-root user is required, when the actual requirements are listening on the correct port and providing model artifacts.

15
MCQeasy

A data scientist has trained a TensorFlow model locally and wants to deploy it to Vertex AI for online predictions. The model accepts a single input tensor of shape (1, 224, 224, 3) and outputs a probability distribution over 10 classes. The data scientist wants to minimize deployment effort and ensure the model is served with low latency. What is the simplest way to deploy this model on Vertex AI?

A.Convert the model to ONNX format and use a custom container with ONNX Runtime for serving.
B.Use Vertex AI Batch Prediction with a pre-built TensorFlow container, and then set up a Cloud Function to serve online requests.
C.Export the model as a SavedModel and upload it to Vertex AI Model Registry, then deploy to an endpoint with a pre-built TensorFlow Serving container.
D.Package the model in a Docker container with a Flask app that loads the model and exposes a REST API, then deploy as a custom container on Vertex AI.
AnswerC

Vertex AI supports deploying TensorFlow SavedModels directly using a pre-built TensorFlow Serving container. This requires no custom code, and the serving container handles the model signature. It is the simplest and most efficient way to deploy a standard TensorFlow model for online predictions with low latency.

Why this answer

The simplest and most efficient way to deploy a standard TensorFlow model on Vertex AI for online predictions is to export it as a SavedModel and use the pre-built TensorFlow Serving container. This requires no custom serving code, leverages Vertex AI's managed infrastructure, and provides low-latency predictions. It also integrates with Vertex AI Model Registry for versioning and monitoring.

Exam trap

The trap here is overcomplicating the deployment by considering custom containers or format conversions when a pre-built container for TensorFlow is readily available.

16
Multi-Selecthard

Your team is deploying a large model on edge devices and needs to reduce its size by 80% while maintaining reasonable accuracy. Which THREE techniques should they consider? (Choose 3.)

Select 3 answers
A.Quantisation to INT8
B.Transfer learning from a larger model
C.Knowledge distillation
D.Increasing model capacity with more layers
E.Pruning of redundant connections
AnswersA, C, E

Reduces model size by reducing precision of weights.

Why this answer

Quantisation to INT8 reduces the precision of model weights and activations from 32-bit floating point to 8-bit integers, cutting memory usage by approximately 75% (4x compression). This directly addresses the 80% size reduction target while often preserving accuracy within 1-2% through careful calibration and scaling, making it a primary technique for edge deployment.

Exam trap

Google Cloud often tests the misconception that transfer learning reduces model size, when in fact it only transfers learned features and does not compress the model; candidates may confuse it with knowledge distillation.

17
MCQmedium

A data scientist wants to train a PyTorch model on Vertex AI using a pre-built container for GPU training. She needs to use 4 NVIDIA A100 GPUs on a single machine. Which machine configuration should she select?

A.n1-highmem-16 with 4 NVIDIA V100 GPUs
B.n1-standard-16 with 4 NVIDIA T4 GPUs
C.a2-highgpu-4g (4 A100 GPUs)
D.a2-megagpu-16g (16 A100 GPUs)
AnswerC

The a2-highgpu-4g machine type provides exactly four NVIDIA A100 GPUs on a single node, matching the stem's requirement for four A100s on one machine. This satisfies the GPU count and single-machine constraint, letting the pre-built PyTorch container train without custom configuration or distributed multi-node setup.

Why this answer

The a2-highgpu-4g machine type is purpose-built for GPU workloads and provides exactly 4 NVIDIA A100 GPUs on a single machine, matching the requirement precisely. Vertex AI supports this machine type for custom training with pre-built containers, so the data scientist can select it directly in the training configuration. Choosing it avoids over-provisioning or mismatching GPU counts.

Exam trap

PMLE often tests machine-type specificity — candidates pick a generic n1 machine with the right GPU count, missing that A100 GPUs require the A2 machine family, not N1.

How to eliminate wrong answers

Option A is wrong because n1-highmem-16 with 4 V100 GPUs uses older V100 GPUs, not the requested A100s, and n1 machines are general-purpose, not GPU-optimized. Option B is wrong because n1-standard-16 with 4 T4 GPUs provides T4 GPUs, which are inference-oriented and far less powerful than A100s for training. Option D is wrong because a2-megagpu-16g provides 16 A100 GPUs, four times the requested count, wasting cost and exceeding the stated requirement.

18
Multi-Selecthard

An engineer is designing a distributed training job on Vertex AI for a TensorFlow model that uses the MultiWorkerMirroredStrategy. They need to ensure proper communication between workers. Which environment variable must be set correctly for each worker?

Select 1 answer
A.CLUSTER_SPEC
B.TF_CPP_MIN_LOG_LEVEL
C.TF_CONFIG_JSON
D.TF_DISTRIBUTED_STRATEGY
E.TF_CONFIG
AnswersE

TF_CONFIG is the environment variable that carries the cluster specification and task details to each worker, letting MultiWorkerMirroredStrategy identify the chief and workers and establish collective communication. Without it correctly set per worker, distributed training on Vertex AI cannot coordinate.

Why this answer

In TensorFlow distributed training with MultiWorkerMirroredStrategy, the only required environment variable is `TF_CONFIG`. It provides the cluster topology and task identity, enabling gRPC communication between workers. The distribution strategy is defined in code, not via an environment variable. `TF_DISTRIBUTED_STRATEGY` is not a standard TensorFlow environment variable.

Exam trap

The exam may confuse candidates with plausible but incorrect environment variable names like TF_CONFIG_JSON or TF_DISTRIBUTED_STRATEGY, but only TF_CONFIG is required.

19
MCQmedium

You are deploying a scikit-learn model to Vertex AI for online predictions. The model requires a custom preprocessing step that transforms raw JSON input into a feature vector before calling predict. You want to minimize latency and avoid managing infrastructure. What should you do?

A.Export the scikit-learn model to a TensorFlow SavedModel and deploy it with the TensorFlow pre-built container, using a preprocessing layer in the model.
B.Deploy the scikit-learn model using the pre-built scikit-learn container and implement preprocessing in a Cloud Function that calls the endpoint.
C.Create a custom container that includes the scikit-learn model and a Flask app that performs preprocessing and calls predict, then deploy it to a Vertex AI endpoint.
D.Use Vertex AI Batch Prediction with a pre-built scikit-learn container and a custom preprocessing script.
AnswerC

Vertex AI supports custom containers for prediction, allowing you to package the model, preprocessing logic, and a web server. By building a container with a Flask app that implements the preprocessing and prediction logic, you can deploy it to a Vertex AI endpoint. This approach minimizes latency because preprocessing and prediction happen in the same process, and Vertex AI manages the infrastructure, including scaling and health checks.

Why this answer

Vertex AI custom containers for prediction allow you to bundle the scikit-learn model, custom preprocessing code, and an HTTP server such as Flask. This keeps preprocessing and inference in the same process, reducing latency, and Vertex AI handles scaling and infrastructure management. Pre-built containers do not support arbitrary preprocessing, and batch prediction or model conversion would not meet the low-latency online requirement.

Exam trap

The trap here is assuming that a pre-built scikit-learn container can execute custom preprocessing, when in fact it only serves the model's predict method and requires raw inputs to be already preprocessed.

20
MCQmedium

An ML team is using Vertex AI to train a deep learning model on a large dataset. To reduce costs, they want to use preemptible VMs for training jobs. However, training must complete within a bounded time. Which strategy should they use?

A.Use Cloud TPU instead of GPU; TPUs are not preemptible.
B.Use Vertex AI Training without spot VMs, because preemptible VMs are not supported for training.
C.Use Vertex AI Training with spot VMs and ensure the training code saves checkpoints periodically to Cloud Storage.
D.Use a single powerful non-preemptible VM to avoid interruptions.
AnswerC

Spot VMs suit this scenario because Vertex AI automatically restarts preempted training jobs, satisfying the bounded-time constraint. Periodic checkpointing to Cloud Storage preserves progress across interruptions, so restarts resume from the last saved state rather than beginning again. This combination absorbs preemption while keeping costs low.

Why this answer

Vertex AI Training supports spot VMs (preemptible instances) for cost savings, and periodic checkpointing to Cloud Storage ensures that training can resume from the last saved state if a VM is preempted, allowing the job to complete within a bounded time despite interruptions.

Exam trap

A common misconception is that preemptible VMs are not supported in Vertex AI Training, but they are fully supported as spot VMs. The key to bounded-time completion is checkpointing to Cloud Storage for resumability.

How to eliminate wrong answers

Option A is wrong because Cloud TPUs are not inherently non-preemptible; they can also be preempted, and using TPUs does not address the cost-reduction goal with preemptible VMs. Option B is wrong because Vertex AI Training does support spot VMs (preemptible VMs) for training jobs, so the claim that they are not supported is incorrect. Option D is wrong because using a single powerful non-preemptible VM increases costs significantly and does not leverage the cost savings of preemptible instances, while still being susceptible to other failures without checkpointing.

21
Multi-Selecteasy

A company wants to use Vertex AI JumpStart to deploy a pre-trained image classification model and later fine-tune it on their own data. Which TWO statements are true about Vertex AI JumpStart?

Select 2 answers
A.JumpStart requires users to build custom Docker containers for all models
B.JumpStart only supports text-based models
C.JumpStart allows you to fine-tune foundation models like Gemma
D.JumpStart only supports tabular data models
E.JumpStart provides one-click deployment of pre-trained models and ML solutions
AnswersC, E

JumpStart supports fine-tuning of foundation models such as Gemma.

Why this answer

Vertex AI JumpStart supports fine-tuning of foundation models like Gemma, allowing users to adapt pre-trained models to their specific datasets. This capability is built into JumpStart's managed environment, which handles the underlying infrastructure for training and deployment.

Exam trap

In the Google PMLE exam, candidates often mistakenly think that JumpStart only supports a narrow set of model types (e.g., text-only or tabular-only), when in fact it supports a broad range including image, text, and tabular models, and provides one-click deployment and fine-tuning capabilities.

22
Multi-Selecthard

A machine learning team is building a feature engineering pipeline using Dataflow. They need to compute features from streaming data and store them in Vertex AI Feature Store for online serving. The features must be updated within 5 seconds of the event. Which TWO services should they combine? (Select 2)

Select 2 answers
A.Cloud Dataflow for stream processing and feature computation
B.Cloud Pub/Sub for event ingestion
C.Cloud Storage for feature store
D.Cloud Functions for feature transformation
E.BigQuery for feature storage
AnswersA, B

Dataflow can compute features in near real-time and write to Feature Store.

Why this answer

Cloud Dataflow is correct because it provides unified stream and batch processing with exactly-once semantics, enabling low-latency feature computation from streaming data. It integrates natively with Vertex AI Feature Store for online serving, ensuring features are updated within the required 5-second SLA.

Exam trap

The exam often tests the distinction between general-purpose storage services (Cloud Storage, BigQuery) and the dedicated online feature store (Vertex AI Feature Store) required for real-time ML serving, leading candidates to pick a storage option instead of the correct streaming ingestion (Pub/Sub) and processing (Dataflow) pair.

23
MCQhard

You have an edge device with limited compute resources. You need to deploy a deep learning model for real-time inference. Which model compression technique should you apply to reduce the model size and latency with minimal accuracy loss?

A.Pruning only
B.Post-training quantization to INT8
C.Knowledge distillation only
D.Use full precision FP32 to maintain accuracy
AnswerB

Post-training quantization to INT8 converts trained weights and activations from FP32 to 8-bit integers, cutting model size roughly fourfold and accelerating inference on constrained edge hardware. It requires no retraining, satisfying the limited-compute constraint, while typically preserving accuracy with minimal degradation for real-time deployment.

Why this answer

Post-training quantization to INT8 reduces model weights and activations from 32-bit floats to 8-bit integers, cutting model size by roughly 4x and speeding up inference on edge hardware with minimal accuracy loss. It requires no retraining and works with most trained models. This makes it the best fit for limited-compute edge deployment.

Exam trap

The trap is assuming pruning or distillation alone delivers the same size/latency reduction as quantization, when quantization is the fastest, most reliable post-training compression for edge inference.

How to eliminate wrong answers

Option A is wrong because pruning alone removes weights but does not reduce numeric precision, so latency and size gains are smaller and often require fine-tuning to recover accuracy. Option C is wrong because knowledge distillation requires training a separate smaller student model, which is more complex and not a pure compression technique applied post-hoc. Option D is wrong because FP32 is the uncompressed baseline and offers no size or latency reduction, defeating the purpose of edge deployment.

24
MCQmedium

A data scientist needs to train a large PyTorch model on a custom dataset using Vertex AI. The training script expects data from Cloud Storage and uses GPU acceleration. Which option correctly configures a custom training job with a pre-built container for PyTorch and attaches a single NVIDIA V100 GPU?

A.Use a custom container built from PyTorch base image and specify accelerator_count=1 in the machine spec
B.Use the pre-built container 'us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-12:latest' and in worker_pool_specs set machine_type='n1-standard-4', accelerator_type='NVIDIA_TESLA_V100', accelerator_count=1
C.Use the AI Platform Training service with gcloud ai-platform jobs submit training and --scale-tier BASIC_GPU
D.Create a training pipeline with AutoML and select GPU runtime
AnswerB

The PyTorch GPU pre-built container supplies the training runtime, while worker_pool_specs declares machine_type, accelerator_type='NVIDIA_TESLA_V100' and accelerator_count=1, attaching exactly one V100. This satisfies both the custom PyTorch training and single-GPU constraints without building a custom image.

Why this answer

Vertex AI custom training jobs use worker_pool_specs to define the machine type, container image, and accelerator configuration. The correct configuration uses the pre-built PyTorch GPU container and specifies machine_type='n1-standard-4', accelerator_type='NVIDIA_TESLA_V100', and accelerator_count=1 in the worker pool spec. This is the documented Vertex AI pattern for attaching a single V100 GPU.

Exam trap

PMLE often tests the exact API field names and the distinction between Vertex AI custom training and legacy AI Platform Training — candidates pick 'custom container' or 'AI Platform Training' because they sound plausible, but only the worker_pool_specs configuration with the pre-built PyTorch GPU image and correct accelerator_type is valid on Vertex AI.

How to eliminate wrong answers

Option A is wrong because while a custom container is valid, the answer omits the required worker_pool_specs structure and the correct accelerator_type string — 'accelerator_count=1 in the machine spec' is not a valid Vertex AI API field; the field is accelerator_count inside worker_pool_specs.machine_spec. Option C is wrong because AI Platform Training (the legacy service) uses --scale-tier BASIC_GPU, which does not let you specify a V100 or a pre-built PyTorch container in the Vertex AI manner — it's a different, deprecated service. Option D is wrong because AutoML does not support custom PyTorch training scripts or GPU runtime selection; AutoML is for tabular/image/text models with managed training, not custom code.

25
MCQmedium

You are designing a distributed training job on Vertex AI for a PyTorch model using DataDistributedParallel (DDP). You have 4 nodes, each with 4 GPUs. What is the total number of workers that should be configured in the TF_CONFIG equivalent for PyTorch?

A.4
B.8
C.1
D.16
AnswerD

DataDistributedParallel assigns one worker process per GPU, so 4 nodes multiplied by 4 GPUs each yields 16 workers. Configuring 16 matches the total GPU count and ensures every device participates in the distributed training job.

Why this answer

In PyTorch DDP on Vertex AI, the number of workers equals the total number of processes across all nodes, which is nodes × GPUs per node. With 4 nodes and 4 GPUs each, the total is 16 workers, so the TF_CONFIG-equivalent worker count should be 16. Each GPU runs one DDP process, and all 16 processes participate in the all-reduce synchronization.

Exam trap

PMLE often tests the worker-count calculation — candidates multiply nodes by GPUs incorrectly or forget that each GPU runs its own DDP process, leading them to pick the node count instead of the total process count.

How to eliminate wrong answers

Option A is wrong because 4 corresponds only to the number of nodes, ignoring the 4 GPUs per node that each run a separate DDP process. Option B is wrong because 8 would correspond to 2 GPUs per node or 2 nodes × 4 GPUs, which does not match the stated topology. Option C is wrong because 1 would mean a single-process job, which contradicts the distributed 4×4 GPU configuration described.

26
MCQhard

You need to preprocess a large dataset (terabytes) for training a TensorFlow model. The preprocessing includes scaling and bucketizing features, and the same transformations must be applied during serving. Which tool should you use?

A.Dataflow with Apache Beam and tf.Transform
B.Dataproc with Spark ML
C.Vertex AI Feature Store
D.BigQuery ML
AnswerA

Dataflow with Apache Beam and tf.Transform applies identical scaling and bucketising transformations across terabyte-scale preprocessing and serving. tf.Transform exports the transform graph so training and prediction use consistent feature engineering, avoiding training-serving skew that separate pipelines would introduce.

Why this answer

tf.Transform is a library built on Apache Beam that allows you to define preprocessing functions (like scaling and bucketizing) once and then apply them consistently in both training and serving. Dataflow provides a scalable, fully managed runner for Apache Beam pipelines, making it ideal for processing terabyte-scale datasets. This combination ensures that the exact same transformations are used during training and serving, avoiding training-serving skew.

Exam trap

PMLE often tests the distinction between tools that can handle large-scale preprocessing and those that are meant for feature storage or model training, so candidates might incorrectly choose Vertex AI Feature Store or BigQuery ML thinking they handle preprocessing, but the key is the need for consistent transformations during serving with TensorFlow.

How to eliminate wrong answers

Option B is wrong because Spark ML on Dataproc does not provide a built-in mechanism to export and apply the same transformations during serving; it would require custom code to replicate the preprocessing logic, increasing the risk of skew. Option C is wrong because Vertex AI Feature Store is designed for online serving of features and does not handle large-scale preprocessing or transformation of raw data; it stores precomputed features. Option D is wrong because BigQuery ML is for building and training models directly in BigQuery using SQL, and while it can do some preprocessing, it does not provide a way to apply the same transformations during serving outside of BigQuery, and it is not designed for TensorFlow model serving.

27
MCQhard

You have a very large language model that does not fit on a single GPU. You need to train it efficiently across multiple GPUs on a single machine. Which approach should you use?

A.Data parallelism with MirroredStrategy
B.Data parallelism with MultiWorkerMirroredStrategy
C.Use TPU training as TPUs have more memory
D.Model parallelism using pipeline parallelism
AnswerD

Pipeline parallelism splits the model's layers across GPUs so each device holds only a subset of parameters, with activations passed between stages. This fits a model too large for one GPU while training efficiently within a single machine.

Why this answer

When a model does not fit on a single GPU, model parallelism is required because it partitions the model itself across devices. Pipeline parallelism is a specific model parallelism technique that splits the model into stages across GPUs and pipelines micro-batches to maintain utilization, making it the appropriate approach for training a very large model on multiple GPUs in one machine.

Exam trap

PMLE often tests the misconception that data parallelism can handle models too large for one GPU, when in fact data parallelism replicates the model and only model parallelism (including pipeline parallelism) partitions it across devices.

How to eliminate wrong answers

Option A is wrong because data parallelism with MirroredStrategy replicates the entire model on each GPU, which is impossible if the model does not fit on one GPU. Option B is wrong because MultiWorkerMirroredStrategy is also data parallelism, just across multiple machines, and still requires the full model to fit on each worker. Option C is wrong because TPUs are not a given in a single-machine multi-GPU scenario, and switching hardware does not address the architectural need for model partitioning; also, TPUs have their own memory limits.

28
MCQhard

You are deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that involves scaling numerical features and one-hot encoding categorical features. You have packaged the preprocessing and model into a single Python script that uses a custom prediction routine. You need to ensure that the online prediction service can handle varying input formats and provide low-latency responses. What should you do?

A.Use Vertex AI's custom prediction routine with a preprocessor that accepts raw input and transforms it, and deploy the model as a model artifact with the custom routine.
B.Deploy the model as a custom container on Vertex AI, and implement the preprocessing logic inside the container's HTTP server.
C.Use Vertex AI's feature store to perform preprocessing, and then feed the features directly to the model.
D.Deploy the model using Vertex AI's built-in scikit-learn container, and perform preprocessing on the client side before sending requests.
AnswerA

Vertex AI supports custom prediction routines where you can define a preprocessor that handles raw input and transforms it before passing to the model. This allows you to encapsulate preprocessing and model logic in a single Python package. It leverages Vertex AI's managed infrastructure for scaling and low-latency serving, and is the recommended approach for scikit-learn models with custom preprocessing.

Why this answer

Using a custom prediction routine with a preprocessor allows you to encapsulate preprocessing and model logic in a single package, ensuring consistency between training and serving. Vertex AI manages the infrastructure, providing low-latency and scalability. This is the most efficient and maintainable approach for scikit-learn models with custom preprocessing.

Exam trap

The trap here is assuming that a custom container is necessary for custom preprocessing, when Vertex AI's custom prediction routines already provide this capability without the overhead of managing a container.

29
MCQhard

A team is deploying a scikit-learn model to Vertex AI for online prediction. The model requires a custom preprocessing step that scales numerical features using statistics computed from the training set. The preprocessing must be identical between training and serving. The team wants to minimize latency and ensure consistency. What should they do?

A.Wrap the preprocessing and model into a single scikit-learn Pipeline object, and deploy that pipeline as a custom model on Vertex AI.
B.Use a Vertex AI pipeline to preprocess the data before training, and then deploy the model without preprocessing, assuming the input data at serving time will already be scaled.
C.Save the scikit-learn model and the scaler as separate artifacts, and in the prediction script, load both and apply the scaler before calling model.predict.
D.Use Vertex AI Feature Store to store the scaled features and serve them online, bypassing the need for preprocessing in the model.
AnswerA

A scikit-learn Pipeline encapsulates all preprocessing and the final estimator, ensuring that the exact same transformations are applied during training and serving. Deploying the pipeline as a custom model on Vertex AI guarantees consistency and reduces the risk of training-serving skew, while keeping the serving logic simple.

Why this answer

Wrapping preprocessing and the model into a single scikit-learn Pipeline ensures that the exact same transformations are applied during both training and serving. This eliminates training-serving skew and simplifies deployment. Vertex AI supports deploying custom models with custom prediction routines, but using a native scikit-learn Pipeline is more straightforward and less error-prone, as it encapsulates all steps and can be serialized as one artifact.

Exam trap

The trap here is assuming that preprocessing can be handled separately or that the client will send preprocessed data, overlooking the need for a single artifact that guarantees consistency.

30
MCQeasy

You have a TensorFlow training script that runs on a single machine. To speed up training on Vertex AI with 8 GPUs on a single machine, which strategy should you use?

A.tf.distribute.ParameterServerStrategy
B.tf.distribute.MirroredStrategy
C.tf.distribute.TPUStrategy
D.tf.distribute.MultiWorkerMirroredStrategy
AnswerB

MirroredStrategy performs synchronous, all-reduce data-parallel training across multiple GPUs within one machine, replicating the model on each device and aggregating gradients. This directly satisfies the stem's constraint of 8 GPUs on a single machine, where MultiWorkerMirroredStrategy would add unnecessary cross-machine networking overhead.

Why this answer

tf.distribute.MirroredStrategy is designed for synchronous, data-parallel training across multiple GPUs on a single machine. It replicates the model on each GPU, splits each batch across replicas, and uses all-reduce (via NCCL) to aggregate gradients, which is exactly the scenario described: 8 GPUs on one machine. This is the canonical strategy for single-node multi-GPU TensorFlow training.

Exam trap

The trap is confusing single-machine multi-GPU (MirroredStrategy) with multi-machine multi-GPU (MultiWorkerMirroredStrategy) — candidates who skim the question miss the 'single machine' qualifier and pick the multi-worker variant.

How to eliminate wrong answers

Option A is wrong because ParameterServerStrategy is designed for asynchronous or synchronous training across many machines with dedicated parameter servers, which is overkill and inappropriate for a single machine with 8 GPUs. Option C is wrong because TPUStrategy targets Google Cloud TPUs, not GPUs — using it on a GPU machine will fail or fall back incorrectly. Option D is wrong because MultiWorkerMirroredStrategy is for multi-node, multi-GPU training across multiple machines; on a single machine it adds unnecessary coordination overhead and is not the intended strategy.

31
MCQmedium

You are training a scikit-learn random forest on a dataset that fits in memory using Vertex AI custom training. The prototype notebook took 15 minutes, but the Vertex AI job takes over an hour and occasionally fails with a resource error. You want the production job to complete reliably without changing the model or preprocessing. What should you do?

A.Switch the training job to use a preemptible VM to reduce contention on shared resources.
B.Select a machine type with more memory and vCPUs, and set the appropriate boot disk size for the training job.
C.Configure the training job to run on a single n1-standard-4 machine with no accelerator.
D.Increase the number of worker replicas and use a distributed reduction strategy for the random forest.
AnswerB

Vertex AI custom training lets you choose a machine type and boot disk for the worker pool. A scikit-learn job that fits in memory but fails on a small default machine benefits from more RAM and vCPUs. Increasing available memory prevents out-of-memory failures, and a larger boot disk ensures the container and any temporary files have sufficient space, matching the reliability goal without changing the model.

Why this answer

The job fits in memory but fails on the default Vertex AI worker configuration, which points to insufficient RAM or vCPUs on the training machine. Choosing a larger machine type and an adequate boot disk gives the scikit-learn process the resources it needs. Distributed training is unnecessary and would require changing the algorithm, while preemptible VMs reduce reliability rather than improve it.

Exam trap

The trap here is assuming that any training job that fails must be fixed with distributed training, when a single-machine resource increase is the correct and simpler remedy.

32
Multi-Selectmedium

You are deploying a prototype model to Vertex AI for online prediction. The model was trained with a custom preprocessing step that must run on raw JSON input before inference. You need the endpoint to return predictions with minimal latency and to support rolling updates of new model versions without downtime. (Choose two.)

Select 2 answers
A.Deploy the model to an endpoint with a traffic split between the current version and a new version, gradually shifting traffic to the new version.
B.Deploy the preprocessing step as a separate Vertex AI endpoint and call it from the client before sending data to the model endpoint.
C.Use a Vertex AI batch prediction job for the initial rollout and switch to online prediction after the model is validated.
D.Store the preprocessing logic in a Cloud Function and invoke it from the model's prediction container at request time.
E.Package the preprocessing and the model into a single custom container that implements the Vertex AI HTTP prediction contract.
AnswersA, E

Vertex AI endpoints support deploying multiple models and assigning a traffic split percentage to each deployed model. This enables canary or rolling updates where a small share of traffic goes to the new version first, and the split is adjusted as confidence grows. It provides zero-downtime updates because both versions serve simultaneously during the transition.

Why this answer

Bundling preprocessing and model inference into one custom container that follows the Vertex AI HTTP prediction contract keeps the request path short and self-contained, which minimizes latency. Deploying model versions to the same endpoint with a traffic split enables canary or rolling updates, so new versions can be validated with a small share of live traffic before full promotion, achieving zero-downtime releases.

Exam trap

The trap here is separating preprocessing into its own service or using batch prediction for rollout, when the latency and zero-downtime goals are best met by a single container plus endpoint traffic splitting.

33
MCQmedium

You are deploying a custom model to Vertex AI for online prediction. The model requires a preprocessing step that normalizes input features. You want to ensure that the same preprocessing is applied during both training and serving to avoid training-serving skew. What should you do?

A.Use Vertex AI Feature Store to serve features and apply preprocessing at training time only.
B.Perform preprocessing on the client side before sending data to the model for prediction.
C.Export the preprocessing as a TensorFlow Transform (tf.Transform) function and include it in both training and serving graphs.
D.Include the preprocessing logic in the training script and replicate it in the serving container.
AnswerC

tf.Transform allows you to define preprocessing as a function that can be applied consistently during training and serving. By exporting the transform function and including it in both the training and serving graphs, you ensure identical preprocessing. This is a best practice for avoiding training-serving skew, as the same code is used in both phases.

Why this answer

Using TensorFlow Transform (tf.Transform) to define preprocessing ensures that the same transformations are applied during both training and serving. The transform function is included in the model graph, so the serving container automatically applies the same logic. This eliminates manual replication and reduces the risk of training-serving skew.

Exam trap

The trap here is thinking that duplicating preprocessing code in both training and serving is sufficient, but it often leads to skew due to maintenance issues; using a shared transform function is more reliable.

34
MCQhard

You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?

A.Model parallelism using pipeline parallelism
B.Data parallelism with MultiWorkerMirroredStrategy
C.Mixed precision training (FP16)
D.Data parallelism with tf.distribute.MirroredStrategy
AnswerA

Pipeline parallelism splits the model's layers across GPUs, with each device holding a subset and passing activations onward. This addresses the constraint that seven billion parameters exceed single-GPU memory, unlike data parallelism which replicates the full model per device.

Why this answer

When a model is too large to fit on a single GPU, model parallelism is required to split the model's layers across multiple devices. Pipeline parallelism is a specific form of model parallelism that partitions layers into stages and pipelines micro-batches across devices, enabling training of models like a 7B-parameter LLM across multiple GPUs. This is the correct approach when memory, not throughput, is the binding constraint.

Exam trap

PMLE often tests the misconception that data parallelism solves memory constraints, when in fact it replicates the model and only model/pipeline parallelism addresses models too large for one GPU.

How to eliminate wrong answers

Option B is wrong because data parallelism with MultiWorkerMirroredStrategy replicates the full model on each worker — if the model does not fit on one GPU, it will not fit on any, so this approach fails. Option C is wrong because mixed precision (FP16) reduces memory footprint but does not solve the fundamental problem of a model that exceeds single-GPU memory; it is an optimization, not a distribution strategy. Option D is wrong because MirroredStrategy is single-machine data parallelism that also replicates the full model on each GPU, which is impossible when the model exceeds single-GPU memory.

35
MCQhard

You are training a TensorFlow model on Vertex AI using a custom container with a single Tesla T4 GPU. You notice that training is slower than expected, and GPU utilization is consistently below 20%. Profiling shows that the input pipeline is the bottleneck. Which change should you make to improve GPU utilization?

A.Switch to a larger GPU instance with more memory to reduce data loading overhead.
B.Move the dataset to a local SSD on the training VM to reduce I/O latency.
C.Increase the batch size to fully utilize GPU memory.
D.Use tf.data with prefetching and parallel data extraction to overlap data loading with GPU computation.
AnswerD

The tf.data API with prefetching and parallel extraction allows data preprocessing to occur on CPU while the GPU computes on the previous batch, effectively overlapping I/O and compute. This directly addresses the input pipeline bottleneck, increasing GPU utilization and reducing training time. It is the recommended approach for optimizing input pipelines in TensorFlow.

Why this answer

When the input pipeline is the bottleneck, the GPU waits for data. Using tf.data with prefetching and parallel extraction overlaps CPU preprocessing with GPU computation, ensuring the GPU is continuously fed. This is the standard TensorFlow optimization for such scenarios and directly improves GPU utilization.

Exam trap

The trap here is assuming that a faster GPU or more memory will fix low GPU utilization, when the issue is actually data starvation from an inefficient input pipeline.

36
MCQhard

A machine learning team is deploying a PyTorch model on Vertex AI Prediction for real-time inference. The model was trained with preprocessing that includes tokenization and normalization. They want to embed the preprocessing logic in the model to reduce prediction latency and avoid additional service calls. Which approach should they take?

A.Deploy the preprocessing logic as a Cloud Function and invoke it before calling the prediction endpoint
B.Wrap the preprocessing logic in a Flask application and deploy it as a separate microservice in front of the prediction endpoint
C.Use TorchScript to trace the preprocessing steps and export the entire pipeline as a single scripted model
D.Use TensorFlow Transform to convert preprocessing into a SavedModel and call it from the PyTorch model
AnswerC

TorchScript tracing captures the tokenisation and normalisation operations as graph nodes, fusing them with the PyTorch model into one serialised artefact. Vertex AI Prediction then serves this single scripted model, eliminating the separate preprocessing service call and satisfying the stem's latency-reduction constraint.

Why this answer

TorchScript allows you to trace or script a PyTorch model, including preprocessing operations like tokenization and normalization, into a single serialized artifact. By embedding preprocessing in the TorchScript model, the entire pipeline runs in one forward pass on the Vertex AI endpoint, eliminating extra service calls and reducing latency. This is the standard approach for consolidating preprocessing with a PyTorch model for real-time inference.

Exam trap

PMLE often tests the misconception that preprocessing must be a separate service — candidates pick Cloud Functions or Flask microservices because they are familiar patterns, but the question explicitly asks to embed preprocessing in the model, and TorchScript is the PyTorch-native way to do that.

How to eliminate wrong answers

Option A is wrong because a Cloud Function adds a network hop and cold-start latency, increasing prediction latency rather than reducing it — the goal is to avoid additional service calls, not add one. Option B is wrong because a separate Flask microservice in front of the endpoint introduces another network call and operational overhead, directly contradicting the requirement to reduce latency and avoid additional service calls. Option D is wrong because TensorFlow Transform produces a SavedModel for TensorFlow, not PyTorch — it cannot be directly called from a PyTorch model without a cross-framework bridge, which adds complexity and latency.

37
MCQhard

An ML engineer is using Vertex AI distributed training for a TensorFlow model that uses the MirroredStrategy. They notice that the training throughput drops significantly when moving from a single GPU to multiple GPUs on the same machine. What is the most likely cause?

A.The GPUs are not properly configured in TF_CONFIG.
B.The batch size is too small, causing each GPU to complete its forward pass quickly, but the sync wait dominates.
C.The learning rate is too high, causing instability.
D.The model uses TensorFlow 1.x instead of 2.x.
AnswerB

MirroredStrategy performs an all-reduce gradient sync across GPUs after every step. With a small batch size, each GPU's forward and backward pass finishes quickly, so communication overhead dominates and throughput falls rather than scaling with added GPUs.

Why this answer

In MirroredStrategy, all GPUs must finish their forward/backward pass before the all-reduce gradient synchronization can occur, so throughput is bounded by the slowest replica plus the sync overhead. With a small batch size, each GPU's compute time is very short, so the fixed cost of the NCCL all-reduce (and the sync wait) dominates the step time, making multi-GPU training slower than single-GPU. Increasing the per-replica batch size amortizes the synchronization cost over more compute, restoring scaling efficiency.

Exam trap

The trap here is assuming any multi-GPU slowdown must be a configuration error (TF_CONFIG, framework version) rather than recognizing that synchronization overhead in data-parallel training is amortized by batch size — a classic distributed-training performance pitfall.

How to eliminate wrong answers

Option A is wrong because TF_CONFIG is used for multi-worker distributed training (MultiWorkerMirroredStrategy, ParameterServerStrategy) to specify cluster/worker/task info, not for single-machine MirroredStrategy, which auto-discovers local GPUs via NCCL. Option C is wrong because a high learning rate causes divergence or NaN loss, not a throughput drop — it affects convergence quality, not step time. Option D is wrong because TensorFlow 1.x vs 2.x affects API style and graph/eager execution, not the fundamental synchronization overhead that causes the observed throughput regression.

38
Multi-Selectmedium

An ML team is optimizing an inference model for deployment on edge devices. They need to reduce the model size and improve latency while maintaining accuracy as much as possible. Which two techniques should they use? (Choose TWO.)

Select 2 answers
A.Use a larger pre-trained model as a starting point.
B.Post-training quantization to INT8.
C.Use half-precision (FP16) instead of INT8.
D.Apply weight pruning to remove small weights.
E.Increase the number of layers in the model.
AnswersB, D

Reduces size and latency with minimal accuracy loss.

Why this answer

Post-training quantization to INT8 reduces model size by converting 32-bit floating-point weights and activations to 8-bit integers, which also speeds up inference on edge devices with integer-optimized hardware. This technique typically maintains accuracy within 1-2% of the original model while significantly lowering memory footprint and latency.

Exam trap

Candidates often think that FP16 is always better than INT8 for edge devices, but INT8 offers greater size reduction and is more widely supported on edge hardware, including Google's Edge TPU.

39
Multi-Selectmedium

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced and that artifacts are tracked. Which two of the following practices should they follow? (Choose two.)

Select 2 answers
A.Version all pipeline components and store them in a version control system.
B.Use the same container image tags for all pipeline steps to ensure consistency.
C.Manually copy all artifacts to a Cloud Storage bucket after each run.
D.Use Vertex AI Metadata to record parameters, metrics, and artifacts for each pipeline run.
E.Disable caching for all pipeline steps to ensure fresh executions.
AnswersA, D

Versioning pipeline components ensures that changes are tracked and that specific versions can be reproduced. Storing them in version control allows the team to revert to previous versions and understand the history of changes. This is essential for reproducibility and artifact tracking.

Why this answer

Versioning pipeline components and using Vertex AI Metadata are key practices for reproducibility and artifact tracking. Versioning ensures that the exact code and dependencies are captured, while Metadata automatically records parameters, metrics, and artifacts for each run, providing lineage and enabling reproducibility.

Exam trap

The trap here is thinking that manual artifact management or disabling caching is necessary for reproducibility, but automated metadata tracking and component versioning are the core practices.

40
MCQeasy

A machine learning engineer wants to use Vertex AI Vizier to tune three hyperparameters: learning rate (log scale), number of layers (integer), and optimizer (categorical). They have 50 parallel trials available. Which parameter specification types should they define?

A.learning_rate: CATEGORICAL, layers: INTEGER, optimizer: CATEGORICAL
B.learning_rate: DOUBLE (unit_log_scale), layers: INTEGER (unit_linear_scale), optimizer: CATEGORICAL
C.learning_rate: DOUBLE (unit_log_scale), layers: DOUBLE (unit_linear_scale), optimizer: DISCRETE
D.learning_rate: DOUBLE (unit_linear_scale), layers: INTEGER (unit_linear_scale), optimizer: CATEGORICAL
AnswerB

Correct types and scales for the parameters.

Why this answer

Vertex AI Vizier requires parameter specifications that match the nature of each hyperparameter. Learning rate is best explored on a logarithmic scale, so it should be a DOUBLE parameter with unit_log_scale. Number of layers is a discrete integer count, so it should be an INTEGER parameter with unit_linear_scale.

Optimizer is a categorical choice among named algorithms, so it should be CATEGORICAL. This combination correctly reflects the mathematical and structural properties of each hyperparameter.

Exam trap

PMLE often tests the confusion between DISCRETE and CATEGORICAL parameter types, and between linear and log scales, causing candidates to pick specifications that do not match the hyperparameter's mathematical nature.

How to eliminate wrong answers

Option A is wrong because treating learning rate as CATEGORICAL discards the continuous, ordered nature of the value and prevents Vizier from exploring intermediate values on a log scale. Option C is wrong because number of layers should be an INTEGER, not a DOUBLE, since layers are discrete counts, and optimizer should be CATEGORICAL, not DISCRETE, because DISCRETE is for ordered numeric values rather than named categories. Option D is wrong because learning rate on a unit_linear_scale is inefficient; log scale is standard for learning rates that span orders of magnitude.

41
MCQeasy

You have a prototype ML model that you want to scale to production on Vertex AI. The model is a Python function that performs simple data preprocessing and then calls a pre-trained scikit-learn model. You need to deploy this as a batch prediction job that runs weekly on a large dataset stored in BigQuery. What is the most efficient way to accomplish this?

A.Deploy the model to a Vertex AI Endpoint and write a script that sends all BigQuery rows as individual prediction requests, then store the results in BigQuery.
B.Create a custom container that includes the preprocessing code and the scikit-learn model, push it to Artifact Registry, and use it to run a Vertex AI Batch Prediction job with BigQuery as the input source.
C.Use a Vertex AI Pipelines job that runs a Dataflow transform to preprocess the BigQuery data, then calls a pre-built scikit-learn model for prediction and writes results back to BigQuery.
D.Export the BigQuery data to Cloud Storage as CSV files, then run a Vertex AI Batch Prediction job using a pre-built scikit-learn container that reads from Cloud Storage.
AnswerB

Vertex AI Batch Prediction supports custom containers, allowing you to package preprocessing logic and the model together. Using BigQuery as the input source is natively supported, and the job can scale to handle large datasets. This approach is efficient because it leverages Vertex AI's managed batch prediction infrastructure without needing to export data manually.

Why this answer

For batch prediction with custom preprocessing, a custom container is required because pre-built containers do not include your code. Vertex AI Batch Prediction natively supports BigQuery as both input and output, so you can avoid data movement. This method is efficient, scalable, and integrates with Vertex AI's managed infrastructure.

Exam trap

The trap here is assuming that a pre-built container can handle custom preprocessing, or that data must be exported from BigQuery, when Vertex AI Batch Prediction supports custom containers and direct BigQuery input.

42
MCQeasy

An ML engineer has a prototype that trains a TensorFlow model on a single CPU machine using Vertex AI custom training. The job now needs to train on a larger dataset and must use multiple GPUs on one machine. The training script already uses tf.distribute.MirroredStrategy. What change is required to scale the job?

A.Add a second worker pool with one replica and set the distribution strategy to MultiWorkerMirroredStrategy.
B.Enable TPU training by specifying a TPU machine type and changing the script to use TPUStrategy.
C.Set the worker pool machine type to a GPU machine and specify the number of accelerators and accelerator type.
D.Increase the boot disk size and set the training container to use the GPU-enabled base image.
AnswerC

Vertex AI custom training uses the worker pool specification to select the machine type, accelerator type, and accelerator count. Because the script already uses MirroredStrategy, it will automatically detect the local GPUs and replicate training across them. Providing a GPU machine type with the desired accelerator count is the only configuration change needed to scale this single-machine job.

Why this answer

Scaling a single-machine TensorFlow job that already uses MirroredStrategy to multiple GPUs is primarily a resource configuration task. You set the machine type to a GPU-capable machine and specify the accelerator type and count in the worker pool spec. The strategy code detects the local GPUs automatically.

Multi-worker strategies, TPU strategies, and disk or image changes are not required for this scenario.

Exam trap

The trap here is assuming that multi-GPU training always requires a multi-worker strategy, when a single machine with multiple accelerators works with MirroredStrategy.

43
MCQeasy

A data scientist wants to use a pre-trained ResNet model from Keras Applications and fine-tune it on a small custom dataset. Which approach should they take to avoid overfitting?

A.Freeze the first few layers and train the rest.
B.Add more convolutional layers to the model.
C.Use a larger learning rate to speed up training.
D.Train the entire model from scratch on the custom dataset.
AnswerA

Freezing early convolutional layers preserves the generic features ResNet learned on ImageNet, so only the later, task-specific layers train. Restricting updates to fewer parameters on a small dataset reduces overfitting while still adapting the model.

Why this answer

Freezing the earlier layers (which capture general features) and only training the later layers is a common transfer learning approach for small datasets, reducing overfitting.

44
MCQmedium

An engineer is using TensorFlow Transform (tf.Transform) to preprocess training data. They want to ensure that the same preprocessing logic is applied during inference without code duplication. Which approach should they take?

A.Use tf.Transform at prediction time by running a separate Beam pipeline
B.Use Dataflow to preprocess data for both training and serving
C.Use tf.Transform to generate a transform_fn and save it as a SavedModel; then use tf.saved_model.load to apply it in the serving pipeline
D.Write separate preprocessing code for training and serving in Python
AnswerC

Saving the transform_fn as a SavedModel exports the full preprocessing graph, including constants such as vocabulary and mean values computed during training. Loading it with tf.saved_model.load in the serving pipeline applies identical logic at inference, satisfying the no-duplication constraint without reimplementing feature engineering.

Why this answer

TensorFlow Transform outputs a SavedModel that contains the preprocessing graph. This can be exported as a transform_fn and embedded in the serving model, ensuring consistency between training and serving.

45
MCQmedium

You are training a scikit-learn random forest model on a large tabular dataset using a Vertex AI custom training job. The training script reads a CSV file from a Cloud Storage bucket and writes the trained model artifact to the same bucket. You need to ensure the training job can access the Cloud Storage bucket without embedding long-lived credentials in the container. What should you do?

A.Grant the Vertex AI Service Agent role to the default Compute Engine service account and rely on the default credentials provided by the training environment to access the bucket.
B.Configure the training job to use a user-managed service account that has the Editor role on the project, ensuring it has sufficient permissions to read and write to any Cloud Storage bucket.
C.Attach a custom service account to the Vertex AI training job that has the Storage Object Viewer and Storage Object Creator roles on the bucket, and let the Vertex AI Training service use the service account's credentials automatically.
D.Create a service account key file, store it in Secret Manager, and mount it into the training container as a volume so the scikit-learn script can read the key and authenticate to Cloud Storage.
AnswerC

Vertex AI training jobs run as a service account you specify. By attaching a custom service account with least-privilege Cloud Storage roles, the training container automatically obtains credentials via the metadata server. This avoids embedding keys and follows Google-recommended security practices for accessing Cloud Storage from training jobs.

Why this answer

The correct approach is to attach a custom service account with least-privilege Cloud Storage roles to the Vertex AI training job. Vertex AI training jobs automatically use the credentials of the attached service account via the metadata server, so no key files are needed. This ensures secure, short-lived authentication and follows Google Cloud best practices for access control.

Exam trap

The trap here is assuming that a service account key file or default credentials are required, when Vertex AI training jobs natively support attaching a custom service account for automatic authentication.

46
MCQhard

A machine learning engineer is deploying a TensorFlow model on an edge device with limited memory and compute. The model needs to perform inference with low latency. The engineer has a trained float32 model. Which model compression technique should be applied first to reduce the model size and improve inference speed without significant accuracy loss?

A.Post-training quantization to INT8
B.Knowledge distillation
C.Quantization-aware training
D.Weight pruning
AnswerA

Post-training quantization converts the trained float32 weights and activations to INT8, cutting model size roughly fourfold and accelerating inference on edge hardware with limited memory and compute. It requires no retraining and typically preserves accuracy, making it the appropriate first compression step.

Why this answer

Post-training quantization to INT8 is the correct first step because it directly reduces the model size by approximately 4x (from 32-bit floats to 8-bit integers) and speeds up inference on edge devices by leveraging integer-optimized hardware (e.g., ARM NEON or Qualcomm Hexagon). This technique requires no retraining and typically yields minimal accuracy loss for most TensorFlow models, making it the fastest path to deploy on resource-constrained devices.

Exam trap

Google Cloud often tests the misconception that quantization-aware training is always required for INT8 deployment, but the trap here is that post-training quantization is the simplest and most effective first step for reducing model size and latency on edge devices, with quantization-aware training reserved only for cases where accuracy drops below acceptable thresholds.

How to eliminate wrong answers

Option B (Knowledge distillation) is wrong because it requires training a smaller student model from scratch using the teacher model's outputs, which is computationally expensive and not a quick compression technique for an already trained model. Option C (Quantization-aware training) is wrong because it simulates quantization effects during training to preserve accuracy, but it requires retraining the model, making it a second step after post-training quantization if accuracy loss is unacceptable. Option D (Weight pruning) is wrong because it removes individual weights (often via magnitude-based pruning), which can reduce model size but typically requires retraining to recover accuracy and does not directly improve inference speed on standard edge hardware without sparse matrix support.

47
MCQmedium

A company has a TensorFlow model for image classification that must run on edge devices with limited memory. They need to reduce the model size without significant accuracy loss. Which technique should they use?

A.Post-training quantization using TensorFlow Lite.
B.Knowledge distillation to train a smaller student model.
C.Pruning the model weights to zero out unimportant connections.
D.Use a larger VM for training.
AnswerA

Post-training quantization converts the trained model's float32 weights to 8-bit integers via TensorFlow Lite, shrinking size roughly fourfold and lowering memory use on constrained edge hardware, with minimal accuracy loss since no retraining is required.

Why this answer

Post-training quantization with TensorFlow Lite converts a trained model's weights from 32-bit floats to 8-bit integers, reducing model size by ~4x and speeding up inference on edge devices with minimal accuracy loss. TensorFlow Lite is purpose-built for edge deployment, making this the most direct and practical technique for the stated constraint.

Exam trap

PMLE often tests the distinction between model compression techniques — candidates confuse quantization (reduces precision) with pruning (removes weights) and distillation (trains a smaller model), picking the wrong one for the 'reduce size without retraining' constraint.

How to eliminate wrong answers

Option B is wrong because knowledge distillation requires training a smaller student model from scratch, which is more complex and time-consuming than post-training quantization, and it does not leverage the existing trained model directly. Option C is wrong because pruning alone reduces the number of weights but does not necessarily shrink the model file size dramatically unless combined with compression, and it often requires retraining to recover accuracy. Option D is wrong because using a larger VM for training does not reduce the deployed model size at all — it only affects training speed.

48
MCQeasy

You want to use Vertex AI JumpStart to quickly deploy a pre-built foundation model for text summarization. Which action is required?

A.Select the model from Model Garden and deploy it to a Vertex AI endpoint
B.Train the model from scratch using Vertex AI Training
C.Export the model to a Cloud Storage bucket and use batch prediction
D.Build a custom Docker container with the model and deploy to Vertex AI
AnswerA

JumpStart models are accessed through Model Garden, and deploying the selected foundation model to a Vertex AI endpoint is what provisions it for inference. This satisfies the stem's requirement to quickly deploy a pre-built summarisation model.

Why this answer

Vertex AI JumpStart provides pre-built foundation models in Model Garden that can be deployed directly to a Vertex AI endpoint with minimal configuration. Selecting the model from Model Garden and deploying it to an endpoint is the standard JumpStart workflow for getting a foundation model into production for tasks like text summarization. No training, containerization, or batch export is required because JumpStart handles the deployment plumbing.

Exam trap

PMLE often tests the distinction between JumpStart's one-click Model Garden deployment and the manual custom-container or training-from-scratch paths, so candidates overthink and pick the container option.

How to eliminate wrong answers

Option B is wrong because training from scratch is unnecessary and prohibitively expensive when JumpStart already offers pre-trained foundation models. Option C is wrong because exporting to Cloud Storage for batch prediction does not satisfy the goal of deploying a model for interactive inference and skips the JumpStart deployment path entirely. Option D is wrong because building a custom Docker container is the manual custom-container deployment route, not the JumpStart one-click deployment workflow.

49
Multi-Selectmedium

An ML engineer is preparing to train a large model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket as a set of TFRecord files. The engineer wants to optimize the training job to reduce cost and improve performance. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Use a TPU or GPU accelerator and ensure the input pipeline is optimized to keep the accelerator busy.
B.Enable Vertex AI Model Monitoring on the training job to detect anomalies during training.
C.Store the training data in a Cloud Storage bucket in the same region as the training job.
D.Use the default Compute Engine service account with broad permissions to simplify access to Cloud Storage.
E.Increase the number of training steps to improve model accuracy, regardless of cost.
AnswersA, C

Accelerators like GPUs or TPUs can significantly speed up training for large models. However, they must be fed data efficiently; otherwise, utilization drops. Optimizing the input pipeline with tf.data, prefetching, and parallel reads ensures the accelerator is not idle. This combination reduces training time and cost, as you pay for the accelerator only while it is used effectively.

Why this answer

To optimize training cost and performance, the engineer should use accelerators and ensure the input pipeline is efficient, and store data in the same region as the training job. These actions reduce training time and avoid unnecessary network overhead. Other options like using broad service accounts or increasing steps do not contribute to efficiency and may introduce security or cost issues.

Exam trap

The trap here is thinking that using a default service account or enabling Model Monitoring on training jobs are valid optimizations, when they are not.

50
MCQhard

You are running a distributed training job on Vertex AI using PyTorch and the DistributedDataParallel (DDP) strategy across 4 nodes, each with 8 GPUs. You notice that the training loss is not decreasing as expected and the job occasionally hangs. You suspect a communication issue between nodes. Which of the following should you check first?

A.Switch to using the Horovod framework for distributed training.
B.Reduce the number of nodes to 2 to decrease communication overhead.
C.Increase the batch size per GPU to improve gradient synchronization.
D.Ensure that the MASTER_ADDR and MASTER_PORT environment variables are correctly set and that all nodes can communicate over the network.
AnswerD

In distributed training with PyTorch DDP, the MASTER_ADDR and MASTER_PORT are used for initializing the process group and coordinating communication. If these are misconfigured or if there are network issues preventing nodes from reaching each other, the job may hang or fail to synchronize gradients. Checking these first is essential for diagnosing communication problems in multi-node setups.

Why this answer

Communication issues in multi-node PyTorch DDP training often stem from incorrect MASTER_ADDR or MASTER_PORT settings, or network connectivity problems. These variables are critical for establishing the rendezvous point for all processes. Checking them first is a logical diagnostic step before considering other changes, as they are common causes of hangs and synchronization failures.

Exam trap

The trap here is assuming that changing batch size or framework will fix communication hangs, when the first step should be verifying the distributed training environment variables and network connectivity.

51
MCQhard

You need to perform a large-scale feature computation on streaming data from Pub/Sub, transforming raw events into features, and writing results to Vertex AI Feature Store for online serving. Which Google Cloud architecture is most appropriate?

A.Use Dataproc with Spark Streaming to read from Pub/Sub and write to Feature Store
B.Use Cloud Functions triggered by Pub/Sub to compute features and update Feature Store
C.Use Dataflow streaming pipeline with Apache Beam to read from Pub/Sub, compute features, and write to Feature Store
D.Use Cloud Run to consume Pub/Sub messages and update Feature Store via a service
AnswerC

Dataflow runs Apache Beam, providing the horizontal scaling and windowing needed for high-volume Pub/Sub streams, and its native Feature Store sink writes computed features directly for low-latency online serving, meeting the streaming transformation and serving constraints.

Why this answer

Dataflow with Apache Beam is the purpose-built managed service for streaming ETL on Google Cloud, natively integrating with Pub/Sub as a source and Vertex AI Feature Store as a sink. Beam's windowing, watermarks, and exactly-once semantics handle the large-scale, continuous feature computation required for online serving. It provides autoscaling and managed infrastructure, which is exactly what a production streaming feature pipeline needs.

Exam trap

PMLE often tests the misconception that any compute service (Cloud Functions, Cloud Run) can handle streaming feature engineering, when in fact managed Apache Beam on Dataflow is the canonical answer for large-scale streaming transformations into Feature Store.

How to eliminate wrong answers

Option A is wrong because Dataproc with Spark Streaming requires manual cluster management and lacks first-class integration with Vertex AI Feature Store, adding operational overhead for a streaming pipeline. Option B is wrong because Cloud Functions has execution timeouts (up to 60 minutes for 2nd gen but typically short), limited concurrency, and is not designed for sustained high-throughput streaming feature computation. Option D is wrong because Cloud Run is a request/container-based service without native streaming windowing semantics, making it unsuitable for continuous large-scale feature computation from Pub/Sub.

52
MCQeasy

An ML engineer has a prototype scikit-learn model that must be served on Vertex AI. The model requires a custom preprocessing step that cannot be expressed in a scikit-learn Pipeline. The engineer wants to package the model with this preprocessing logic and deploy it to a Vertex AI Endpoint for online predictions. Which approach should they take?

A.Use Vertex AI Model Monitoring to apply the preprocessing transformations before the model receives the request.
B.Save the model with joblib, upload it as a Vertex AI Model, and deploy it to an Endpoint; Vertex AI automatically applies the preprocessing.
C.Deploy the model as a batch prediction job, which automatically applies preprocessing from a saved scikit-learn Pipeline.
D.Create a custom container that includes the scikit-learn model and the preprocessing code, push it to Artifact Registry, import it as a Vertex AI Model, and deploy to an Endpoint.
AnswerD

A custom container allows the engineer to bundle the model artifact, the preprocessing logic, and the required dependencies into a single image. Vertex AI runs this container for online predictions, so the preprocessing executes exactly as written before the model inference. This is the supported and recommended way to serve models with custom preprocessing on Vertex AI.

Why this answer

When a model requires custom preprocessing that is not part of the model artifact itself, the most reliable way to serve it on Vertex AI is to build a custom container. The container can include the model, the preprocessing code, and all dependencies, ensuring consistent behavior between training and serving. Vertex AI then deploys this container as a Model resource and makes it available via an Endpoint.

Exam trap

The trap here is assuming that Vertex AI automatically applies preprocessing defined in a scikit-learn Pipeline or that Model Monitoring can transform requests.

53
MCQmedium

A team is training a large TensorFlow model that requires more memory than a single GPU provides. They have access to multiple GPUs on a single machine. Which distributed training strategy should they use to split the model layers across GPUs?

A.tf.distribute.experimental.MultiWorkerMirroredStrategy
B.tf.distribute.experimental.ParameterServerStrategy
C.Manual device placement using tf.device to assign layers to specific GPUs
D.tf.distribute.MirroredStrategy
AnswerC

tf.device lets you pin individual layers to named GPUs, so a model too large for one device is partitioned layer-by-layer across the machine's GPUs. This directly satisfies the stem's requirement to split model layers rather than replicate the model.

Why this answer

When a single model's layers exceed one GPU's memory, the model itself must be partitioned across devices — this is model parallelism. Manual device placement with tf.device('/GPU:0'), tf.device('/GPU:1'), etc. is the TensorFlow-native way to assign specific layers or operations to specific GPUs, splitting the model across them.

Exam trap

The trap is assuming any 'distributed strategy' solves memory limits — most strategies (Mirrored, MultiWorker, ParameterServer) replicate the model and only help with speed, not with fitting an oversized model.

How to eliminate wrong answers

Option A is wrong because MultiWorkerMirroredStrategy replicates the full model on each worker and synchronizes gradients — it addresses throughput across machines, not the case where the model does not fit on one device. Option B is wrong because ParameterServerStrategy also replicates the model across workers with parameter servers coordinating updates; it does not split layers within a single model. Option D is wrong because MirroredStrategy performs data parallelism — it creates a full replica of the model on every GPU, which is impossible when the model exceeds single-GPU memory.

54
MCQmedium

A company is using Vertex AI Vizier for hyperparameter tuning of a model with 5 integer hyperparameters, each with a range of 10-100. They have a budget of 50 trials and want to maximize the chance of finding the best configuration. Which Vizier algorithm should they use?

A.Grid search
B.Simulated annealing
C.Bayesian optimization (GP bandit)
D.Random search
AnswerC

Bayesian optimisation with a Gaussian process bandit models the objective surface and selects trials that balance exploration against exploitation, converging efficiently within a limited budget. With 50 trials across five integer parameters, it maximises the chance of locating the best configuration.

Why this answer

Bayesian optimization (GP bandit) is Vizier's default and most sample-efficient algorithm, using a Gaussian Process surrogate model to balance exploration and exploitation across the 5-dimensional hyperparameter space. With only 50 trials over a large search space, it converges on promising regions far faster than uninformed methods. This makes it the best choice for maximizing the chance of finding the optimal configuration within a limited budget.

Exam trap

The trap here is assuming that random search is 'good enough' for high-dimensional tuning or that grid search guarantees coverage — PMLE often tests whether candidates understand that Bayesian optimization is the sample-efficient choice when the trial budget is constrained.

How to eliminate wrong answers

Option A is wrong because grid search scales exponentially with dimensionality and would exhaust the 50-trial budget long before covering a meaningful fraction of the 5-dimensional space. Option B is wrong because simulated annealing is a local search heuristic not natively offered as a Vizier algorithm and is poorly suited to expensive black-box hyperparameter evaluations. Option D is wrong because random search ignores prior trial results and requires many more samples than Bayesian optimization to find competitive configurations.

55
MCQhard

You are scaling a prototype recommendation model to production on Vertex AI. The model is trained with a custom container and uses a large embedding table that must be updated frequently. You need to serve predictions with low latency and also support online updates to the embeddings without redeploying the model. Which Vertex AI feature should you use?

A.Deploy the model to a Vertex AI Endpoint with a custom container that loads the embedding table from a shared store such as Cloud Storage or a database at prediction time, and update the shared store independently.
B.Deploy the model to a Vertex AI Endpoint with a pre-built container and use Vertex AI Feature Store to serve the embeddings as features at prediction time.
C.Deploy the model to a Vertex AI Endpoint and use the endpoint's built-in support for model updates through Vertex AI Model Registry versioning.
D.Use Vertex AI Batch Prediction with a large batch size so that embeddings are computed in bulk and can be updated between batches.
AnswerA

A custom container can load or refresh the embedding table from an external store, allowing the embeddings to be updated without redeploying the model. This supports low-latency serving if the store is fast and the container caches or periodically refreshes the table. It is the standard way to combine a custom model with frequently changing embeddings on Vertex AI.

Why this answer

To serve low-latency predictions while allowing online embedding updates, the model should run in a custom container that can read and refresh the embedding table from an external store. This keeps the model deployed while the embeddings change independently. The other options either do not update internal embeddings, require redeployment, or are batch-oriented and therefore unsuitable for low-latency online serving.

Exam trap

The trap here is confusing feature serving with model-internal embedding updates, or assuming that Model Registry versioning gives you online weight updates without redeployment.

56
MCQeasy

You are deploying a scikit-learn model to Vertex AI for online prediction. The model expects a JSON payload with a single feature vector. You need to ensure the endpoint can handle bursts of traffic up to 1000 requests per second while maintaining low latency. What should you do?

A.Deploy the model to a Vertex AI Endpoint with a single machine type and enable autoscaling based on CPU utilization.
B.Deploy the model to a Vertex AI Endpoint with a fixed number of replicas equal to the expected peak traffic.
C.Use a batch prediction job to process incoming requests every minute.
D.Deploy the model to a Vertex AI Endpoint with a GPU machine type to accelerate inference.
AnswerA

Vertex AI Endpoint supports autoscaling, which automatically adjusts the number of replicas based on traffic. Setting a minimum and maximum replica count with CPU utilization as the metric allows the endpoint to handle bursts while keeping latency low. This is the standard and recommended approach for scaling online predictions.

Why this answer

Vertex AI Endpoints provide autoscaling to dynamically adjust replicas based on traffic. By enabling autoscaling with CPU utilization as the metric, the endpoint can scale out during bursts and scale in during low traffic, ensuring low latency and cost efficiency. This is the correct way to handle variable online prediction traffic.

Exam trap

The trap here is thinking that provisioning for peak load or using GPUs will solve latency issues, when the key is elastic scaling based on actual demand.

57
MCQmedium

A data scientist has a TensorFlow 2.x model trained on a single GPU. They want to scale training to multiple GPUs on a single Vertex AI machine without code changes. Which strategy should they use?

A.MultiWorkerMirroredStrategy
B.TPUStrategy
C.CentralStorageStrategy
D.MirroredStrategy
AnswerD

MirroredStrategy replicates the model across all GPUs on one machine using all-reduce synchronisation, distributing the existing single-GPU code across devices with no code changes. This satisfies the stem's single-machine, multi-GPU scaling requirement, unlike distributed strategies spanning multiple hosts.

Why this answer

MirroredStrategy is TensorFlow's default single-machine, multi-GPU strategy. It replicates the model on each GPU and synchronizes gradients via all-reduce, and it can be enabled with minimal code changes (typically just wrapping model creation in strategy.scope()), making it the right choice for scaling on one Vertex AI machine.

Exam trap

The trap is choosing MultiWorkerMirroredStrategy for a single-machine multi-GPU scenario — candidates confuse 'multiple GPUs' with 'multiple workers' and overlook that MirroredStrategy is the single-host default.

How to eliminate wrong answers

Option A is wrong because MultiWorkerMirroredStrategy is designed for multiple machines (workers), not a single machine with multiple GPUs. Option B is wrong because TPUStrategy targets Tensor Processing Units, not GPUs, and requires TPU-specific infrastructure. Option C is wrong because CentralStorageStrategy places variables on CPU or one GPU and is intended for scenarios where variables are too large to replicate — it is not the standard no-code-change path for multi-GPU scaling on one machine.

58
MCQmedium

You have a Python training script that reads a 500 GB CSV dataset from a Cloud Storage bucket. You submit a Vertex AI custom training job using a pre-built container, specifying a machine with 16 vCPUs and 60 GB RAM. The job fails after a few minutes with an out-of-memory error. You need to scale the prototype to handle this dataset without changing the model architecture. What should you do?

A.Increase the machine type to one with 128 GB RAM and retry the job.
B.Convert the CSV files to TFRecord format and use the tf.data API to stream the data in batches during training.
C.Split the CSV files into smaller files and use a larger number of workers with data parallelism.
D.Use Vertex AI Pipelines to preprocess the data and store it in BigQuery, then read from BigQuery during training.
AnswerB

Converting to TFRecord and using tf.data enables efficient streaming from Cloud Storage, so the full dataset is never loaded into memory. This directly addresses the out-of-memory error by decoupling data size from machine memory. It is the recommended approach for large-scale training on Vertex AI and requires only changing the input pipeline, not the model.

Why this answer

The out-of-memory error occurs because the training script attempts to load the entire 500 GB dataset into memory. Using TFRecord and tf.data allows streaming data in batches, so memory usage remains bounded. This is the standard method for scaling data input on Vertex AI and requires minimal changes to the training code.

Exam trap

The trap here is assuming that simply increasing machine memory or splitting files will solve the problem, when the real fix is to stream data instead of loading it all at once.

59
MCQmedium

You are preparing a Vertex AI custom training job that uses a custom Docker image built on top of the PyTorch pre-built training container. The image must be able to read training data from a Cloud Storage bucket without embedding credentials in the image. You want the job to run on a single NVIDIA T4 GPU. Which approach should you take?

A.Mount the Cloud Storage bucket as a local file system using Cloud Storage FUSE inside the container and rely on the bucket's default IAM permissions for anonymous access.
B.Hard-code a service account JSON key file into the Docker image and set GOOGLE_APPLICATION_CREDENTIALS in the Dockerfile.
C.Specify a user-managed service account with the Storage Object Viewer role when submitting the custom job, and let the container use Application Default Credentials to read from Cloud Storage.
D.Grant the Vertex AI Custom Code Service Agent the Storage Object Viewer role and let the training container use the default credentials from the metadata server.
AnswerC

When you submit a Vertex AI custom training job, you can attach a user-managed service account. The container then obtains credentials from the Compute Engine metadata server through Application Default Credentials. Granting that service account the Storage Object Viewer role allows the training code to read objects from Cloud Storage without embedding keys, and the job can be configured with a single T4 GPU accelerator.

Why this answer

A user-managed service account attached to the custom job provides the identity that the training container uses. Vertex AI injects credentials via the metadata server, so Application Default Credentials in the container can authenticate to Cloud Storage. Granting the least-privilege Storage Object Viewer role satisfies the data access requirement without embedding keys.

The job can also be configured with a single T4 GPU accelerator to meet the compute requirement.

Exam trap

The trap here is confusing the Vertex AI Custom Code Service Agent, which acts on behalf of the Vertex AI service, with the user-managed service account that actually runs inside the training container and accesses Cloud Storage.

60
MCQmedium

You are fine-tuning a pre-trained BERT model from Hugging Face on a custom text classification dataset using Vertex AI Training. You want to speed up training by using mixed precision. What should you do?

A.Modify the model to use half-precision layers
B.Use a custom container with TensorFlow instead of PyTorch
C.Enable mixed precision via Vertex AI hyperparameter tuning
D.Set fp16=True in the TrainingArguments
AnswerD

Setting fp16=True in TrainingArguments enables NVIDIA AMP mixed precision, using float16 for forward and backward passes while keeping master weights in float32. This reduces memory and speeds training on compatible GPUs without manual loss scaling.

Why this answer

Hugging Face's Trainer API exposes mixed precision through the TrainingArguments parameter fp16=True (or bf16=True for bfloat16). Setting this flag enables automatic mixed precision (AMP) via PyTorch's torch.cuda.amp, which casts eligible operations to FP16 while keeping master weights in FP32 for numerical stability. This is the standard, supported way to accelerate fine-tuning on GPU without rewriting the model.

Exam trap

The trap here is confusing framework-level training configuration (fp16=True in TrainingArguments) with infrastructure-level tuning (Vertex AI hyperparameter tuning) or model surgery (half-precision layers), when mixed precision is purely a Trainer API flag.

How to eliminate wrong answers

Option A is wrong because manually converting model layers to half-precision (e.g., .half()) discards FP32 master weights and typically causes NaN losses and degraded accuracy during training. Option B is wrong because switching frameworks from PyTorch to TensorFlow has no bearing on mixed precision and would require rewriting the training code. Option C is wrong because Vertex AI hyperparameter tuning searches over hyperparameter values; it does not enable mixed precision, which is a training configuration flag, not a tunable hyperparameter.

61
MCQmedium

An ML engineer is using Vertex AI Vizier to tune hyperparameters for a PyTorch model. They want to maximise the chance of finding the global optimum within a fixed trial budget of 50 trials. Which algorithm should they select?

A.Random search
B.Bayesian optimisation
C.Grid search
D.Evolutionary algorithm
AnswerB

Bayesian optimisation builds a probabilistic surrogate model of the objective and uses an acquisition function to pick each next trial, balancing exploration against exploitation. Over a fixed 50-trial budget this converges on the global optimum faster than grid or random search, directly satisfying the budget constraint.

Why this answer

Bayesian optimisation (option B) is the correct choice because it builds a probabilistic surrogate model of the objective function and uses an acquisition function to balance exploration and exploitation, making it highly sample-efficient. With only 50 trials, Bayesian optimisation maximises the probability of finding the global optimum by focusing trials on the most promising hyperparameter regions, unlike random or grid search which waste trials on unpromising areas.

Exam trap

The trap here is that candidates often choose random search (option A) because they recall it is better than grid search for high-dimensional spaces, but they overlook that Bayesian optimisation is strictly more sample-efficient and is the default recommendation in Vertex AI Vizier for maximising global optimum discovery under a fixed trial budget.

How to eliminate wrong answers

Option A is wrong because random search, while better than grid search in high-dimensional spaces, does not use past trial results to guide future trials, so it wastes trials on suboptimal regions and has a lower probability of finding the global optimum within a fixed budget of 50 trials. Option C is wrong because grid search exhaustively evaluates a fixed set of points, which scales exponentially with the number of hyperparameters and is extremely inefficient for more than a few parameters, often missing the global optimum entirely within a limited budget. Option D is wrong because evolutionary algorithms (e.g., genetic algorithms) are population-based and require many generations to converge, typically needing hundreds or thousands of trials to be effective, making them impractical for a tight budget of 50 trials.

62
MCQhard

Your team is training a very large transformer model that does not fit on a single GPU. They are using Vertex AI custom training with PyTorch. Which distributed training approach should they use?

A.Data parallelism using PyTorch DistributedDataParallel (DDP)
B.Horovod with allreduce
C.Model parallelism using pipeline parallelism
D.Vertex AI distributed training with TF_CONFIG
AnswerC

Pipeline parallelism splits the transformer's layers across GPUs, so each device holds only a subset of parameters and activations. This addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

Why this answer

When a transformer model is too large to fit on a single GPU, model parallelism (specifically pipeline parallelism) is required because it splits the model's layers across multiple devices, with each device holding a subset of the model's parameters. Data parallelism (DDP) replicates the entire model on each GPU, which fails if the model exceeds a single GPU's memory. Pipeline parallelism allows training very large models by partitioning the model into stages and passing activations and gradients sequentially between devices.

Exam trap

Google Cloud often tests the distinction between data parallelism (which replicates the model) and model parallelism (which splits the model), and the trap here is that candidates assume any distributed training framework (like DDP or Horovod) can handle oversized models, ignoring the fundamental memory constraint that data parallelism cannot overcome.

How to eliminate wrong answers

Option A is wrong because PyTorch DistributedDataParallel (DDP) implements data parallelism, which requires the entire model to fit on each GPU; if the model is too large for one GPU, DDP cannot be used. Option B is wrong because Horovod with allreduce is also a data-parallel approach that replicates the model on every worker, suffering the same memory limitation as DDP. Option D is wrong because Vertex AI distributed training with TF_CONFIG is a configuration mechanism for TensorFlow-based distributed training (using MirroredStrategy or MultiWorkerMirroredStrategy), not a PyTorch-native approach, and it still relies on data parallelism unless combined with model parallelism; the question specifies PyTorch, making this option technically incompatible.

63
MCQhard

A data engineering team needs to compute rolling window features (7-day average, 30-day sum) from a high-volume stream of e-commerce events stored in BigQuery. They must output the features to Vertex AI Feature Store for online serving. Which approach is MOST cost-effective and scalable?

A.Use Cloud Composer (Airflow) with a daily DAG to run SQL queries on BigQuery and export results
B.Schedule a query in BigQuery using scheduled queries and export results to Feature Store
C.Use Cloud Functions triggered by Pub/Sub to compute features on the fly
D.Use Dataflow with Apache Beam, reading from BigQuery, computing windowed aggregations, and writing to Vertex AI Feature Store
AnswerD

Dataflow with Apache Beam provides managed, autoscaling stream processing that reads from BigQuery, applies windowed aggregations for the 7-day and 30-day features, and writes directly to Vertex AI Feature Store. This avoids custom serving infrastructure while meeting the online-serving and scalability constraints.

Why this answer

Dataflow (Apache Beam) is ideal for processing large-scale batch and streaming data. It can read from BigQuery, perform windowing computations, and write to Feature Store's online store. Cloud Functions have timeouts, Cloud Composer is not optimal for streaming, and BigQuery scheduled queries are not designed for streaming-feature computation.

64
MCQmedium

You have a trained XGBoost model that you want to deploy on Vertex AI for online prediction. The model expects input features in a specific order and requires a custom preprocessing step that normalizes numerical features using statistics computed during training. You need to ensure that the same preprocessing is applied at serving time. What should you do?

A.Use Vertex AI's built-in preprocessing feature by specifying a preprocessing function in the model's metadata when uploading.
B.Export the model as a SavedModel and include the preprocessing logic in the model's serving signature using TensorFlow's preprocessing layers.
C.Preprocess the input data on the client side before sending requests to the Vertex AI endpoint, using the same statistics.
D.Deploy the model using a custom container that includes the preprocessing code and the trained model artifacts, and implement the preprocessing in the container's prediction handler.
AnswerD

A custom container gives you full control over the serving logic. You can load the model and the training statistics (e.g., mean and standard deviation) and apply the exact same preprocessing in the prediction handler before passing data to the model. This ensures consistency between training and serving, which is critical for model performance.

Why this answer

For custom preprocessing with non-TensorFlow models like XGBoost, a custom container is the most reliable way to ensure that the exact same preprocessing is applied at serving time. You can embed the training statistics and logic in the container, maintaining consistency and simplifying client requests.

Exam trap

The trap here is assuming that Vertex AI provides built-in preprocessing for custom models, when in fact you must implement it yourself, often via a custom container.

65
MCQmedium

You are training a large tabular model on Vertex AI using a custom training job. The dataset is stored in BigQuery and is several terabytes in size. Training reads the same data for many epochs, and reading directly from BigQuery each epoch is slow and expensive. You want to maximize training throughput while keeping the data accessible to the training container. What should you do?

A.Enable Vertex AI Pipelines caching so that the BigQuery read step is not re-executed on subsequent training runs.
B.Export the BigQuery table to Cloud Storage in TFRecord or Parquet format, then read the exported files from the training container using the appropriate data loader.
C.Increase the number of worker replicas in the training job so that each worker reads a smaller portion of the BigQuery table in parallel.
D.Keep reading directly from BigQuery each epoch and rely on the BigQuery Storage Read API to stream rows to the training container.
AnswerB

Exporting to Cloud Storage once and reading the files repeatedly avoids repeated BigQuery scans and network round trips, and Cloud Storage offers high aggregate throughput for large sequential reads. This is the standard pattern for scaling prototype training on Vertex AI when the source is BigQuery and the job needs multiple epochs over multi-terabyte data.

Why this answer

The most efficient approach is to materialize the BigQuery data in Cloud Storage once and then have the training container read those files for every epoch. This removes repeated BigQuery scans, reduces cost and latency, and lets the training job use high-throughput file readers. The other choices either keep the expensive repeated reads or add complexity without fixing the core bottleneck.

Exam trap

The trap here is assuming that a newer BigQuery API or more workers will make repeated full-table reads cheap, when the real win is exporting once and reusing the data locally.

66
MCQeasy

An ML team wants to use Vertex AI Hyperparameter Tuning to tune a custom training job. They have a budget of 50 trials and want to use an algorithm that balances exploration and exploitation. Which algorithm should they choose?

A.Random search
B.Grid search
C.Bayesian optimization (Vizier default)
D.Manual search
AnswerC

Bayesian optimization (Vizier's default) models the objective probabilistically, using prior trial results to pick promising configurations while still sampling uncertain regions — directly balancing exploration and exploitation. It converges within the 50-trial budget far more efficiently than grid or random search, satisfying the stated constraint.

Why this answer

Bayesian optimization (the default algorithm in Vertex AI Vizier) is the correct choice because it explicitly balances exploration and exploitation by building a probabilistic model of the objective function and using an acquisition function to select the next hyperparameter configuration. With a budget of 50 trials, this algorithm efficiently converges to optimal regions while still exploring uncertain areas, making it ideal for tuning custom training jobs where each trial is computationally expensive.

Exam trap

A common pitfall is assuming that random search is the best default for balancing exploration and exploitation. However, random search lacks any exploitation mechanism, making Bayesian optimization the correct choice for efficient tuning within a constrained budget, as emphasized in Google PMLE.

How to eliminate wrong answers

Option A is wrong because random search does not balance exploration and exploitation; it samples hyperparameters uniformly at random without using past trial results to guide future selections, which wastes budget on suboptimal regions. Option B is wrong because grid search exhaustively evaluates a fixed set of hyperparameter combinations, which is computationally inefficient for a budget of 50 trials and does not incorporate any exploitation mechanism. Option D is wrong because manual search relies on human intuition and ad-hoc adjustments, which is not an automated algorithm and cannot systematically balance exploration and exploitation within a defined trial budget.

67
MCQeasy

You want to use a pre-trained model from TensorFlow Hub for image classification, but you need to adapt it to classify your own custom categories with a small dataset. Which Vertex AI approach is most appropriate?

A.Write a custom training script that loads the pre-trained model and fine-tunes it on your dataset
B.Deploy the pre-trained model as-is via Vertex AI JumpStart
C.Build a custom container with the pre-trained model and deploy to Vertex AI Endpoints
D.Use Vertex AI AutoML for image classification
AnswerA

Fine-tuning loads the TensorFlow Hub pre-trained model's weights into a custom training script and continues training on the small labelled dataset, adapting output categories. This suits limited data far better than training from scratch, which would require far more examples.

Why this answer

Fine-tuning a pre-trained model on your custom dataset is the most appropriate approach when you have a small dataset and need to adapt the model to new categories. This leverages transfer learning, where the pre-trained weights are used as a starting point and updated with your data.

Exam trap

The trap is confusing deployment with adaptation. Candidates may think that deploying a pre-trained model via JumpStart or a custom container will somehow adapt it to new categories, but adaptation requires training/fine-tuning.

How to eliminate wrong answers

Option B is wrong because deploying the pre-trained model as-is would not classify your custom categories; it would only recognize the original classes. Option C is wrong because building a custom container with the pre-trained model and deploying it does not adapt the model to your categories; it just serves the original model. Option D is wrong because AutoML for image classification builds a model from scratch (though it may use some transfer learning internally), but it does not allow you to start from a specific pre-trained model from TensorFlow Hub.

68
MCQmedium

You want to reduce training costs by using preemptible VMs on Vertex AI for a fault-tolerant distributed training job that uses checkpointing. Which machine type should you choose in the worker pool configuration?

A.Use spot VMs by setting 'spot' to true in the machine spec
B.Use custom machine types with preemptible flag
C.Use standard VMs and rely on Vertex AI auto-restart
D.Use TPU VMs because they are cheaper
AnswerA

Setting spot to true in the machine spec provisions spot (preemptible) VMs, which cost substantially less than standard VMs. Checkpointing makes the fault-tolerant job resilient to preemption, satisfying the requirement to reduce training costs while tolerating interruptions.

Why this answer

Vertex AI supports spot VMs (the successor to preemptible VMs) by setting the 'spot' field to true in the machine spec of the worker pool. Spot VMs offer up to 60-91% discounts and are suitable for fault-tolerant jobs with checkpointing, since they can be preempted with 30 seconds' notice. This is the documented, supported configuration for cost-optimized training.

Exam trap

The trap is using the legacy GCE term 'preemptible' or assuming auto-restart alone provides cost savings, when Vertex AI requires the explicit 'spot: true' field in the machine spec and checkpointing for fault tolerance.

How to eliminate wrong answers

Option B is wrong because 'preemptible' is the legacy GCE term; Vertex AI custom training uses the 'spot' boolean in the machine spec, and custom machine types are about vCPU/memory sizing, not preemption. Option C is wrong because standard VMs are on-demand and do not provide the cost savings of spot VMs; auto-restart does not change the billing model. Option D is wrong because TPU VMs are a different accelerator family and are not inherently cheaper than spot GPU/CPU VMs; choosing TPUs requires code compatibility with XLA and is not a drop-in cost optimization.

69
MCQmedium

You are deploying a pre-trained BERT model for inference on edge devices. The model must be under 500 MB and inference latency under 50 ms. Which approach should you take?

A.Use a larger model like BERT-Large and deploy on GPU
B.Apply post-training INT8 quantization using TensorFlow Lite
C.Prune 50% of the model weights and fine-tune
D.Use knowledge distillation to train a smaller student model from scratch
AnswerB

Post-training INT8 quantization shrinks the BERT weights roughly fourfold, bringing the model comfortably under the 500 MB ceiling, while integer arithmetic accelerates inference to meet the 50 ms latency budget. TensorFlow Lite's runtime is purpose-built for edge deployment, so no retraining is required.

Why this answer

Post-training INT8 quantization reduces model size by approximately 75% (from ~440 MB to ~110 MB for BERT-Base) and accelerates inference on edge devices via integer arithmetic, easily meeting the 500 MB and 50 ms constraints. TensorFlow Lite provides hardware-optimized kernels for ARM CPUs and NPUs, making it ideal for edge deployment without requiring retraining.

Exam trap

A common pitfall is assuming that only pruning or distillation can reduce model size, but post-training INT8 quantization directly shrinks the model and speeds inference without retraining, which is ideal for deploying a pre-trained model on edge devices.

How to eliminate wrong answers

Option A is wrong because BERT-Large is ~1.3 GB, far exceeding the 500 MB limit, and GPU deployment is not feasible on most edge devices due to power and thermal constraints. Option C is wrong because pruning 50% of weights without fine-tuning would cause catastrophic accuracy loss, and fine-tuning requires the original training data and compute, which may not be available; even with fine-tuning, pruned models often need specialized hardware for speedup. Option D is wrong because knowledge distillation requires training a smaller student model from scratch, which demands significant compute, time, and access to the teacher model's logits, making it impractical for a quick deployment scenario where a pre-trained BERT model is already available.

70
Multi-Selectmedium

You are using tf.Transform to preprocess data at scale. Which TWO services are required to run tf.Transform on Google Cloud? (Choose 2)

Select 2 answers
A.Cloud Functions
B.Dataflow
C.Cloud Storage
D.Vertex AI Training
E.BigQuery
AnswersB, C

Dataflow is required to execute the Apache Beam pipeline that tf.Transform generates, running the preprocessing at scale. It satisfies the stem's requirement by providing the distributed runner that applies the transform over large datasets, complementing the service that analyses and materialises the transform artefacts.

Why this answer

tf.Transform requires Apache Beam for execution, which on GCP is typically run on Dataflow. The processed data and transform artifacts are stored in Cloud Storage.

71
MCQhard

A machine learning engineer is training a model on Vertex AI using a custom container. The training job uses a large dataset stored in BigQuery. The engineer wants to minimize data transfer costs and maximize training speed. Which of the following approaches is most efficient?

A.Use the BigQuery Storage API to read data directly into the training program with parallel streams.
B.Run a query in BigQuery to extract the data, then save it to a local SSD on the training VM.
C.Export the BigQuery table to Cloud Storage in TFRecord format, then read it using the tf.data API.
D.Use the BigQuery client library to run a query and fetch results into memory using the to_dataframe() method.
AnswerA

The BigQuery Storage API provides high-throughput, parallel access to BigQuery data directly from the training program. It avoids exporting data to Cloud Storage, reducing costs and latency. It also supports column filtering and predicate pushdown, which can reduce the amount of data read. This is the most efficient method for training on large BigQuery datasets.

Why this answer

The BigQuery Storage API is designed for high-performance data reading, offering parallel streams and efficient data transfer. It avoids the overhead of exporting data to Cloud Storage or loading into memory all at once. This makes it the most efficient and cost-effective approach for training on large BigQuery datasets.

Exam trap

The trap here is assuming that exporting to Cloud Storage is always necessary, but the BigQuery Storage API allows direct, efficient access.

72
MCQeasy

You want to use Vertex AI Vizier for hyperparameter tuning. You have 2 categorical parameters and 3 continuous parameters. Which algorithm is best suited for this mixed parameter space?

A.Evolutionary algorithm
B.Random search
C.Bayesian optimization
D.Grid search
AnswerC

Bayesian optimisation handles mixed categorical and continuous search spaces by building a probabilistic surrogate model and selecting promising trials. It suits Vertex AI Vizier's 2 categorical and 3 continuous parameters, converging on good configurations with fewer trials than grid or random search.

Why this answer

Bayesian optimization is the default and best-suited algorithm in Vertex AI Vizier for mixed parameter spaces containing both categorical and continuous parameters. It builds a probabilistic surrogate model of the objective function and uses an acquisition function to intelligently select the next hyperparameter configuration, handling categorical and continuous dimensions natively. This makes it far more sample-efficient than random or grid search when the search space is mixed and potentially high-dimensional.

Exam trap

The trap here is assuming that any algorithm works equally well for mixed parameter spaces; candidates often pick random search because it is simple, but the exam expects recognition that Bayesian optimization is the default and most efficient choice in Vizier for mixed categorical/continuous spaces.

How to eliminate wrong answers

Option A is wrong because the evolutionary algorithm in Vizier is designed for large-scale, high-dimensional search spaces and is typically recommended when the number of trials is very large or the search space is complex; it is not the default best choice for a small mixed space of 5 parameters. Option B is wrong because random search ignores the structure of the search space and does not learn from previous trials, making it far less sample-efficient than Bayesian optimization. Option D is wrong because grid search is not a Vizier-supported algorithm and would scale exponentially with the number of parameters, making it impractical for continuous parameters.

73
MCQmedium

An ML engineer is training a PyTorch model on Vertex AI using a custom training job. The dataset is stored in a Cloud Storage bucket with 500,000 small JPEG images. The training job is configured with an n1-standard-8 machine and a single NVIDIA T4 GPU. The engineer observes that GPU utilization is very low (around 15%) and training is slow. The model code is not the bottleneck. What is the most likely cause and the best solution?

A.The NVIDIA T4 GPU is not powerful enough for this model; upgrade to an NVIDIA A100 GPU.
B.The training job is using a single worker and not distributing the load; switch to distributed training with multiple workers.
C.The Cloud Storage bucket is in a different region than the training job; move the bucket to the same region.
D.The data loading pipeline is inefficient due to many small files; convert the dataset to TFRecord or WebDataset format and use a larger batch size with prefetching.
AnswerD

Reading many small JPEGs individually from Cloud Storage introduces high per-file latency and I/O overhead, starving the GPU. Converting to a sharded, sequential format like TFRecord or WebDataset allows efficient streaming, and using prefetching and larger batches keeps the GPU fed, increasing utilization.

Why this answer

Low GPU utilization with many small files typically points to an I/O-bound data pipeline. Cloud Storage has high per-file access latency, so reading hundreds of thousands of small JPEGs individually creates a bottleneck. Converting to a sharded format like TFRecord or WebDataset enables efficient sequential reads, and combining with prefetching and larger batches ensures the GPU is continuously fed, improving utilization and training speed.

Exam trap

The trap here is assuming that low GPU utilization always means insufficient GPU compute and upgrading the GPU, instead of diagnosing the data input pipeline.

74
MCQmedium

An ML team is building a feature pipeline with Dataflow that reads from BigQuery, computes features, and writes to Vertex AI Feature Store. They need to ensure that features are available for both training and serving with low latency. Which Feature Store option should they use?

A.Create a featurestore with only offline serving
B.Store features directly in BigQuery
C.Use Cloud SQL as a feature store
D.Create a featurestore with online serving enabled
AnswerD

Enabling online serving on the featurestore provisions a low-latency endpoint that serves features in real time, while the same store retains the data for training retrieval. This dual capability satisfies both training and serving requirements from one featurestore.

Why this answer

To ensure features are available for both training and serving with low latency, the featurestore must have online serving enabled. Vertex AI Feature Store provides both online serving (low-latency reads for real-time predictions) and offline serving (batch reads for training). Enabling online serving allows the same feature values to be served at low latency during prediction, ensuring consistency between training and serving.

Exam trap

The trap is assuming that storing features in BigQuery or Cloud SQL is sufficient for low-latency serving; candidates may overlook that Vertex AI Feature Store's online serving is specifically designed for millisecond-latency access and integration with training pipelines.

How to eliminate wrong answers

Option A is wrong because a featurestore with only offline serving cannot provide low-latency online serving for real-time predictions, failing the requirement. Option B is wrong because storing features directly in BigQuery does not provide the low-latency online serving needed for real-time inference; BigQuery is optimized for analytical queries, not millisecond-latency lookups. Option C is wrong because Cloud SQL is a relational database not designed as a feature store; it lacks the integration with Vertex AI and the optimized online serving capabilities of Feature Store.

75
Multi-Selectmedium

You are designing a distributed training job for a very large neural network that does not fit on a single machine. You need to split the model across multiple devices. Which TWO techniques can you use?

Select 2 answers
A.ParameterServerStrategy
B.Pipeline parallelism
C.Operator-level model parallelism
D.Data parallelism with MirroredStrategy
E.MultiWorkerMirroredStrategy
AnswersB, C

Pipeline parallelism divides the model into sequential stages placed on different devices, with micro-batches flowing between them. This splits the model across devices, satisfying the stem's constraint that the neural network does not fit on a single machine.

Why this answer

Pipeline parallelism (B) is correct because it splits the model's layers into sequential stages placed on different devices, so a model too large for one machine can be distributed across devices while micro-batches flow through the pipeline. Operator-level model parallelism (C) is also correct because it partitions individual operations (e.g., splitting a large matrix multiplication or a layer's weights) across devices, which directly addresses a model that does not fit on a single machine. ParameterServerStrategy (A) is a data-parallel approach that replicates the full model on each worker and only shards the parameter updates, so it does not solve the memory problem of a model too large for one device.

Data parallelism with MirroredStrategy (D) likewise replicates the entire model on every device, which is infeasible when the model does not fit on one machine. MultiWorkerMirroredStrategy (E) is also a data-parallel strategy that requires the full model to fit on each worker, so it does not satisfy the requirement.

Exam trap

The trap is confusing data parallelism with model parallelism; candidates may pick ParameterServerStrategy or MirroredStrategy because they are familiar distributed training strategies, but the exam expects recognition that only pipeline and operator-level parallelism split the model across devices.

Page 1 of 2 · 103 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Pmle Scaling Models questions.