Courseiva

CCNA Scaling Prototypes into ML Models Questions

28 of 103 questions · Page 2/2 · Scaling Prototypes into ML Models · Answers revealed

76
MCQmedium

An engineer is training a model on Vertex AI using a custom container. The training job fails with an error indicating that the container exited with a non-zero status. The engineer wants to debug the issue. What is the best way to access the logs?

A.SSH into the training container using Vertex AI's SSH feature
B.View logs in Cloud Storage under the job's output directory
C.Use Cloud Debugger to inspect the container
D.Check the logs in Cloud Logging (Logs Explorer)
AnswerD

Vertex AI writes custom container training output, including stderr and the non-zero exit trace, to Cloud Logging. Logs Explorer surfaces those entries for the failed job, satisfying the debugging requirement by exposing the container's actual error rather than only the job's status.

Why this answer

Vertex AI automatically streams all container stdout and stderr to Cloud Logging (Logs Explorer). When a custom container exits with a non-zero status, the detailed error messages, stack traces, and application logs are captured there, making it the primary and most comprehensive debugging tool. Cloud Logging provides structured, searchable logs without requiring direct access to the container.

Exam trap

A common trap is the misconception that you can SSH into a training container or that logs are stored in Cloud Storage, when in fact Cloud Logging is the centralized, default logging solution for all Vertex AI training jobs.

How to eliminate wrong answers

Option A is wrong because Vertex AI does not provide an SSH feature for training containers; training jobs run in ephemeral, isolated environments with no interactive shell access. Option B is wrong because Cloud Storage under the job's output directory stores artifacts like model checkpoints and metrics, not real-time container logs or error messages. Option C is wrong because Cloud Debugger is designed for debugging running applications in production by capturing snapshots and variable states, not for inspecting container exit errors or retrieving logs from a failed training job.

77
Multi-Selecteasy

A company wants to use Vertex AI for hyperparameter tuning. Which three components are required to configure a hyperparameter tuning job? (Choose THREE.)

Select 3 answers
A.Algorithm (e.g., Bayesian, grid, random)
B.Machine type for each trial
C.List of hyperparameters with types and ranges
D.Objective metric name and goal (minimize or maximize)
E.Training container image
AnswersA, C, D

Required to specify how to search.

Why this answer

Vertex AI hyperparameter tuning requires specifying the search algorithm (Bayesian, grid, or random) to determine how the hyperparameter space is explored. Bayesian optimization is the default and most efficient for continuous spaces, while grid search is exhaustive and random search is simple. Without an algorithm, Vertex AI cannot decide how to sample trials.

Exam trap

Candidates often mistakenly include machine type or container image as mandatory tuning parameters, but these are optional training job settings.

78
MCQmedium

An ML engineer is using Vertex AI Training to fine-tune a large image classification model on a dataset stored in Cloud Storage. The training job uses a custom container and runs on a single NVIDIA V100 GPU. The engineer notices that GPU utilization is consistently low (around 20%) and training is slow. The data is stored as many small JPEG files. What should the engineer do to improve GPU utilization and training speed?

A.Enable Vertex AI Vizier to automatically tune hyperparameters and improve training speed.
B.Use a larger GPU instance with more memory, such as an NVIDIA A100.
C.Convert the dataset to TFRecord format and use the tf.data API with parallel interleaving and prefetching.
D.Increase the batch size to the maximum that fits in GPU memory.
AnswerC

TFRecord is a binary format that stores data efficiently and allows sequential reads, which is much faster than reading many small JPEG files individually. Using tf.data with interleave, map with num_parallel_calls, and prefetch overlaps data loading and preprocessing with GPU computation. This directly addresses the input pipeline bottleneck, leading to higher GPU utilization and faster training.

Why this answer

Low GPU utilization during training often indicates that the GPU is starved for data. When data is stored as many small files, the I/O overhead dominates. Converting to TFRecord, a binary format optimized for TensorFlow, and using tf.data with parallel reads and prefetching can dramatically speed up data loading.

This ensures the GPU receives data as fast as it can process it, improving utilization and reducing training time.

Exam trap

The trap here is attributing low GPU utilization to insufficient GPU power or needing hyperparameter tuning, rather than diagnosing the input pipeline bottleneck.

79
MCQeasy

You need to run a custom training job on Vertex AI using a pre-built container for scikit-learn. Which container image should you specify?

A.us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.1-9
B.us-docker.pkg.dev/vertex-ai/training/scikit-learn-cpu.0.23-0
C.us-docker.pkg.dev/vertex-ai/training/xgboost-cpu.1-3
D.us-docker.pkg.dev/vertex-ai/training/tf-cpu.2-6
AnswerB

Vertex AI provides pre-built training containers hosted in its Artifact Registry; the scikit-learn CPU image at that path bundles the framework and Python dependencies needed to run custom scikit-learn training code. Specifying it satisfies the stem's requirement to use a pre-built container rather than building your own.

Why this answer

For a scikit-learn custom training job on Vertex AI, you must specify a pre-built container image that includes scikit-learn. The correct image is us-docker.pkg.dev/vertex-ai/training/scikit-learn-cpu.0.23-0, which is the official Vertex AI pre-built training container for scikit-learn version 0.23 on CPU. Other images correspond to different frameworks (PyTorch, XGBoost, TensorFlow).

Exam trap

PMLE often tests the ability to match the correct pre-built container image to the framework, causing candidates to confuse scikit-learn with XGBoost or TensorFlow images.

How to eliminate wrong answers

Option A is wrong because pytorch-gpu.1-9 is a PyTorch GPU training container, not scikit-learn. Option C is wrong because xgboost-cpu.1-3 is an XGBoost training container, not scikit-learn. Option D is wrong because tf-cpu.2-6 is a TensorFlow CPU training container, not scikit-learn.

80
MCQhard

You are training a PyTorch model on Vertex AI using a custom container. The training script uses DistributedDataParallel (DDP) with NCCL backend across 4 nodes, each with 8 GPUs. You notice that training throughput is low and GPUs are underutilized. After profiling, you find that the data loading is the bottleneck. You need to improve throughput without changing the model. What should you do?

A.Increase the number of DataLoader worker processes and use a larger prefetch factor.
B.Reduce the batch size per GPU to decrease memory pressure and allow more concurrent kernels.
C.Switch from NCCL to Gloo backend for inter-node communication.
D.Use a larger machine type with more vCPUs to increase data loading throughput.
AnswerA

Increasing DataLoader workers and prefetch factor allows more parallel data loading and preloading of batches, reducing GPU idle time. This directly addresses the data loading bottleneck. It is a standard PyTorch optimization that requires minimal code changes and can significantly improve throughput in distributed training.

Why this answer

The data loading bottleneck in PyTorch distributed training is often alleviated by increasing the number of DataLoader worker processes and prefetch factor. This enables parallel data fetching and preloading of batches, keeping GPUs busy. Other options either do not address the bottleneck or could worsen performance.

Exam trap

The trap here is assuming that more vCPUs or a different communication backend will solve the issue, when the real fix is to tune the DataLoader to feed data faster.

81
MCQmedium

You have a prototype model trained on a single machine using scikit-learn. You now need to scale training to a larger dataset that does not fit in memory on one machine. You want to use Vertex AI training with minimal changes to your existing scikit-learn code. What should you do?

A.Use Vertex AI Training with a custom container that runs your scikit-learn script, and use a distributed framework such as Dask or Ray to partition the data across multiple workers.
B.Convert the scikit-learn model to a TensorFlow model and use MirroredStrategy to train across multiple GPUs.
C.Use Vertex AI Training with the pre-built scikit-learn container and set the worker count to more than one; Vertex AI will automatically distribute the scikit-learn training.
D.Use Vertex AI distributed training with a ParameterServerStrategy and wrap your scikit-learn estimator in a TensorFlow estimator.
AnswerA

Vertex AI custom training lets you bring your own container and code. By using a distributed framework like Dask or Ray inside the container, you can partition the dataset across workers and train a scikit-learn model in a distributed manner with relatively small changes to the original script. This directly addresses the out-of-memory issue while staying close to the existing code.

Why this answer

The best approach is to keep the scikit-learn code and run it in a custom container on Vertex AI, using a distributed framework like Dask or Ray to partition the data across workers. This scales beyond a single machine's memory while minimizing changes to the original training script. The other options either rely on incompatible distribution strategies, assume automatic distribution that does not exist, or require a full rewrite.

Exam trap

The trap here is assuming that Vertex AI automatically distributes any training code when you set multiple workers, or that TensorFlow distribution strategies apply to scikit-learn.

82
MCQhard

A team is training a large recommendation model on Vertex AI using a custom container. They need to log training metrics and visualize them in Vertex AI TensorBoard. The training code is written in PyTorch and runs on multiple worker nodes. Which of the following is the correct way to enable TensorBoard logging?

A.Install TensorBoard in the custom container and run it as a sidecar process on each worker node.
B.Use the Vertex AI SDK to create a TensorBoard instance and pass its resource name to the training job using the --tensorboard flag.
C.Write training metrics to a Cloud Storage bucket and then manually upload them to Vertex AI TensorBoard after training.
D.Use the torch.utils.tensorboard.SummaryWriter to write logs to a local directory, and Vertex AI will automatically sync it to TensorBoard.
AnswerB

The Vertex AI SDK allows creating a TensorBoard instance, and the training job can be configured with the tensorboard resource name. This enables automatic logging of metrics from the training container to the TensorBoard instance. The training code should write logs to the directory specified by the AIP_TENSORBOARD_LOG_DIR environment variable, which Vertex AI sets.

Why this answer

To integrate with Vertex AI TensorBoard, you must create a TensorBoard instance and pass its resource name to the training job. The training code should write logs to the directory specified by the AIP_TENSORBOARD_LOG_DIR environment variable. This allows Vertex AI to automatically collect and display metrics from all workers in the managed TensorBoard service.

Exam trap

The trap here is assuming that any local logging will be automatically synced to Vertex AI TensorBoard, but explicit configuration of the TensorBoard instance is required.

83
MCQhard

A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?

A.Use Vertex AI Hyperparameter Tuning to find optimal model partitioning
B.Use tf.distribute.MirroredStrategy across all GPUs
C.Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM
D.Use Vertex AI distributed training with TF_CONFIG to set up multi-worker mirrored strategy and rely on XLA to partition the model
AnswerC

Pipeline parallelism partitions consecutive model layers into stages across GPUs, so each device holds only a fraction of parameters. This directly addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

Why this answer

When a model does not fit into a single GPU's memory, model parallelism is required. Pipeline parallelism splits the model layers across multiple GPUs, and frameworks like PyTorch's RPC or Megatron-LM provide the necessary primitives to implement it. This is the correct strategy for fitting a large model like LLaMA 2 across multiple GPUs.

Exam trap

PMLE often tests the confusion between data parallelism (replicating the model) and model parallelism (splitting the model), causing candidates to choose MirroredStrategy when the model does not fit on a single GPU.

How to eliminate wrong answers

Option A is wrong because Vertex AI Hyperparameter Tuning is for tuning hyperparameters, not for partitioning a model across GPUs. Option B is wrong because tf.distribute.MirroredStrategy replicates the entire model on each GPU (data parallelism), which does not solve the memory issue when the model does not fit on one GPU. Option D is wrong because TF_CONFIG with multi-worker mirrored strategy also replicates the model and relies on XLA, which does not automatically partition a model across GPUs for model parallelism.

84
MCQeasy

You are fine-tuning a BERT model from Hugging Face Transformers on Vertex AI. You want to minimise cost for a short experiment. Which compute configuration should you use?

A.A custom training job with a single NVIDIA T4 GPU using spot VMs
B.A custom training job with a TPU v3-8 pod
C.A custom training job with 8 NVIDIA V100 GPUs using regular VMs
D.A standard n1-highmem-8 machine with no accelerator
AnswerA

A single T4 GPU on spot VMs gives the lowest cost for a short fine-tuning experiment. Spot capacity suits interruptible, brief jobs, and T4 provides sufficient memory and compute for BERT-scale training without paying for premium accelerators.

Why this answer

A single NVIDIA T4 GPU with spot VMs is the most cost-effective choice for a short BERT fine-tuning experiment on Vertex AI. T4 GPUs are inexpensive and well-suited for moderate training workloads, and spot VMs offer up to 60-70% discount over regular VMs. Since the experiment is short, the risk of preemption is acceptable, and the cost savings are significant.

This configuration balances performance and cost effectively.

Exam trap

PMLE often tests the misconception that more powerful hardware (TPUs or multiple high-end GPUs) is always better, but the key is matching compute to workload scale and cost constraints; candidates may overlook spot VMs as a cost-saving option for short, fault-tolerant jobs.

How to eliminate wrong answers

Option B is wrong because a TPU v3-8 pod is significantly more expensive and overpowered for a short BERT fine-tuning experiment; TPUs are optimized for large-scale training and require specific code adjustments, adding complexity and cost. Option C is wrong because using 8 NVIDIA V100 GPUs with regular VMs incurs very high costs and is unnecessary for a short experiment; V100s are powerful but overkill, and regular VMs lack the cost savings of spot instances. Option D is wrong because a standard n1-highmem-8 machine without any accelerator would be extremely slow for BERT fine-tuning, as BERT training benefits greatly from GPU acceleration; using only CPU would prolong training time and increase overall cost despite lower hourly rates.

85
MCQhard

You are training a TensorFlow model on Vertex AI using a custom container. The training job uses a single node with 4 GPUs and a global batch size of 1024. You notice that the training is slower than expected and GPU utilization is low. You suspect the input pipeline is the bottleneck. Which of the following should you do to improve training throughput?

A.Reduce the global batch size to 512 to decrease memory pressure on the GPUs.
B.Use tf.data to prefetch data and parallelize data extraction and transformation with num_parallel_calls and prefetch.
C.Increase the number of GPUs to 8 and double the global batch size to 2048.
D.Switch to using TPUs instead of GPUs for the training job.
AnswerB

Optimizing the input pipeline with tf.data by using prefetch and parallel calls allows data to be prepared while the GPU is training. This overlaps I/O and preprocessing with computation, increasing GPU utilization and throughput. It directly addresses the bottleneck by ensuring the GPU is not waiting for data, which is a common cause of low utilization in multi-GPU training.

Why this answer

Low GPU utilization in multi-GPU training often indicates the input pipeline cannot supply data fast enough. Using tf.data with prefetch and parallel calls overlaps data preprocessing with model training, keeping the GPUs busy. This is a standard optimization for TensorFlow input pipelines and directly targets the bottleneck without changing hardware or batch size.

Exam trap

The trap here is assuming that adding more hardware or changing batch size will solve low GPU utilization, when the real issue is data starvation from an inefficient input pipeline.

86
MCQeasy

A data engineer wants to compute feature aggregates over a large dataset stored in BigQuery and write the results to Vertex AI Feature Store. The pipeline must handle both batch and streaming data. Which Google Cloud service should they use?

A.BigQuery scheduled queries
B.Cloud Functions triggered by Pub/Sub
C.Cloud Dataproc with Spark
D.Cloud Dataflow with Apache Beam
AnswerD

Cloud Dataflow with Apache Beam provides unified batch and streaming pipelines, reading from BigQuery and writing aggregates to Vertex AI Feature Store. This satisfies the stem's requirement to handle both data modes within one managed service.

Why this answer

Cloud Dataflow with Apache Beam is a unified stream and batch data processing service. It can read from BigQuery, compute aggregates, and write to Vertex AI Feature Store, handling both batch and streaming data with the same pipeline code. This makes it the ideal choice for the requirement.

Exam trap

PMLE often tests the choice between batch-only and unified processing services, and candidates may pick BigQuery scheduled queries or Dataproc for streaming, missing Dataflow's unified capability.

How to eliminate wrong answers

Option A is wrong because BigQuery scheduled queries only handle batch processing and cannot process streaming data. Option B is wrong because Cloud Functions triggered by Pub/Sub is for lightweight event-driven processing, not large-scale feature aggregation. Option C is wrong because Cloud Dataproc with Spark is primarily for batch processing and does not natively handle streaming as seamlessly as Dataflow.

87
MCQhard

An ML engineer is training a TensorFlow model on Vertex AI using a custom training job with a single worker and multiple GPUs. The training script uses tf.distribute.MirroredStrategy. After a few epochs, the job fails with a NCCL timeout error. The engineer confirms the GPUs are healthy and the batch size is reasonable. What should they do to resolve the error?

A.Increase the number of workers in the worker pool configuration to match the number of GPUs.
B.Set the environment variable NCCL_SOCKET_IFNAME to the correct network interface and ensure the container image includes a compatible NCCL version.
C.Set the environment variable NCCL_DEBUG=INFO and reduce the per-GPU batch size to lower memory pressure.
D.Ensure the training job is configured with a single replica and that the container sets the TF_CONFIG environment variable appropriately for a single-worker, multi-GPU setup.
AnswerB

NCCL timeouts on a single machine with multiple GPUs often occur when NCCL cannot determine the correct network interface or when there is a version mismatch between NCCL and the GPU driver or framework. Setting NCCL_SOCKET_IFNAME to the primary network interface (e.g., eth0) and using a pre-built container with a validated NCCL version ensures proper inter-GPU communication and resolves the timeout.

Why this answer

NCCL timeout errors in a single-worker, multi-GPU Vertex AI training job usually stem from NCCL being unable to select the correct network interface or from incompatible NCCL versions. Explicitly setting NCCL_SOCKET_IFNAME to the primary interface and using a container image with a validated NCCL build ensures that inter-GPU communication initializes correctly. MirroredStrategy does not require TF_CONFIG, so adding workers or TF_CONFIG would not help.

Exam trap

The trap here is assuming that NCCL timeout errors indicate a need for more workers or TF_CONFIG, when in a single-machine multi-GPU setup the issue is typically network interface selection or NCCL version compatibility.

88
Multi-Selectmedium

You are preparing to scale a prototype ML model to production on Vertex AI. The model is trained with a custom training job, and you want to ensure that the training is reproducible and that you can compare different runs. Which two practices should you follow? (Choose two.)

Select 2 answers
A.Use Vertex AI Experiments to log parameters, metrics, and artifacts for each training run.
B.Specify a fixed random seed in your training code and record it along with the training configuration.
C.Always use the latest pre-built training container without pinning the version so that you get the newest features.
D.Store the trained model artifacts in a versioned location such as a Cloud Storage bucket with object versioning or a Vertex AI Model Registry entry.
E.Train on the full dataset without any sampling to ensure that all runs see the same data.
AnswersA, B

Vertex AI Experiments provides a structured way to track parameters, metrics, and artifacts across training runs, making it easy to compare runs and reproduce results. It integrates with Vertex AI Training and helps you maintain a record of what was used for each model, which is a key practice for scaling prototypes to production.

Why this answer

To make training reproducible and comparable, you should control randomness by setting and recording a fixed seed, and you should use Vertex AI Experiments to log parameters, metrics, and artifacts for each run. These two practices together let you reproduce a run and compare it with others. The other options either harm reproducibility by not pinning the environment or do not directly address run comparison.

Exam trap

The trap here is thinking that using the latest container or training on all data guarantees reproducibility, when the real keys are controlling randomness and systematically logging run metadata.

89
MCQmedium

You are scaling a prototype ML model to production on Vertex AI. The model is a TensorFlow model that you want to train on a large dataset using distributed training across multiple nodes. You need to minimize training time and ensure the job can recover from node failures. Which approach should you take?

A.Use Vertex AI Training with a custom training job that uses the TensorFlow DistributionStrategy with CentralStorageStrategy, and enable automatic restart of the job on failure without checkpointing.
B.Use Vertex AI Training with a custom training job that uses the TensorFlow DistributionStrategy with MirroredStrategy across multiple GPUs on a single node, and enable checkpointing to Cloud Storage.
C.Use Vertex AI Training with a custom training job that uses the TensorFlow DistributionStrategy with TPUStrategy on a TPU pod, and rely on TPU's built-in fault tolerance to handle node failures.
D.Use Vertex AI Training with a custom training job that uses the TensorFlow DistributionStrategy with MultiWorkerMirroredStrategy across multiple nodes, and configure the job to save checkpoints to Cloud Storage and resume from the latest checkpoint on failure.
AnswerD

MultiWorkerMirroredStrategy is designed for multi-node, multi-GPU training in TensorFlow. It synchronizes gradients across workers. Saving checkpoints to Cloud Storage and resuming from them allows the job to recover from node failures without restarting from scratch, minimizing training time and ensuring fault tolerance.

Why this answer

MultiWorkerMirroredStrategy enables multi-node, multi-GPU distributed training in TensorFlow. Combined with checkpointing to Cloud Storage and automatic resume, it provides both scalability and fault tolerance. This approach minimizes training time and ensures recovery from node failures.

Exam trap

The trap here is confusing single-node strategies like MirroredStrategy with multi-node strategies like MultiWorkerMirroredStrategy, or assuming that automatic restart without checkpointing is sufficient for fault tolerance.

90
MCQeasy

A team is building a feature pipeline for an ML model. They need to compute aggregate features over a sliding time window from streaming data. Which Google Cloud service is most appropriate for this task?

A.Cloud Dataflow with sliding windows
B.Cloud Pub/Sub for windowing logic
C.BigQuery scheduled queries
D.Cloud Functions with Pub/Sub triggers
AnswerA

Cloud Dataflow supports sliding windows (e.g., via the SlidingWindows transform in Beam), making it suitable for computing aggregates over overlapping time intervals. The option text says 'fixed windows', but the service itself can handle sliding windows.

Why this answer

Cloud Dataflow, based on Apache Beam, natively supports sliding time windows, making it the appropriate service for computing aggregate features over overlapping time intervals from streaming data. Fixed windows alone would not satisfy the sliding-window requirement, and the other services lack native sliding window aggregation capabilities.

Exam trap

Candidates might confuse fixed windows with sliding windows. Cloud Dataflow supports both, but to meet the requirement the option must explicitly specify sliding windows, not fixed windows.

How to eliminate wrong answers

Option B is wrong because Cloud Pub/Sub is a messaging service that handles event ingestion and delivery, not windowing logic or stateful aggregation; it has no built-in capability to compute sliding window aggregates. Option C is wrong because BigQuery scheduled queries operate on batch data in tables, not on streaming data in real time, and they lack the low-latency, per-event windowing needed for a streaming feature pipeline. Option D is wrong because Cloud Functions with Pub/Sub triggers are stateless and ephemeral, with a maximum timeout of 9 minutes (or 60 minutes with 2nd gen), making them unsuitable for maintaining sliding window state or performing continuous aggregation over streaming data.

91
MCQmedium

You are training a model on Vertex AI using a custom training job. The training data is stored in a Cloud Storage bucket in the us-central1 region, and the training job runs in the us-central1 region. You notice that the training job takes significantly longer than expected due to data loading. You want to improve data loading performance without changing the model architecture. What should you do?

A.Copy the training data to the local SSD of the training VM before training begins, and read from there.
B.Convert the training data to TFRecord format and use the tf.data.TFRecordDataset with parallel interleave and prefetching.
C.Enable streaming reads from Cloud Storage by using the tf.data.experimental.make_csv_dataset function with a large buffer size.
D.Use the Vertex AI Training reduction server to accelerate data loading.
AnswerB

TFRecord is a binary format that is more efficient to read than many small files. Using tf.data.TFRecordDataset with parallel interleave and prefetching allows overlapping data loading with model training, significantly improving throughput. This is a recommended practice for optimizing input pipelines on Vertex AI, especially when data is stored in Cloud Storage.

Why this answer

Converting data to TFRecord and using tf.data with parallel interleave and prefetching optimizes the input pipeline by reducing the number of read operations and overlapping data loading with computation. This is a standard best practice for improving training performance on Vertex AI when data is in Cloud Storage.

Exam trap

The trap here is assuming that simply increasing buffer size or copying data locally will solve the data loading bottleneck without addressing the format and parallelization of reads.

92
MCQmedium

You have a TensorFlow model that you want to deploy on Vertex AI for online prediction. The model requires a custom preprocessing step that transforms raw input features into the format expected by the model. You need to ensure that the same preprocessing is applied both during training and serving, and that it is maintained as part of the model artifact. What should you do?

A.Write a separate Python script for preprocessing and include it in the custom container for both training and serving.
B.Use Vertex AI Feature Store to store preprocessed features and serve them during prediction.
C.Perform preprocessing in a Cloud Function that triggers before sending data to the Vertex AI endpoint.
D.Implement the preprocessing as a TensorFlow Transform (tf.Transform) preprocessing function and attach it to the model using a SavedModel signature.
AnswerD

tf.Transform allows you to define preprocessing as part of the TensorFlow graph, ensuring consistency between training and serving. By exporting the transform function with the SavedModel, the preprocessing is encapsulated in the model artifact and automatically applied during online prediction. This is the recommended approach for maintaining preprocessing consistency.

Why this answer

Using tf.Transform to define preprocessing and attaching it to the SavedModel ensures that the same transformation is applied during both training and serving. The preprocessing becomes part of the model artifact, simplifying deployment and maintenance. Other options either risk inconsistency or add unnecessary complexity.

Exam trap

The trap here is thinking that any shared code or external service can guarantee consistency, when the most reliable method is to embed preprocessing directly into the model artifact.

93
Multi-Selecthard

You are fine-tuning a large language model (LLM) from Vertex AI Model Garden using a custom dataset. You need to minimize training cost while maintaining reasonable throughput. Which THREE strategies should you combine?

Select 3 answers
A.Use spot VM instances for training
B.Use parameter-efficient fine-tuning (PEFT) such as LoRA
C.Use full fine-tuning of all model parameters
D.Use TPU v4 pods for training
E.Use mixed precision training (FP16)
AnswersA, B, E

Spot VMs are significantly cheaper than regular VMs and are suitable for fault-tolerant fine-tuning jobs.

Why this answer

Spot VM instances are significantly cheaper than on-demand instances, reducing training cost. They can be preempted, but for fine-tuning tasks that can checkpoint and resume, this trade-off is acceptable for cost savings.

Exam trap

The Google PMLE exam often tests the misconception that higher-performance hardware (like TPU pods) is always the best choice for cost optimization, when in reality, cost-minimization strategies prioritize cheaper compute and efficient training methods over raw throughput.

94
MCQmedium

A machine learning engineer is training a TensorFlow model on Vertex AI using distributed training with the MultiWorkerMirroredStrategy. The training job uses 4 workers with 4 GPUs each. The engineer notices that the training is not scaling linearly. What is the most likely cause?

A.The model architecture is too simple to benefit from distribution
B.The workers are not using the same version of TensorFlow
C.Communication overhead due to gradient synchronization
D.The GPUs are not configured correctly
AnswerC

MultiWorkerMirroredStrategy performs all-reduce gradient synchronisation across workers each step, so inter-worker communication overhead grows with worker count and prevents linear scaling. With 4 workers and 16 GPUs, network latency and bandwidth for exchanging gradients, not compute, become the bottleneck limiting throughput gains.

Why this answer

With MultiWorkerMirroredStrategy, each worker computes gradients independently on its local batch, then all-reduces gradients across workers via collective communication (e.g., NCCL or gRPC). As the number of workers increases, the communication overhead for gradient synchronization grows, often dominating the per-step time and preventing linear scaling. This is the most common bottleneck in distributed TensorFlow training, especially with many workers or small batch sizes per worker.

Exam trap

The trap here is that candidates often assume more workers always means linear speedup, ignoring the fixed overhead of gradient synchronization that becomes the dominant factor in distributed training.

How to eliminate wrong answers

Option A is wrong because even a simple model can suffer from communication overhead if the compute-to-communication ratio is low; the issue is not model simplicity but the cost of synchronizing gradients across workers. Option B is wrong because TensorFlow enforces version consistency across workers in a distributed job; mismatched versions would cause a job failure, not sublinear scaling. Option D is wrong because GPU misconfiguration (e.g., incorrect driver or CUDA version) would typically cause errors or zero utilization, not gradual scaling degradation; the observed symptom of sublinear scaling points to communication, not hardware misconfiguration.

95
MCQhard

A research team is training a very large Transformer model that does not fit into the memory of a single GPU. They have access to multiple GPUs on a single machine and want to split the model layers across GPUs. Which distributed training strategy should they use?

A.MultiWorkerMirroredStrategy
B.Parameter server strategy
C.MirroredStrategy (data parallelism)
D.Pipeline parallelism (model parallelism)
AnswerD

Pipeline parallelism partitions the model's layers across GPUs, so each device holds only a subset of parameters and activations. This directly addresses the constraint that the model exceeds single-GPU memory. Data parallelism would replicate the full model on every GPU, which remains impossible.

Why this answer

Pipeline parallelism (model parallelism) is the correct choice because it splits the model's layers across multiple GPUs, allowing a model too large for a single GPU's memory to be trained. Each GPU holds a subset of layers and processes micro-batches in a pipelined fashion, enabling training of models that exceed single-device memory. This directly addresses the memory constraint by distributing the model parameters, not just the data.

Exam trap

The trap here is confusing data parallelism (which replicates the model) with model parallelism (which splits the model), and overlooking that the problem explicitly states the model does not fit in a single GPU's memory, making any data-parallel strategy incorrect.

How to eliminate wrong answers

Option A is wrong because MultiWorkerMirroredStrategy is designed for multi-worker (multi-machine) synchronous data parallelism, not for splitting a single model across GPUs on one machine. Option B is wrong because parameter server strategy is an asynchronous data-parallel approach where workers compute gradients and a parameter server updates weights; it does not split model layers across devices and is typically used for large-scale distributed training across many machines, not for fitting a single large model on one machine. Option C is wrong because MirroredStrategy (data parallelism) replicates the entire model on each GPU, which fails when the model does not fit in a single GPU's memory.

96
MCQmedium

You have a prototype model trained on a small sample of data. You now want to train on the full dataset using Vertex AI, but the dataset is stored in BigQuery and is several terabytes. You need to minimize data movement and avoid exporting the full dataset to Cloud Storage. What should you do?

A.Export the BigQuery table to Cloud Storage as CSV and pass the Cloud Storage path to the training job.
B.Use the BigQuery reader in the training application to read data directly from BigQuery during training.
C.Create a Dataproc cluster to convert the BigQuery table to Parquet in Cloud Storage, then train from that path.
D.Load the BigQuery table into a Vertex AI Feature Store and train from the feature store online serving endpoint.
AnswerB

Vertex AI custom training supports reading data directly from BigQuery using client libraries or the BigQuery Storage API. This avoids exporting the dataset and keeps data in place, which minimizes movement and storage duplication. The training container can stream or batch records from BigQuery, and you can control the query to select only needed columns or rows, reducing I/O further.

Why this answer

When training data lives in BigQuery and the goal is to avoid moving it, the training application should read from BigQuery directly using the available client libraries or Storage API. This keeps the data in place and avoids exporting terabytes to Cloud Storage. Exporting to CSV or Parquet, or routing through Feature Store, adds unnecessary copies or is not designed for bulk training reads.

Exam trap

The trap here is assuming that training data must always be staged in Cloud Storage, when Vertex AI training can read directly from BigQuery.

97
MCQmedium

You need to run a distributed training job on Vertex AI using TensorFlow with MirroredStrategy on a single machine with 4 GPUs. Which training configuration should you use?

A.Use MirroredStrategy with a single workerPoolSpec containing a machine_type with 4 GPUs
B.Use MultiWorkerMirroredStrategy with multiple workerPools
C.Use MirroredStrategy with two workerPoolSpecs, each with 2 GPUs
D.Use ParameterServerStrategy with a chief and a parameter server
AnswerA

MirroredStrategy replicates variables across all four GPUs on one host and synchronises gradients via NVIDIA NCCL all-reduce, so a single workerPoolSpec with a four-GPU machine_type satisfies the single-machine constraint. Adding extra replicas would introduce multi-worker distribution, which MirroredStrategy does not handle.

Why this answer

MirroredStrategy is TensorFlow's single-machine, multi-GPU strategy that replicates the model across all GPUs on one host using all-reduce for gradient synchronization. On Vertex AI, this maps to a single workerPoolSpec whose machine type has 4 GPUs attached. No additional workers are needed because MirroredStrategy operates within one machine.

Exam trap

PMLE often tests the boundary between single-machine and multi-machine strategies — candidates who see '4 GPUs' and assume 'distributed across machines' incorrectly choose MultiWorkerMirroredStrategy instead of MirroredStrategy on one worker pool.

How to eliminate wrong answers

Option B is wrong because MultiWorkerMirroredStrategy is for multi-machine (multi-node) training and requires multiple worker pools — it is unnecessary and incorrect for a single machine with 4 GPUs. Option C is wrong because MirroredStrategy does not span multiple worker pools; splitting into two workerPoolSpecs would imply two machines, which contradicts the single-machine requirement and would require MultiWorkerMirroredStrategy instead. Option D is wrong because ParameterServerStrategy is a different distributed paradigm (parameter servers + workers) used for asynchronous or large-scale training, not for single-machine mirrored training.

98
MCQhard

An ML engineer is preparing to train a large recommendation model on Vertex AI. The model uses a custom training loop in PyTorch and requires a multi-node cluster with 8 A100 GPUs per node. The engineer wants to minimize training time and ensure the job can recover from a node failure without restarting from scratch. Which combination of Vertex AI features should the engineer use?

A.Use a custom training job with a worker pool of 8 A100 GPUs per replica and multiple replicas, enable distributed training with the PyTorch distributed launcher, and write checkpoints to Cloud Storage periodically.
B.Use Vertex AI Pipelines to orchestrate the training as a series of steps, with each step training on a subset of data, and use pipeline caching to avoid retraining on failure.
C.Use a Vertex AI custom training job with a single replica of 8 A100 GPUs and enable automatic machine restart, then rely on the default checkpointing behavior of the PyTorch training loop.
D.Use a Vertex AI hyperparameter tuning job with a single replica of 8 A100 GPUs, relying on early stopping to reduce training time.
AnswerA

A custom training job with multiple replicas and 8 A100 GPUs each provides the required multi-node cluster. Vertex AI handles the distributed environment variables and network setup, and the PyTorch distributed launcher coordinates ranks across nodes. Periodic checkpoints to Cloud Storage allow the job to resume from the latest checkpoint after a node failure, satisfying both performance and recovery requirements.

Why this answer

The requirement is a multi-node A100 cluster with fault tolerance. A Vertex AI custom training job with multiple replicas, each with 8 A100 GPUs, supplies the cluster, and the PyTorch distributed launcher coordinates training across nodes. Writing checkpoints to Cloud Storage at intervals lets the job resume after a node failure without losing all progress, which is the standard pattern for large distributed training on Vertex AI.

Exam trap

The trap here is confusing orchestration or tuning features with the distributed training runtime, when only a properly configured custom training job with multiple replicas provides the required multi-node GPU cluster.

99
Multi-Selectmedium

You are scaling a prototype ML model to production on Vertex AI. The model is trained with a custom training job and you want to ensure reproducibility and traceability of each training run. Which two practices should you implement? (Choose two.)

Select 2 answers
A.Store the trained model artifacts in a Cloud Storage bucket with a unique timestamped folder for each run.
B.Store the training code in a Git repository and pass the commit hash as a hyperparameter or label to the Vertex AI training job.
C.Use the same machine type and accelerator for every training run to ensure identical hardware.
D.Enable Vertex AI Model Monitoring on the endpoint to detect training-serving skew.
E.Use Vertex AI Experiments to log parameters, metrics, and artifacts for each training run.
AnswersB, E

Recording the Git commit hash with the training job creates a direct link between the code version and the resulting model. This enables reproducibility because you can check out the exact code that produced a model, and it provides traceability for audits. Vertex AI training jobs support labels and hyperparameters, so you can include the commit hash as metadata without affecting the training logic.

Why this answer

To ensure reproducibility and traceability, you need to record the exact code version and the training parameters and metrics. Storing the Git commit hash with the training job links the model to the code, while Vertex AI Experiments logs parameters, metrics, and artifacts for each run. Together, these practices allow you to reproduce any model and trace its lineage.

Model monitoring, hardware consistency, and timestamped folders do not provide the necessary metadata linkage.

Exam trap

The trap here is focusing on operational practices like hardware consistency or artifact organization, which do not capture the code and parameter metadata required for true reproducibility and traceability.

100
MCQhard

You are deploying a model to a Vertex AI Endpoint that must serve predictions with a strict 50 ms latency SLA. The model is a custom container that loads a large model file from Cloud Storage at startup. You notice that the first few predictions after a new deployment are slow, and sometimes the endpoint scales up and the new replicas also have slow first predictions. What should you do to reduce this cold-start latency?

A.Bake the model file into the custom container image so that it does not need to be downloaded from Cloud Storage at startup.
B.Increase the minimum number of replicas and enable autoscaling with a lower target utilization so that replicas are always warm.
C.Configure the endpoint to use a pre-built container instead of a custom container so that the model is loaded more efficiently.
D.Use a larger machine type for the endpoint so that the model file is downloaded faster and the model loads more quickly.
AnswerA

Embedding the model file in the container image removes the Cloud Storage download from the startup path, which is often the dominant cause of cold-start latency. The container image is cached on the host, so new replicas can start faster and serve predictions within the SLA. This is a common technique for latency-sensitive custom containers on Vertex AI.

Why this answer

The most direct way to reduce cold-start latency is to eliminate the need to download the large model file at startup by baking it into the container image. This removes a major source of delay for new replicas. The other options either do not address the startup download, add cost without solving the root cause, or are not applicable to custom models.

Exam trap

The trap here is focusing on autoscaling or machine size when the real cold-start bottleneck is the model file download at container startup.

101
MCQmedium

A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?

A.Cloud Composer
B.Dataproc
C.Cloud Functions
D.Dataflow
AnswerD

Dataflow runs Apache Beam pipelines that read from BigQuery, compute aggregates at scale, and write to Vertex AI Feature Store for both online serving and offline training. Its managed batch processing satisfies the daily large-scale computation requirement without provisioning servers.

Why this answer

Dataflow is Google Cloud's managed Apache Beam service, purpose-built for batch and streaming data pipelines that read from BigQuery, transform data, and write to sinks like Vertex AI Feature Store. It handles autoscaling, sharding, and windowing natively, making it the canonical choice for daily batch feature engineering at scale. Its native BigQuery and Feature Store I/O connectors mean minimal glue code.

Exam trap

PMLE often tests the boundary between orchestration (Cloud Composer/Airflow) and computation (Dataflow/Beam) — candidates who pick Composer because it 'runs pipelines' miss that it does not process the data itself.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is an orchestration service (managed Airflow) — it schedules and coordinates jobs but does not itself perform large-scale data transformation. Option B is wrong because Dataproc is a managed Spark/Hadoop cluster; while it can process BigQuery data, it requires cluster management and is better suited to Spark/Hadoop workloads than Beam-style pipelines. Option C is wrong because Cloud Functions is a serverless, event-driven compute service with short execution timeouts (up to 60 minutes for HTTP, 10 minutes for events) and is unsuitable for large-scale batch aggregation.

102
MCQmedium

Your PyTorch training script uses DistributedDataParallel (DDP) across 4 vertices each with 4 GPUs (16 GPUs total). You submit a Vertex AI custom training job. How should you configure the worker pool spec?

A.Create one worker pool with 4 replicas, each with machine type having 4 GPUs
B.Create a chief worker pool with 1 replica (4 GPUs) and a parameter server pool with 4 replicas (no GPUs)
C.Create 4 separate jobs, each with 1 replica and 4 GPUs
D.Create one worker pool with 16 replicas, each with 1 GPU
AnswerA

This matches the requirement: 4 workers, each with 4 GPUs.

Why this answer

For DDP across multiple machines, use MultiWorkerMirroredStrategy equivalent in PyTorch: set replicas to 4, each with machine type having 4 GPUs. The TF_CONFIG env var is not needed; Vertex AI sets necessary environment variables for distributed training.

103
MCQmedium

You are using TensorFlow Transform (tf.Transform) to preprocess data for a model that will be deployed on Vertex AI. What is the primary benefit of using tf.Transform over Dataflow alone?

A.Support for GPUs during preprocessing
B.Built-in feature store integration
C.Faster data processing
D.Training/serving skew prevention through a consistent transformation graph
AnswerD

tf.Transform records the full preprocessing graph and applies it identically at training and prediction, eliminating training/serving skew. Dataflow alone executes transforms but does not preserve that reusable graph, so the same feature engineering cannot be replayed consistently at serving time on Vertex AI.

Why this answer

tf.Transform computes statistics (e.g., min, max) on the full dataset, then generates a TensorFlow graph that applies the same transformation consistently at training and serving time. Dataflow alone does not ensure this consistency.

← PreviousPage 2 of 2 · 103 questions total

Ready to test yourself?

Try a timed practice session using only Scaling Prototypes into ML Models questions.