Courseiva

PMLE · topic practice

Scaling Prototypes into ML Models practice questions

This domain covers turning a working prototype into a production-grade ML system on Google Cloud: distributed training on Vertex AI, tuning and serving large models, and packaging preprocessing with the model. Questions present concrete scenarios (custom training jobs, Model Garden fine-tuning, Vertex AI Prediction) and ask you to choose configurations and strategies that balance cost, throughput, and correctness.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Scaling Prototypes into ML Models

What the exam tests

What to know about Scaling Prototypes into ML Models

Be able to translate a training or serving scenario into concrete Vertex AI configuration: correct worker pool specs for distributed training, cost-aware fine-tuning choices, and packaging preprocessing with the deployed model. The single most important thing is matching replica counts and resources to the actual distributed topology.

Configuring Vertex AI custom training worker pools, replica counts, and machine specs for DistributedDataParallel jobs

Using Vertex AI Model Garden and Vertex AI Training for fine-tuning and adapting pre-trained models

Reducing LLM fine-tuning cost via parameter-efficient tuning, smaller machines, and checkpointing strategies

Deploying on Vertex AI Prediction with custom containers and embedding preprocessing in the serving graph

Watch out for

Common Scaling Prototypes into ML Models exam traps

  • ▸Setting the worker pool replica count to the number of GPUs instead of the number of machines, which breaks DDP process groups
  • ▸Reimplementing preprocessing in the client instead of embedding it in the model or serving container, causing training-serving skew
  • ▸Assuming a pre-trained model can classify new categories without retraining the head or fine-tuning on labeled data

Practice set

Scaling Prototypes into ML Models questions

20 questions · select your answer, then reveal the explanation

You are deploying a deep learning model on edge devices with limited computational resources. The model must run inference in <10 ms and the model size must be under 50 MB. Currently, your trained model is 200 MB and runs in 50 ms. Which combination of model compression techniques should you apply?

You are running a Vertex AI custom training job with pre-built TensorFlow container. You want to use TPU v3 pods for faster training. Which configuration is required?

You are fine-tuning a large language model using Vertex AI Training with spot VMs to reduce cost. Your training job keeps getting preempted, causing delays. Which THREE strategies can help mitigate the impact of preemption?

You are building a machine learning pipeline on Google Cloud. You need to perform feature engineering on large datasets stored in BigQuery and store the resulting features in Vertex AI Feature Store for both online and offline use. Which TWO Google Cloud services should you use?

A data scientist wants to train a TensorFlow model on Vertex AI using a pre-built container. Which of the following pre-built containers is NOT available for custom training in Vertex AI?

You are performing hyperparameter tuning on Vertex AI with Vizier. You want to maximize the accuracy of your model, and you have a budget of 50 trials. Which algorithm should you choose to best explore the search space?

You are fine-tuning a pre-trained BERT model from Hugging Face for a sentiment analysis task using Vertex AI training. The dataset has 100k examples. To avoid catastrophic forgetting, which layer freezing strategy should you apply?

Which Vertex AI service allows you to discover, fine-tune, and deploy foundation models with a few clicks, including models like Llama and Gemma?

You want to deploy a trained scikit-learn model to Vertex AI for online predictions. The model file is 2 GB. Which option should you use?

You need to reduce the cost of training a large model on Vertex AI while maintaining fault tolerance. Which THREE actions should you take? (Choose 3)

You are fine-tuning a Gemma model using Vertex AI JumpStart. You want to combine the fine-tuned model with a custom output layer for a unique task. Which TWO components are required to deploy the combined model? (Choose 2)

A machine learning engineer is preparing to train a Transformer-based model using TensorFlow on a single TPU v3-8 pod slice. The training script uses tf.distribute.TPUStrategy. Which environment variable must be set in Vertex AI to enable TPU training with the appropriate topology?

A team wants to perform hyperparameter tuning on a Vertex AI custom training job with 100 trials. They require an algorithm that efficiently explores the search space by learning from previous trials. Which algorithm should they select in the study configuration?

A company wants to bring their own Docker container to Vertex AI for training a model with a custom framework. They need to ensure the container is compatible with the Vertex AI training service. What is the minimum requirement for the container?

A data scientist wants to quickly experiment with a pre-trained Vision Transformer model from Hugging Face and fine-tune it on a custom dataset using Vertex AI. They want to use a managed environment with minimal setup. Which Vertex AI service should they use?

A company is deploying a deep learning model on edge devices with limited storage and computational resources. They need to reduce the model size by 80% while maintaining acceptable accuracy. Which two techniques should they combine?

A company is fine-tuning a large language model (Gemma 7B) using Vertex AI JumpStart. They want to reduce the model's memory footprint for deployment on edge devices. Which THREE model compression techniques should they consider?

Your team is deploying a large language model (LLM) on Vertex AI for online prediction. The model exceeds the maximum request size for Vertex AI Prediction. Which approach should you take to serve this model?

An ML engineer wants to use Vertex AI Model Garden to deploy a pre-trained foundation model for text summarisation. What is the quickest way to achieve this?

You are performing post-training quantisation of a trained TensorFlow model to INT8 for deployment on edge devices. Which technique should you use to minimise accuracy loss?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Scaling Prototypes into ML Models sessions

Start a Scaling Prototypes into ML Models only practice session

Every question in these sessions is drawn from the Scaling Prototypes into ML Models domain — nothing else.

Related practice questions

Related PMLE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PMLE exam test about Scaling Prototypes into ML Models?
Be able to translate a training or serving scenario into concrete Vertex AI configuration: correct worker pool specs for distributed training, cost-aware fine-tuning choices, and packaging preprocessing with the deployed model. The single most important thing is matching replica counts and resources to the actual distributed topology.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Scaling Prototypes into ML Models questions in a focused session?
Yes — the session launcher on this page draws every question from the Scaling Prototypes into ML Models domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PMLE topics?
Use the topic links above to move to related areas, or go back to the PMLE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PMLE exam covers. They are not copied from any real exam or dump site.