Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A machine learning engineer is deploying a TensorFlow model on an edge device with limited memory and compute. The model needs to perform inference with low latency. The engineer has a trained float32 model. Which model compression technique should be applied first to reduce the model size and improve inference speed without significant accuracy loss?

⚠ Common exam trap

Google Cloud often tests the misconception that quantization-aware training is always required for INT8 deployment, but the trap here is that post-training quantization is the simplest and most effective first step for reducing model size and latency on edge devices, with quantization-aware training reserved only for cases where accuracy drops below acceptable thresholds.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Post-training quantization to INT8

Post-training quantization to INT8 is the correct first step because it directly reduces the model size by approximately 4x (from 32-bit floats to 8-bit integers) and speeds up inference on edge devices by leveraging integer-optimized hardware (e.g., ARM NEON or Qualcomm Hexagon). This technique requires no retraining and typically yields minimal accuracy loss for most TensorFlow models, making it the fastest path to deploy on resource-constrained devices.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Post-training quantization to INT8

    Why this is correct

    Post-training quantization converts the trained float32 weights and activations to INT8, cutting model size roughly fourfold and accelerating inference on edge hardware with limited memory and compute. It requires no retraining and typically preserves accuracy, making it the appropriate first compression step.

  • ✗

    Knowledge distillation

    Why it's wrong here

    Knowledge distillation trains a smaller student model to mimic a larger teacher model, but this requires a separate training phase and a dataset, which the stem does not provide—the engineer has only a trained float32 model and needs an immediate deployment technique. It is tempting because distillation can reduce model size and latency in scenarios where retraining is feasible, such as when a large, accurate teacher model exists and the deployment environment allows additional training compute.

  • ✗

    Quantization-aware training

    Why it's wrong here

    Quantisation-aware training is not the initial technique to apply here because the engineer already possesses a *trained* float32 model. This method necessitates modifying the training loop or fine-tuning the model to simulate quantisation effects, which is not a post-training compression step. It is tempting, however, as QAT is highly effective for minimising accuracy degradation when quantising, making it the preferred choice if retraining or fine-tuning is feasible and desired for optimal performance.

  • ✗

    Weight pruning

    Why it's wrong here

    Pruning can help but may cause accuracy loss without retraining; quantization is usually preferred first.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

8 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?

easy
  • ✓ A.Vertex AI Model Optimization
  • B.Vertex AI Model Monitoring
  • C.Vertex AI Matching Engine
  • D.Vertex AI Continuous Training

Why A: Vertex AI Model Optimization is the correct feature because it provides model quantization, pruning, and distillation techniques specifically designed to reduce model size and improve inference latency on edge devices. This service applies post-training quantization (e.g., FP32 to INT8) and structured weight pruning to shrink the model footprint while maintaining accuracy within acceptable thresholds, directly addressing the need for efficient deployment on resource-constrained hardware.

Variation 2. You have an edge device with limited compute resources. You need to deploy a deep learning model for real-time inference. Which model compression technique should you apply to reduce the model size and latency with minimal accuracy loss?

hard
  • A.Pruning only
  • ✓ B.Post-training quantization to INT8
  • C.Knowledge distillation only
  • D.Use full precision FP32 to maintain accuracy

Why B: Post-training quantization to INT8 reduces model weights and activations from 32-bit floats to 8-bit integers, cutting model size by roughly 4x and speeding up inference on edge hardware with minimal accuracy loss. It requires no retraining and works with most trained models. This makes it the best fit for limited-compute edge deployment.

Variation 3. A company has a TensorFlow model for image classification that must run on edge devices with limited memory. They need to reduce the model size without significant accuracy loss. Which technique should they use?

medium
  • ✓ A.Post-training quantization using TensorFlow Lite.
  • B.Knowledge distillation to train a smaller student model.
  • C.Pruning the model weights to zero out unimportant connections.
  • D.Use a larger VM for training.

Why A: Post-training quantization with TensorFlow Lite converts a trained model's weights from 32-bit floats to 8-bit integers, reducing model size by ~4x and speeding up inference on edge devices with minimal accuracy loss. TensorFlow Lite is purpose-built for edge deployment, making this the most direct and practical technique for the stated constraint.

Variation 4. An ML team is optimizing an inference model for deployment on edge devices. They need to reduce the model size and improve latency while maintaining accuracy as much as possible. Which two techniques should they use? (Choose TWO.)

medium
  • A.Use a larger pre-trained model as a starting point.
  • ✓ B.Post-training quantization to INT8.
  • C.Use half-precision (FP16) instead of INT8.
  • ✓ D.Apply weight pruning to remove small weights.
  • E.Increase the number of layers in the model.

Why B: Post-training quantization to INT8 reduces model size by converting 32-bit floating-point weights and activations to 8-bit integers, which also speeds up inference on edge devices with integer-optimized hardware. This technique typically maintains accuracy within 1-2% of the original model while significantly lowering memory footprint and latency.

Variation 5. A company is deploying a deep learning model on edge devices with limited storage and computational resources. They need to reduce the model size by 80% while maintaining acceptable accuracy. Which two techniques should they combine?

medium
  • A.Pruning and distillation
  • B.Quantization and distillation
  • ✓ C.Quantization and pruning

Why C: Quantization reduces model size by converting weights from FP32 to INT8, achieving up to 75% size reduction. Pruning removes redundant weights (e.g., 50% pruning) to further shrink the model. Together, these two techniques can reduce model size by over 80% (e.g., 50% pruning + INT8 quantization yields ~87.5% reduction) while preserving acceptable accuracy on edge devices. Knowledge distillation also reduces size but typically requires a pre-trained teacher and may not achieve the extreme compression needed when combined with quantization alone.

Variation 6. An ML engineer is optimizing a large model for deployment on Vertex AI with GPU acceleration. They want to reduce model size and improve inference latency without significant accuracy loss. Which tool should they use?

hard
  • A.Use gcloud CLI to prune the model.
  • B.Use Cloud TPU for faster inference.
  • ✓ C.Use Vertex AI Model Optimization with TensorRT.
  • D.Use TensorFlow.js converter to optimize the model for web.

Why C: Vertex AI Model Optimization with TensorRT is specifically designed to reduce model size and improve inference latency on NVIDIA GPUs by applying techniques like quantization, pruning, and graph optimizations. TensorRT optimizes the model for the target GPU architecture, enabling faster inference with minimal accuracy loss, which directly addresses the engineer's goals.

Variation 7. You are deploying a deep learning model on edge devices with limited computational resources. The model must run inference in <10 ms and the model size must be under 50 MB. Currently, your trained model is 200 MB and runs in 50 ms. Which combination of model compression techniques should you apply?

hard
  • A.Only apply weight pruning
  • B.Apply quantization-aware training and knowledge distillation
  • C.Use knowledge distillation to train a smaller student model
  • ✓ D.Apply post-training quantization (INT8) and pruning

Why D: Post-training quantization (INT8) reduces model size ~4x and speeds inference by using 8-bit integers instead of 32-bit floats, while pruning removes redundant weights to further shrink the model and reduce compute. Together they can bring a 200 MB / 50 ms model under 50 MB and under 10 ms with minimal accuracy loss, making them the most practical combination for edge deployment. Quantization-aware training and knowledge distillation are heavier lifts that may not be necessary to hit these targets.

Variation 8. A company is fine-tuning a large language model (Gemma 7B) using Vertex AI JumpStart. They want to reduce the model's memory footprint for deployment on edge devices. Which THREE model compression techniques should they consider?

medium
  • ✓ A.Post-training quantization to INT8
  • B.Using a larger batch size during inference
  • ✓ C.Quantization-aware training
  • D.Increasing the number of layers
  • ✓ E.Pruning of weights with small magnitude

Why A: Post-training quantization to INT8 (A) is correct because it converts the model's FP32/FP16 weights to 8-bit integers after training, cutting memory usage roughly 2-4x with minimal accuracy loss, which is ideal for edge deployment. Quantization-aware training (C) is correct because it simulates quantization during fine-tuning so the Gemma 7B model learns to compensate for reduced precision, yielding better accuracy than post-training quantization alone at the same low-bit footprint. Pruning of weights with small magnitude (E) is correct because removing near-zero weights reduces parameter count and model size, directly lowering the memory footprint on edge devices. Using a larger batch size during inference (B) is incorrect because batch size affects throughput and activation memory, not the model's stored weight footprint, and larger batches actually increase memory use. Increasing the number of layers (D) is incorrect because adding layers enlarges the model and increases memory requirements, the opposite of compression.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.