Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A company has a TensorFlow model for image classification that must run on edge devices with limited memory. They need to reduce the model size without significant accuracy loss. Which technique should they use?

⚠ Common exam trap

PMLE often tests the distinction between model compression techniques — candidates confuse quantization (reduces precision) with pruning (removes weights) and distillation (trains a smaller model), picking the wrong one for the 'reduce size without retraining' constraint.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Post-training quantization using TensorFlow Lite.

Post-training quantization with TensorFlow Lite converts a trained model's weights from 32-bit floats to 8-bit integers, reducing model size by ~4x and speeding up inference on edge devices with minimal accuracy loss. TensorFlow Lite is purpose-built for edge deployment, making this the most direct and practical technique for the stated constraint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Post-training quantization using TensorFlow Lite.

    Why this is correct

    Post-training quantization converts the trained model's float32 weights to 8-bit integers via TensorFlow Lite, shrinking size roughly fourfold and lowering memory use on constrained edge hardware, with minimal accuracy loss since no retraining is required.

  • ✗

    Knowledge distillation to train a smaller student model.

    Why it's wrong here

    Knowledge distillation requires training a *new*, smaller student model, rather than directly reducing the size of the existing TensorFlow model through modification. While it effectively creates a compact model suitable for edge devices by transferring knowledge from a larger teacher, the scenario implies optimising the *given* model. This option is tempting because it achieves the goal of a smaller, accurate model for resource-constrained environments, and would be correct if the task was to *develop* a new, efficient model based on a pre-existing complex one.

  • ✗

    Pruning the model weights to zero out unimportant connections.

    Why it's wrong here

    Pruning removes individual weights but leaves the dense tensor structure intact, so memory and file size shrink only marginally unless combined with sparse formats and retraining. Quantisation, which stores weights at lower precision, directly reduces model size for edge deployment.

  • ✗

    Use a larger VM for training.

    Why it's wrong here

    A larger training VM increases compute during training only; the exported model's parameter count and file size are unchanged, so edge memory use is unaffected. Larger VMs suit shortening training time or handling bigger datasets, not deployment-time model compression.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.