Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You have an edge device with limited compute resources. You need to deploy a deep learning model for real-time inference. Which model compression technique should you apply to reduce the model size and latency with minimal accuracy loss?

⚠ Common exam trap

The trap is assuming pruning or distillation alone delivers the same size/latency reduction as quantization, when quantization is the fastest, most reliable post-training compression for edge inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Post-training quantization to INT8

Post-training quantization to INT8 reduces model weights and activations from 32-bit floats to 8-bit integers, cutting model size by roughly 4x and speeding up inference on edge hardware with minimal accuracy loss. It requires no retraining and works with most trained models. This makes it the best fit for limited-compute edge deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pruning only

    Why it's wrong here

    Pruning alone removes weights but leaves the architecture and precision unchanged, so it cannot deliver the size and latency reduction this edge scenario demands. It is tempting because pruning is a genuine compression technique, and it would be the correct choice when the requirement is to reduce redundant parameters with minimal accuracy impact.

  • ✓

    Post-training quantization to INT8

    Why this is correct

    Post-training quantization to INT8 converts trained weights and activations from FP32 to 8-bit integers, cutting model size roughly fourfold and accelerating inference on constrained edge hardware. It requires no retraining, satisfying the limited-compute constraint, while typically preserving accuracy with minimal degradation for real-time deployment.

  • ✗

    Knowledge distillation only

    Why it's wrong here

    Knowledge distillation alone transfers knowledge to a smaller model but does not compress the original architecture or guarantee the latency reduction this edge scenario requires. It is tempting because distillation is a recognised compression technique, and it would be the correct choice when a compact student model must match a larger teacher's accuracy.

  • ✗

    Use full precision FP32 to maintain accuracy

    Why it's wrong here

    Full precision FP32 preserves accuracy but increases model size and inference latency, directly contradicting the edge device's limited compute requirement. It is tempting because FP32 is the default training precision, and it would be the correct choice when accuracy is paramount and hardware resources are unconstrained.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.