Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are deploying a pre-trained BERT model for inference on edge devices. The model must be under 500 MB and inference latency under 50 ms. Which approach should you take?

⚠ Common exam trap

A common pitfall is assuming that only pruning or distillation can reduce model size, but post-training INT8 quantization directly shrinks the model and speeds inference without retraining, which is ideal for deploying a pre-trained model on edge devices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply post-training INT8 quantization using TensorFlow Lite

Post-training INT8 quantization reduces model size by approximately 75% (from ~440 MB to ~110 MB for BERT-Base) and accelerates inference on edge devices via integer arithmetic, easily meeting the 500 MB and 50 ms constraints. TensorFlow Lite provides hardware-optimized kernels for ARM CPUs and NPUs, making it ideal for edge deployment without requiring retraining.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a larger model like BERT-Large and deploy on GPU

    Why it's wrong here

    BERT-Large exceeds the 500 MB budget and GPU deployment contradicts edge hardware constraints, so both size and latency targets fail. Larger GPU-hosted models suit server-side inference where memory and accelerators are plentiful, not constrained edge devices.

  • ✓

    Apply post-training INT8 quantization using TensorFlow Lite

    Why this is correct

    Post-training INT8 quantization shrinks the BERT weights roughly fourfold, bringing the model comfortably under the 500 MB ceiling, while integer arithmetic accelerates inference to meet the 50 ms latency budget. TensorFlow Lite's runtime is purpose-built for edge deployment, so no retraining is required.

  • ✗

    Prune 50% of the model weights and fine-tune

    Why it's wrong here

    Pruning 50% of weights reduces size but unstructured sparsity rarely shrinks the dense tensor footprint below 500 MB or guarantees sub-50 ms latency without hardware support. Pruning suits compressing models where sparse execution is available, not generic edge deployment.

  • ✗

    Use knowledge distillation to train a smaller student model from scratch

    Why it's wrong here

    Training a student from scratch discards the pre-trained BERT weights the stem supplies, costing far more compute and risking accuracy loss. Distillation is correct when no suitable pre-trained model exists and training budget permits, not when a pre-trained model must simply be compressed.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.