PMLE Scaling Prototypes into ML Models Practice Question
You are deploying a pre-trained BERT model for inference on edge devices. The model must be under 500 MB and inference latency under 50 ms. Which approach should you take?
⚠ Common exam trap
A common pitfall is assuming that only pruning or distillation can reduce model size, but post-training INT8 quantization directly shrinks the model and speeds inference without retraining, which is ideal for deploying a pre-trained model on edge devices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply post-training INT8 quantization using TensorFlow Lite
Post-training INT8 quantization reduces model size by approximately 75% (from ~440 MB to ~110 MB for BERT-Base) and accelerates inference on edge devices via integer arithmetic, easily meeting the 500 MB and 50 ms constraints. TensorFlow Lite provides hardware-optimized kernels for ARM CPUs and NPUs, making it ideal for edge deployment without requiring retraining.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger model like BERT-Large and deploy on GPU
Why it's wrong here
BERT-Large exceeds the 500 MB budget and GPU deployment contradicts edge hardware constraints, so both size and latency targets fail. Larger GPU-hosted models suit server-side inference where memory and accelerators are plentiful, not constrained edge devices.
- ✓
Apply post-training INT8 quantization using TensorFlow Lite
Why this is correct
Post-training INT8 quantization shrinks the BERT weights roughly fourfold, bringing the model comfortably under the 500 MB ceiling, while integer arithmetic accelerates inference to meet the 50 ms latency budget. TensorFlow Lite's runtime is purpose-built for edge deployment, so no retraining is required.
- ✗
Prune 50% of the model weights and fine-tune
Why it's wrong here
Pruning 50% of weights reduces size but unstructured sparsity rarely shrinks the dense tensor footprint below 500 MB or guarantees sub-50 ms latency without hardware support. Pruning suits compressing models where sparse execution is available, not generic edge deployment.
- ✗
Use knowledge distillation to train a smaller student model from scratch
Why it's wrong here
Training a student from scratch discards the pre-trained BERT weights the stem supplies, costing far more compute and risking accuracy loss. Distillation is correct when no suitable pre-trained model exists and training budget permits, not when a pre-trained model must simply be compressed.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.