You are deploying a deep learning model on edge devices with limited computational resources. The model must run inference in <10 ms and the model size must be under 50 MB. Currently, your trained model is 200 MB and runs in 50 ms. Which combination of model compression techniques should you apply?
Trap 1: Only apply weight pruning
Pruning alone may not achieve necessary size reduction; quantization is more impactful.
Trap 2: Apply quantization-aware training and knowledge distillation
These require retraining; not the fastest approach if you want to avoid retraining.
Trap 3: Use knowledge distillation to train a smaller student model
Knowledge distillation alone reduces model size by training a smaller student network, but it does not guarantee the student will meet the <10 ms inference latency or the <50 MB size constraint unless the student architecture is explicitly designed for those limits; the stem’s 200 MB, 50 ms model requires additional techniques like pruning or quantisation to achieve both targets. It is tempting because distillation effectively transfers accuracy from a large teacher to a compact student, and would be correct if the only requirement were reducing model size while preserving performance, without strict latency or size thresholds.
- A
Only apply weight pruning
Why it fails: Pruning alone may not achieve necessary size reduction; quantization is more impactful.
- B
Apply quantization-aware training and knowledge distillation
Why it fails: These require retraining; not the fastest approach if you want to avoid retraining.
- C
Use knowledge distillation to train a smaller student model
Why it fails: Knowledge distillation alone reduces model size by training a smaller student network, but it does not guarantee the student will meet the <10 ms inference latency or the <50 MB size constraint unless the student architecture is explicitly designed for those limits; the stem’s 200 MB, 50 ms model requires additional techniques like pruning or quantisation to achieve both targets. It is tempting because distillation effectively transfers accuracy from a large teacher to a compact student, and would be correct if the only requirement were reducing model size while preserving performance, without strict latency or size thresholds.
- D
Apply post-training quantization (INT8) and pruning
Quantization reduces size and latency; pruning reduces complexity; both can be applied post-training.